Comparing automated vs manual purchase order testing: reliability and coverage tradeoffs

I’m evaluating our testing strategy for purchase order workflows and would love to hear experiences from teams using both automated (Selenium) and manual testing approaches. We currently do 70% automated, 30% manual, but I’m questioning if this balance is optimal.

Our automated tests provide excellent regression coverage - we can run 500+ test cases in 2 hours. However, they miss edge cases that manual testers catch, like subtle UI inconsistencies or workflow logic errors that only appear with specific data combinations.

The script maintenance challenges are significant. Every D365 update breaks 10-15% of our automated tests due to UI changes or new validation rules. We spend 2-3 days after each update fixing scripts.

What’s your team’s approach? How do you balance automated vs manual test tradeoffs? Are there specific scenarios where one approach clearly outperforms the other?

Automated vs Manual PO Testing: Tradeoffs in D365 Environments

Your 10–15% script breakage per update is a known pain point with Selenium-based UI automation against D365 — DOM changes, control rendering shifts, and new validation layers in each wave release compound this. Before reassessing your 70/30 split, consider what each bucket is actually covering.

Criteria Comparison

Criterion Automated (Selenium/UI) Manual
Regression depth High — 500+ cases at scale Low — human bandwidth limits repeat coverage
Edge case detection Low — scripts follow happy paths and predefined data High — testers adapt to anomalies in real time
UI/UX inconsistency detection Poor — assertions check elements, not usability Strong — humans notice layout drift, accessibility gaps
Data-combination variance Limited without parameterization investment Natural — testers improvise data scenarios
Maintenance cost High per D365 wave update Low — no code to fix
Execution speed Fast — hours for full suite Slow — days for equivalent coverage
Workflow logic validation Moderate — depends on test design rigor High — exploratory testing surfaces logic gaps
Audit trail/repeatability Strong — scripts are versioned artifacts Weak — relies on tester documentation discipline

Where Automation Clearly Wins

Regression suites for PO creation, three-way match, approval routing, and posting flows — anything stable and high-volume. Automate at the API/OData layer rather than UI where possible (verify in your version). D365 F&O exposes PO entities via OData that are far more stable than UI controls between updates. This directly addresses your maintenance cost problem.

Where Manual Clearly Wins

  • Exploratory testing around new D365 feature waves before scripts are written
  • Vendor-specific data edge cases — unusual currency/tax combinations, partial receipts against blanket POs
  • Cross-module workflow validation — PO to AP invoice to payment where state transitions cross boundaries
  • Post-update smoke testing before re-running the full automated suite

Reducing Your Maintenance Burden

Rather than re-balancing percentages, restructure your automation architecture:

  • Shift UI automation to Power Automate Desktop or the EasyRepro framework (verify in your version), which tracks D365 control IDs more reliably than raw Selenium XPath selectors
  • Parameterize data-combination coverage using data-driven test patterns rather than duplicating scripts
  • Tag tests by stability tier — stable core flows vs. update-sensitive tests — so post-wave triage is scoped, not a full 2–3 day remediation

Your 70/30 ratio isn’t inherently wrong; the issue is likely Selenium’s fragility against a SaaS release cadence, not the automation ratio itself.

Ultimately, optimal balance depends on context / your requirements — release frequency, team size, regulatory audit needs, and how often your PO workflow configuration changes.


This draft is based on general Microsoft Dynamics 365 knowledge. It has not been verified against your specific version and environment. Practitioners: verify the steps and share your experience below.

We use 85% automated for regression and 15% manual for exploratory testing. The key is having a robust test data management strategy. Most script maintenance issues come from hardcoded test data that becomes invalid after updates. We use API-driven test data creation which reduces maintenance significantly. Our scripts generate fresh purchase orders for each test run, so we’re not dependent on pre-existing data states.

I’d argue for more manual testing, especially for purchase order workflows with complex approval chains. Automated tests validate the happy path well but miss the nuanced business logic errors. For example, we caught a critical bug where purchase orders over $50K were routing to the wrong approval queue - the automated tests passed because they only checked that an approval was triggered, not which approval path was used.

The maintenance burden is real. We’ve moved to a hybrid approach where critical happy paths are automated with high investment in maintainability (page object models, abstraction layers), while complex edge cases and new features go through manual exploratory testing first. Once a manual test case has been run successfully 3-4 times across updates, we automate it. This reduces wasted effort on automating scenarios that change frequently.

The data management point is interesting. We do have issues with test data becoming stale. How do you handle scenarios that require specific vendor configurations or approval hierarchies? Creating those via API before each test seems time-consuming. Do you maintain a separate test environment with stable master data, or recreate everything dynamically?

We maintain baseline master data (vendors, approval hierarchies, chart of accounts) in the test environment that’s reset weekly. Transactional data (purchase orders, receipts, invoices) is created dynamically via API at test runtime. This gives us stability for configuration-dependent tests while ensuring transaction tests don’t interfere with each other. The weekly reset catches any data drift issues before they accumulate.

After implementing testing strategies across multiple D365 Finance implementations, I’ve found the optimal balance depends heavily on your specific context, but I can share patterns that work well:

Automated vs Manual Test Tradeoffs:

The 70/30 split you mentioned is actually quite reasonable, but the key is which tests fall into each category. Here’s what I recommend:

Automate:

  • Regression tests for stable core workflows (PO creation, approval routing, receipt posting)
  • Integration points with external systems (EDI, supplier portals)
  • Data validation rules that rarely change
  • Performance and load testing scenarios
  • Smoke tests for deployment verification

Manual testing for:

  • New features in their first 2-3 release cycles
  • Complex business logic with multiple conditional branches
  • UI/UX validation and accessibility testing
  • Exploratory testing for edge cases
  • Scenarios requiring human judgment (approval reasonableness, vendor selection logic)

Regression Coverage:

Your 500+ automated test cases in 2 hours is impressive, but coverage isn’t just about quantity. I’ve seen teams with 1000+ tests that miss critical bugs because they focus on happy paths. Consider:

  1. Risk-based prioritization: Automate high-risk, high-frequency workflows first. A purchase order approval bug affects every PO; a rarely-used report formatting issue doesn’t warrant automation investment.

  2. Business process coverage: Map your tests to actual business processes. Ensure you’re testing complete end-to-end workflows (requisition → PO → receipt → invoice → payment) not just individual transactions.

  3. Negative testing: 30% of your automated tests should verify error handling - invalid amounts, missing approvals, duplicate PO numbers, etc. These catch regression bugs that break validation logic.

Script Maintenance Challenges:

The 10-15% breakage rate after updates is actually lower than industry average (typically 20-30% for UI automation). However, you can reduce it further:

  1. Abstraction layers: Implement page object models with business-focused method names. When D365 changes a field label or ID, you update one place, not 50 test scripts.

  2. Resilient selectors: Use data-testid attributes or stable XPath expressions that survive UI changes. Avoid brittle CSS selectors based on generated class names.

  3. API-first approach: For setup and teardown, use D365 Data Management APIs rather than UI automation. Creating test vendors via API is faster and less brittle than clicking through forms.

  4. Version-aware tests: Tag tests with the D365 version they’re designed for. When upgrading, run version-specific test suites to identify breaking changes systematically.

  5. Maintenance budget: Allocate 20% of your automation team’s time specifically for test maintenance. This prevents technical debt accumulation.

Practical Recommendation:

For purchase order workflows specifically, I suggest:

  • 80% automated for standard PO scenarios (various amounts, vendors, approval paths)
  • 15% manual exploratory testing for complex multi-line POs with mixed item types, partial receipts, change orders
  • 5% manual for UI validation and accessibility

Invest heavily in automating the “boring but critical” scenarios - single-line POs, standard approval routing, simple receipts. These provide stable regression coverage. Use manual testing where human judgment adds value - complex procurement scenarios, vendor selection logic, approval reasonableness.

The maintenance challenge is unavoidable but manageable with proper architecture. Treat your test automation as a software product that needs refactoring and technical debt management, not just a collection of scripts.

Finally, measure what matters: defect escape rate (bugs found in production), test execution time, and time to fix broken tests after updates. These metrics guide whether your automation investment is paying off better than pure optimization of coverage percentages.