API Quality at Scale
Checkout verification across 50–100 merchant sites with Newman automation — for the infrastructure layer of agentic commerce.
API quality engineering for agentic checkout infrastructure.
The client is a universal agentic checkout infrastructure platform that solves a core unsolved problem in AI-driven commerce: AI agents can browse and recommend products, but consistently fail at checkout due to merchant fraud-detection systems that block automated transactions. Its Universal Checkout API breaks through this barrier with a simple contract — provide a product URL and payment token, get back true landed costs and a completed order. The system uses AI browser automation with fraud-mitigation techniques, caching successful checkout flows as deterministic workflows for speed. It achieves 90%+ order reliability on browser-automated flows, sub-35-second offer resolution, 10-second checkout latency, and 99.9% API uptime — working across any web store without requiring merchant integration. Beyond checkout, the platform includes a Product Data API, affiliate commissions, and integrations with leading AI agent and payment ecosystems — positioning it as commerce infrastructure for the agentic web. It is PCI DSS scope-reducing, tokenised-payment-only, and fully security-compliant.
A QA practice that could validate thousands of executions per quarter and scale down gracefully as the platform matured.
The numbers.
Test Cases Created
Q3 Executions
Sites Covered
Defect Closure Rate
Checkout, at the intersection of AI, payments, and live sites.
Testing this checkout infrastructure presents challenges unlike conventional web application QA — the product sits at the intersection of AI automation, payment processing, and live merchant websites, each introducing its own layer of complexity.
Scale: 50–100 Live Merchant Sites
Checkout must work reliably across 50–100 different merchant websites — each with unique page structures, authentication flows, fraud detection thresholds, and checkout behaviours.
Fraud Detection Sensitivity
Because checkout is automated on live sites, testing must be carefully managed to avoid triggering fraud detection or rate-limiting systems, adding constraints on how frequently scenarios can be validated.
Dual-Environment Validation
Every regression cycle requires parallel validation across both Production and Staging. A single session covering 98 sites × 3 scenarios × 2 environments generates 588 executions.
API Coverage Breadth
The API surface spans checkout flows, pricing calculations, tax and shipping accuracy, variant handling, payment processing, and authentication — requiring systematic test design across both POST and GET request types.
From API test design to multi-site automation.
The team built QA capability from the ground up — starting with API test case design, then scaling into automated multi-site regression, and maintaining continuous checkout verification as the product evolved.
API Test Case Design (500+ Cases)
Built a comprehensive API test case library from scratch — 500+ structured cases covering checkout, pricing, tax, shipping, variant handling, payment processing, and authentication flows, across both POST and GET request types.
Newman Automation for Multi-Site Regression
Built automated checkout verification using Newman — the Postman CLI runner — to execute regression across 50–100 merchant sites, generating 588 executions per run without manual effort.
Dual-Environment Checkout Verification
Every regression cycle was executed across both Production and Staging in parallel — validating checkout behaviour, pricing accuracy, and API responses were consistent across environments.
Automation Feature Matrix
Maintained a structured feature matrix tracking 115 product features — automation readiness, coverage status, and prioritisation — giving engineering full visibility into automated vs. pending flows.
Checkout Monitoring & Verification
As the engagement moved to on-demand, maintained daily and weekly checkout verification monitoring — identifying pricing discrepancies and raising defects with clear reproduction evidence.
New Feature & Tool Testing
Beyond regression, tested new product capabilities as they launched — including a new AI-agent tool integration, where 16 bugs were identified and 13 retested and resolved.
What the work delivered.
Newman automation built from scratch
covering 50–100 merchant sites per regression cycle across Production and Staging, generating 588 executions per run and eliminating manual overhead.
10,108 test executions delivered in Q3 alone
across 98 sites, 3 checkout scenarios, and 2 environments — demonstrating the scale of coverage the automated model enables within a reduced 10-hour weekly allocation.
500+ API test cases built, covering the full checkout surface: pricing, tax, shipping, variant handling, payment processing, and authentication flows across both POST and GET request types.
Product stability confirmed through QA data
defect volume dropped from 8 in the foundation quarter to just 2 in Q3, with 100% closure rate across all quarters.
Automation Feature Matrix maintained across 115 features
giving the engineering team a clear, up-to-date view of automation coverage, readiness, and prioritisation at all times.
New product areas validated on launch
a new AI-agent tool integration tested with 16 bugs identified and 13 resolved, extending QE coverage into the platform's expanding agentic product surface.
QA process matured alongside the product
the engagement scaled down naturally from 20h/week to 10h/week to on-demand as the product stabilised, without ever compromising coverage quality.
The stack.
Testing checkout reliability across 50–100 live merchant sites at scale required building bespoke Newman automation from scratch — not just writing test cases. The result was a QA practice that confirmed product stability with data, without ever compromising coverage quality.