Choose representative task classes and compare eligible routes using the same acceptance tests, source packets, and review rubric. Include privacy-restricted work, long context, structured output, tool use, provider timeout, quota exhaustion, schema failure, and an expensive retry path. Verify that hard exclusions hold, cumulative budgets are enforced, fallback is explainable, and old receipts retain the catalog and pricing references that governed their decisions.
Monitor eligibility failures, quote variance, accepted quality, retry-adjusted exposure, fallback frequency, latency, evaluator disagreement, human-review effort, and unresolved settlement. These indicators should lead to owned experiments and policy reviews, not automatic claims of savings or superiority. Revalidate after model, provider, prompt, tool, context, or pricing changes because a previous result does not certify a new route version. Keep failed comparisons and rollback decisions so future teams do not repeat an uneconomic experiment after its headline result has been forgotten.