Trials should include missing data, contradictory instructions, revoked access, tool timeout, duplicate events, partial completion, unsafe requests, and a provider outage. Observe whether the system refuses, retries, escalates, or continues silently. Test idempotency where repeated execution could create duplicate messages, orders, records, or charges. The goal is not to eliminate every error; it is to make failure bounded, visible, and recoverable.
Measure completion quality, unsupported claims, approval burden, exception rate, recovery time, cost per outcome, and operator effort. Use the same cases across finalists and preserve the evidence. A platform may produce fluent outputs while failing the operational test, or it may be reliable but require more setup. Both findings matter. Avoid declaring a winner from a curated vendor scenario that does not resemble the target workflow.
Run the trial long enough to encounter routine variation, but keep its authority bounded. Use representative synthetic or approved data where production access would create unnecessary exposure, and require human approval before external or irreversible actions. A trial cannot prove every future reliability condition. It can reveal failure patterns, operating effort, and evidence gaps that a polished demonstration conceals, giving the buyer a stronger basis for a limited rollout or a decision to stop.