Build an evaluation set with a normal completion, missing source, conflicting source, denied authority, provider timeout, duplicate attempt, partial action, corrected evidence, and unverified outcome. Reviewers should state the strongest supported status and the next safe action. Compare their answers with the contract. Disagreement reveals ambiguous semantics or inadequate views and should become implementation work rather than being averaged away.
Measure retrieval time, missing links, false completion, access exceptions, correction propagation, and reviewer confidence grounded in specific evidence. Do not use confidence alone as proof of quality; a clear but wrong trace can create high confidence. Case sampling should include high-consequence work even if it is rare. The evaluation report names limitations, unresolved owners, and the evidence required before wider authority.
Repeat the evaluation after source, policy, role, or provider changes. A pattern can pass at launch and drift as the operating environment evolves. Versioned test cases reveal whether historical records remain interpretable and whether new cases still stop safely when a required relationship is missing. This turns the pattern into a maintained review instrument rather than a one-time demonstration.