Operational signals can include retrieval latency, ingestion lag, source coverage, empty-result rate, stale-source use, permission denials, unsupported claims, abstentions, reviewer overrides, correction time, model and retrieval cost, and user-reported issues. Segment these by workflow and source collection. A global average can hide a high-risk domain with poor evidence.
Cost should be interpreted with quality and consequence. A smaller context window may reduce model expense while excluding a necessary exception. Aggressive refresh may improve freshness while increasing connector load and provider cost. Extensive human review may prevent harm but eliminate the expected operating benefit. Teams should define the service level and evidence threshold for each workflow, then observe tradeoffs. There is no universal optimum independent of the decision, source volatility, and cost of being wrong.
Value measures should return to the decision, such as reduced repeated searching or faster access to an approved source. Establish a baseline and account for review and maintenance work. A higher answer rate may be harmful if the system stops abstaining. Teams should not claim productivity, accuracy, or return on investment from a deployment without comparable evidence and appropriate attribution.