Record queue wait, execution time, model and tool route, errors, retries, context size, cache use, provider cost, reviewer correction, and terminal outcome. Predict expected cost, latency, quality, and value before the run where proportionate, then compare with the actual result. The objective is not maximal telemetry. It is enough evidence to diagnose failure, reconcile cost, and decide whether the workflow should expand, change, or stop.
Create dashboards and alerts around business state and service health. A queue full of completed tasks can still hide unresolved customer cases. Define support and incident ownership, data retention, model-change review, provider fallback, and budget thresholds. Learning should update a concrete control such as routing, evidence requirements, prompt version, cache policy, or review threshold. If a lesson has no owner or next decision, it is documentation rather than a closed operating loop.
Keep measurement definitions stable across versions. If completion, correction, or cost changes meaning when a new worker is introduced, trend comparisons become misleading. Version material metric definitions and preserve denominators, exclusions, and sampling notes. Use qualitative reviewer evidence alongside aggregate rates, especially during small pilots where a percentage can imply more certainty than the number of observed cases supports.
Review telemetry access and retention as part of the data design. Diagnostic value does not justify copying raw customer content into every log. Separate operational identifiers and metrics from sensitive payloads, and provide a controlled path to inspect protected evidence when investigation requires it. Observability should reduce uncertainty without creating a parallel ungoverned data store.