Build a controlled set of renewal cases with current facts, superseded terms, conflicting notes, restricted data, corrected entities, and missing records. Ask the memory service for task-specific context and verify precision, source coverage, access enforcement, freshness, and explicit uncertainty. Then run the downstream agent and measure whether the context improves the decision packet without increasing unsupported assertions. Retrieval quality and agent quality should be evaluated separately so one does not hide the other.
Test writes and lifecycle behavior as carefully as reads. Confirm that rejected recommendations do not become approved strategy, corrections appear in later retrieval, revoked access removes protected context, and deletion requests reach derived indexes where applicable. Interrupt a workflow and resume it after source updates to see whether material conditions are revalidated. Measure storage and retrieval cost, latency, human correction, and the percentage of context items actually used. More retrieved text is not evidence of better memory.
Review failed retrieval as a product signal. Repeated misses may indicate weak indexing, poor entity resolution, an unavailable source, or an unrealistic workflow expectation. Repeated irrelevant results may indicate overly broad access or an imprecise context contract. Assign each class to an owner and a correction path. Do not tune prompts to conceal a data or authority problem that belongs in the source or memory service.