Context Engineering · Principal
The memory expired yesterday. Why does the agent still act on it today?
The question
Interview question
A support agent stored a temporary preference: “For this incident, contact the customer on the emergency number.” It had a 24-hour expiry. The raw memory row is gone, yet today's prompt summary says “customer prefers the emergency number.” The agent calls it again. Was the TTL implementation broken?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
It may have worked exactly as written for the raw row. The failure is that a derived summary copied the fact and lost its scope. A memory system can have raw events, extracted facts, vector chunks and periodic summaries. Expiring one tier does not automatically invalidate every descendant. Research on deployment-time agent memory measures residue after raw-only deletion and discusses derived copies. That study does not prescribe this hypothetical TTL policy, but it demonstrates why deletion fidelity needs to be checked across memory tiers.
I would trace the fact's origin, subject, creation time, expiry, source revision and every derived artifact that included it. Did the summary carry the phrase without “for this incident”? Was it embedded under a new ID? Did a cached prompt include it even after the storage rows were updated? The agent should not be asked to guess the original expiry from compressed prose. A durable memory item needs provenance and validity metadata, and summary generation should preserve a link to the facts it depends on. At read time, materialized summaries should be invalidated, recomputed or filtered if a contributing fact is no longer valid.
There are different expiry semantics. A short-lived operational instruction should stop governing future actions at a precise boundary. A past event may remain in an audit record but not be promoted into current preference. Legal deletion may require removing or redacting derived text as well. Decide which contract applies before choosing a physical purge, tombstone, version gate or lazy rebuild. Test retrieval through every path, including exact lookup, vector search, cached context and a weekly summary, after the TTL fires. A deletion test that only queries the primary table misses the product failure.
What if recomputing every summary is expensive? Track dependencies and invalidate only affected artifacts, or keep source-linked fact IDs in a summary and verify them before use. Approximate cleanup may be acceptable for low-risk personalization, but not for an expired contact instruction. The user corrected the agent's memory. Why does the old fact return next week? concerns a corrected memory whose old fact returns. Here no correction was made. The original fact was intentionally temporary, and its derived copy escaped the time boundary.
Continue reading
Related questions
Read beyond the question
Explore more context engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →