The memory store returned what it knows now, while the replay asked what the agent could know on Monday. On Tuesday the customer told the assistant “send updates by text, not email.” A Monday incident replay that injects this preference into context can make the earlier agent look more helpful or more policy-compliant than it really was. That is future information leakage. A timestamp on the memory row is useful only if the retrieval query enforces the replay's cutoff.

I would record at least when a fact was learned and when it was intended to apply. Those can differ. A user might say on Tuesday, “I've preferred text messages since January.” The fact has a claimed valid time in January but did not become known to this agent until Tuesday. Historical replay of Monday should not see it if the question is what the agent actually knew. A current answer about the customer's past preference may need the valid-time view. Temporal database work distinguishes transaction time from valid time for exactly this reason. The original bitemporal semantics paper provides that underlying model. We still have to choose which view the product requires.

Store memory revisions with provenance, observed-at time, valid interval when appropriate, and revocation or correction events. Build replay context from an as-of snapshot of information available to that run, including retrieval index and tool results if the evaluation claims exact replay. A plain created_at <= Monday filter is not enough if rows were overwritten in place or a summary created Wednesday compresses earlier and later facts together. Corrections need an explicit rule: evaluating what the agent knew then and evaluating what it should say now are separate tests.

This is not only an offline metric issue. A long-running agent can resume from a checkpoint while a new memory fact appears. It should decide whether to incorporate it at a defined step, not quietly let a mutable summary rewrite the context of an action already approved. Ask which facts were available to this agent decision at this point in its run.