Data and Knowledge Systems · Principal
The document sync cursor expired. Can we just request a fresh one and continue?
The question
Interview question
No, not if "fresh" means start tracking changes from now and forget the gap. A delta cursor says where the source can resume telling us what changed. If that cursor is no longer valid, we cannot infer the missing creates, updates, moves or deletes from the fact that the connector is healthy today. For a Microsoft Graph drive, the [driveItem delta API](https://learn.microsoft.com/en-us/graph/api/driveitem-delta?view=graph-rest-1.0) documents a 410 response and a fresh enumeration link when the service cannot use the old state. The source's exact resync instructions matter.
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
I would move the connector into a resync state and keep the last known-good index generation separate while enumerating the source again. Persist the provider's nextLink after each successfully applied page. A nextLink means there are more pages in the current round. Only a completed round yields a deltaLink to persist for future changes. Do not promote an intermediate page cursor as a completed sync watermark. Graph's delta overview makes that distinction.
Upserts should use stable source IDs and observed versions where available. Deletions are harder. Merely replaying current files does not remove documents that disappeared while we were offline. Build a complete inventory for the resync generation, reconcile old IDs that are absent, and apply the provider's deletion semantics. For a large corpus, do this without exposing an index where half the old files have been removed but the new inventory is only half loaded. Depending on product requirements, queries can use the prior generation with a visible freshness warning, or the affected source can be withheld. Permission checks must still happen at query time so stale index metadata cannot grant access.
The pushback is that a full enumeration might take two days while the source keeps changing. Then I need a source-supported snapshot or delta handoff, or repeated reconciliation until a valid completed delta round gives a new checkpoint. I cannot invent a gap-free snapshot guarantee from ordinary pagination. A bulk snapshot races the change stream covers a planned snapshot/change-stream race. Here the checkpoint was lost, so recovery must prove what happened during the blind interval before declaring freshness restored.
Continue reading
Related questions
Read beyond the question
Explore more data and knowledge systems
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →