Data and Knowledge Systems · Principal
The incident happened last Sunday. Why did the AI report count it this week?
The question
Interview question
An incident was created Sunday at 23:50 in the customer's timezone. Its event arrived at the warehouse on Tuesday after an outage. The assistant answers “How many incidents occurred last week?” from a table partitioned by ingest date. It reports the event in this week. A beautifully worded answer can still have the wrong time semantics.
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
There are at least three clocks here. The event time says when the incident occurred. The processing or ingest time says when this pipeline saw it. A report's as-of time says which facts were known when we published the report. Filtering ingested_at for a question about occurrence substitutes one clock for another. Apache Beam's event-time model makes this distinction explicit and uses watermarks and triggers to decide when windowed results can be emitted and revised.
I would ask the product owner exactly what “last week” means: the customer's timezone, inclusive start and exclusive end, and whether a late arrival should restate an already published report. Store the source occurrence timestamp, original timezone if supplied, normalized instant, ingest timestamp, and source identifier. Partitioning for storage efficiency need not dictate the semantic filter. Compute occurrence-based counts on event time and choose an explicit policy for late data: provisional numbers, a bounded correction window, or a versioned restatement. A watermark is an estimate of completeness, not proof that no older event will arrive.
Now the interviewer asks what the assistant should do when the user requests “the report as it looked Monday morning.” Recomputing today's event-time count is also wrong. You need an as-of view or versioned materialization that preserves what had been observed by Monday, including the report version and the late event's later arrival. A current corrected view and a historical published view answer different questions. Carry the data cutoff and definition into the tool response so the assistant can say which one it used.
I would test a late Sunday event ingested Tuesday, a timezone boundary, a source timestamp correction, and duplicate delivery. Verify the SQL or aggregation against expected membership, then verify the natural-language answer names the cutoff when it matters. This is distinct from The training feature knew about a payment that had not arrived yet. How?, where a later document leaks into a historical answer. Here the source event is old, but its late arrival and the chosen clock alter which reporting window owns it.
Continue reading
Related questions
Read beyond the question
Explore more data and knowledge systems
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →