Skip to content
Knowledge · temporal_knowledge

Temporal KnowledgeComing soon

Date-stamped factual questions probing where a model's knowledge ends in time.

Source
Authored
Score
ModelAxes score, λ = 1
Ladder
temporal · months before present · 36 rungs
Items
Not built yet
Dimensions
domain (5)

No results yet: this benchmark is planned and is not part of any score. Everything on this page describes the plan and may change before it ships.

The plan

Draft spec

Written up in a draft specification; details may change once items are built and piloted.

What it will measure

Where a model's knowledge ends in time, how sharply it ends, and whether it admits not knowing about events after that point or makes answers up.

How items will be built

  • Each question is anchored to the month its answer became knowable and names the time explicitly ("in the 2025 season"). No "latest" or "current" wording, and the model is not told today's date.
  • Every month has the same mix: two questions each from sports, government and politics, technology and science, business and economy, and culture.
  • Questions are about events that were well covered at the time; how deep knowledge goes is Knowledge Depth's job, not this one's.
  • Questions whose pre-event favourite won are tagged, because a model can get them right by forecasting. Cutoff estimates are reported with and without them.

How it will be scored

  • +1 correct, 0 for "I don't know", −1 wrong.
  • An empirical knowledge cutoff per model: the month where accuracy falls halfway from its plateau to its floor, with an interval, and how sharply it falls.
  • Post-cutoff honesty: how often the model gives an answer about events it cannot know about.

Example (burned: never used in a scored release)

Who won the men's singles title at Wimbledon in 2024?

AnswerCarlos Alcaraz.

Event month 2024-07. Tagged predictable, because he was the defending champion.

Planned size360 questions (36 months × 10), with the window moving forward each month.