Skip to content
Discernment · hidden_assumptions

Hidden AssumptionsComing soon

Everyday questions that can't be answered as asked, each paired with an answerable twin, so reflexive "can't tell" is caught too.

Source
Authored
Score
ModelAxes score, λ = 1
Ladder
Not decided yet
Items
Not built yet
Dimensions
—

No results yet: this benchmark is planned and is not part of any score. Everything on this page describes the plan and may change before it ships.

The plan

Early plan

Only the idea is settled. How items are built and scored is still open.

What it will measure

Whether a model notices what an everyday question doesn't say, instead of quietly filling the gap with an assumption.

How items will be built

  • Items rewritten from a seed sheet of everyday questions that can't be answered as asked.
  • Each gets an answerable twin with the same surface, so reflexively saying "can't be determined" is caught.
  • Items whose right answer is a matter of opinion, or whose premise is a joke rather than a missing fact, are dropped.

How it will be scored

  • +1 / 0 / −1, graded by a judge against a labeled truth kind: not enough information, no such thing, a contradiction, or a value.

Planned size40–60 items, each with its twin.