Discernment · hidden_assumptions
Hidden AssumptionsComing soon
Everyday questions that can't be answered as asked, each paired with an answerable twin, so reflexive "can't tell" is caught too.
- Source
- Authored
- Score
- ModelAxes score, λ = 1
- Ladder
- Not decided yet
- Items
- Not built yet
- Dimensions
- —
No results yet: this benchmark is planned and is not part of any score. Everything on this page describes the plan and may change before it ships.
The plan
Early planOnly the idea is settled. How items are built and scored is still open.
What it will measure
Whether a model notices what an everyday question doesn't say, instead of quietly filling the gap with an assumption.
How items will be built
- Items rewritten from a seed sheet of everyday questions that can't be answered as asked.
- Each gets an answerable twin with the same surface, so reflexively saying "can't be determined" is caught.
- Items whose right answer is a matter of opinion, or whose premise is a joke rather than a missing fact, are dropped.
How it will be scored
- +1 / 0 / −1, graded by a judge against a labeled truth kind: not enough information, no such thing, a contradiction, or a value.
Planned size40–60 items, each with its twin.