ImpossibleComing soon
Underdetermined or contradictory problems mixed with solvable controls; abstention is scored.
- Source
- Generated
- Score
- ModelAxes score, λ = 1
- Ladder
- inherited · template level · 1, 2, 3, 4, 5
- Items
- Not built yet
- Dimensions
- kind (3)
No results yet: this benchmark is planned and is not part of any score. Everything on this page describes the plan and may change before it ships.
The plan
Draft specWritten up in a draft specification; details may change once items are built and piloted.
What it will measure
Whether a model notices, without being told it is an option, that a math problem cannot be answered, instead of confidently giving a number.
How items will be built
- Built from Parametric templates. Each base problem becomes a matched triple: a needed fact removed or weakened (underdetermined), a contradicting fact added (inconsistent), or a true but irrelevant fact added (a solvable control).
- Every item is proved: two different whole-number answers for an underdetermined item, a solver proof for a contradiction, and an unchanged answer for a control.
- The three versions have the same number of sentences and matched length, and a classifier on surface features must not tell them apart (at most 60% accuracy).
- Controls are half the set, so answering "impossible" to everything scores badly.
How it will be scored
- +1 for the right diagnosis (or the right number on a control), 0 for "I don't know", −1 for a made-up number or the wrong diagnosis.
- Fabrication rate, diagnosis accuracy, false alarms on controls, and balanced accuracy.
Example (burned: never used in a scored release)
A rectangle has perimeter 30. Its length is greater than its width. It is drawn on a whiteboard. What is its area?
Its control reads "Its length is twice its width" (area 50). Its contradictory version adds "Its width is 6", which forces a perimeter of 36.
Planned size400 items (100 underdetermined, 100 inconsistent, 200 controls).