DeductionComing soon
Long inference chains from premise sets; a third of targets are correctly undetermined.
- Source
- Generated
- Score
- ModelAxes score, λ = 1
- Ladder
- structural · proof depth · 1, 2, 4, 6, 8, 12, 16, 24
- Items
- Not built yet
- Dimensions
- label (3), vocabulary (3)
No results yet: this benchmark is planned and is not part of any score. Everything on this page describes the plan and may change before it ships.
The plan
Draft specWritten up in a draft specification; details may change once items are built and piloted.
What it will measure
Long chains of valid inference from stated premises, including recognizing when a conclusion does not follow.
How items will be built
- Premises about individuals and categories ("Every florp is a zindle. No mave is a quib."), in invented words by default.
- A proof of the rung's depth is built and distractor rules sharing its words are added. A solver derives the label (it follows, its negation follows, or neither), and items with a shorter shortcut are rejected.
- A third of targets are undetermined: the statements don't settle them.
- 20% of structures also come in real categories, once agreeing with world knowledge and once contradicting it, to measure belief bias.
How it will be scored
- +1 correct, 0 for "I don't know", −1 wrong. Concluding either way on an undetermined item is wrong.
- x50 over proof depth, accuracy per label, belief bias and reasoning lift.
Example (burned: never used in a scored release)
Every florp is a zindle. Every zindle is a mave. No mave is a quib. Every quib is a norble. Tessa is a florp. Based only on these statements, what can be concluded about whether Tessa is a norble?
Tessa is not a quib, which says nothing about being a norble. Asked about quib instead, the answer is "Tessa is not a quib" (three steps).
Planned size1,200 items (8 rungs × 150).