Skip to content
Reasoning · deduction

DeductionComing soon

Long inference chains from premise sets; a third of targets are correctly undetermined.

Source
Generated
Score
ModelAxes score, λ = 1
Ladder
structural · proof depth · 1, 2, 4, 6, 8, 12, 16, 24
Items
Not built yet
Dimensions
label (3), vocabulary (3)

No results yet: this benchmark is planned and is not part of any score. Everything on this page describes the plan and may change before it ships.

The plan

Draft spec

Written up in a draft specification; details may change once items are built and piloted.

What it will measure

Long chains of valid inference from stated premises, including recognizing when a conclusion does not follow.

How items will be built

  • Premises about individuals and categories ("Every florp is a zindle. No mave is a quib."), in invented words by default.
  • A proof of the rung's depth is built and distractor rules sharing its words are added. A solver derives the label (it follows, its negation follows, or neither), and items with a shorter shortcut are rejected.
  • A third of targets are undetermined: the statements don't settle them.
  • 20% of structures also come in real categories, once agreeing with world knowledge and once contradicting it, to measure belief bias.

How it will be scored

  • +1 correct, 0 for "I don't know", −1 wrong. Concluding either way on an undetermined item is wrong.
  • x50 over proof depth, accuracy per label, belief bias and reasoning lift.

Example (burned: never used in a scored release)

Every florp is a zindle. Every zindle is a mave. No mave is a quib. Every quib is a norble. Tessa is a florp. Based only on these statements, what can be concluded about whether Tessa is a norble?

AnswerNothing. The statements don't settle it.

Tessa is not a quib, which says nothing about being a norble. Asked about quib instead, the answer is "Tessa is not a quib" (three steps).

Planned size1,200 items (8 rungs × 150).