Reasoning · induction
InductionComing soon
Inducing a transformation rule from few examples on generated grid tasks (text-rendered).
- Source
- Generated
- Score
- ModelAxes score, λ = 1
- Ladder
- structural · rule depth · 1, 2, 3, 4, 5, 6
- Items
- Not built yet
- Dimensions
- —
No results yet: this benchmark is planned and is not part of any score. Everything on this page describes the plan and may change before it ships.
The plan
Draft specWritten up in a draft specification; details may change once items are built and piloted.
What it will measure
Inferring a transformation rule from a few examples and applying it to a new input, in the style of ARC, with grids written out as text.
How items will be built
- Grids of 3 to 15 cells a side with values 0–9. Rules chain primitives: rotate, reflect, recolor, extract objects, fill, tile, gravity, and conditionals.
- Three demonstrations and a test input per item. A bounded search over every rule up to the item's depth checks that only one answer is consistent; demonstrations are added, or the item is redrawn, until it is.
- Rules that leave an input unchanged, or equal a shorter rule, are rejected.
How it will be scored
- Exact grid match graded by code; +1 / 0 / −1. Correct dimensions and cell accuracy are secondary.
- x50 over rule depth, and accuracy by kind of primitive (which kinds of rule does the model miss?).
Example (burned: never used in a scored release)
One demonstration pair (an item gives three or more, then a test input): Input: 1 0 0 / 1 1 0 / 0 0 0 Output: 0 0 2 / 0 2 2 / 0 0 0
AnswerThe rule, depth 2: reflect left to right, then recolor 1 to 2.
Planned size1,200 items (6 depths × 200).