Skip to content
Reasoning · induction

InductionComing soon

Inducing a transformation rule from few examples on generated grid tasks (text-rendered).

Source
Generated
Score
ModelAxes score, λ = 1
Ladder
structural · rule depth · 1, 2, 3, 4, 5, 6
Items
Not built yet
Dimensions
—

No results yet: this benchmark is planned and is not part of any score. Everything on this page describes the plan and may change before it ships.

The plan

Draft spec

Written up in a draft specification; details may change once items are built and piloted.

What it will measure

Inferring a transformation rule from a few examples and applying it to a new input, in the style of ARC, with grids written out as text.

How items will be built

  • Grids of 3 to 15 cells a side with values 0–9. Rules chain primitives: rotate, reflect, recolor, extract objects, fill, tile, gravity, and conditionals.
  • Three demonstrations and a test input per item. A bounded search over every rule up to the item's depth checks that only one answer is consistent; demonstrations are added, or the item is redrawn, until it is.
  • Rules that leave an input unchanged, or equal a shorter rule, are rejected.

How it will be scored

  • Exact grid match graded by code; +1 / 0 / −1. Correct dimensions and cell accuracy are secondary.
  • x50 over rule depth, and accuracy by kind of primitive (which kinds of rule does the model miss?).

Example (burned: never used in a scored release)

One demonstration pair (an item gives three or more, then a test input): Input: 1 0 0 / 1 1 0 / 0 0 0 Output: 0 0 2 / 0 2 2 / 0 0 0

AnswerThe rule, depth 2: reflect left to right, then recolor 1 to 2.

Planned size1,200 items (6 depths × 200).