Skip to content
Reasoning · classic_vs_novel

Classic vs NovelComing soon

Paired canonical and modified puzzles measuring memorized-solution intrusion.

Source
Authored
Score
ModelAxes score, λ = 1
Ladder
Not decided yet
Items
Not built yet
Dimensions
modification (4)

No results yet: this benchmark is planned and is not part of any score. Everything on this page describes the plan and may change before it ships.

The plan

Draft spec

Written up in a draft specification; details may change once items are built and piloted.

What it will measure

Whether a model applies a memorized solution to a famous puzzle that has been changed so the memorized answer is wrong or wasteful.

How items will be built

  • Each item is a pair: the classic puzzle, with its well-known answer cited, and a modified version with a different correct answer.
  • Four kinds of change, balanced: a constraint removed so the puzzle becomes trivial, the goal inverted, the numbers changed, or what someone knows changed.
  • Two reviewers solve each modified version blind. Variants that are already published (such as "Monty Fall") are reported separately.

How it will be scored

  • +1 / 0 / −1 on each version; numbers graded by code, short text by a judge.
  • Memorized-answer intrusion: how often the model gives the classic answer to the modified puzzle. This is the headline number.
  • Overthinking cost: output tokens on a trivialized puzzle relative to its classic version.

Example (burned: never used in a scored release)

You have a 3-liter jug, a 5-liter jug and unlimited water. How do you measure exactly 3 liters?

AnswerFill the 3-liter jug. One step.

The classic puzzle asks for 4 liters, which takes the well-known multi-step procedure.

Planned size100 pairs.