Reasoning · classic_vs_novel
Classic vs NovelComing soon
Paired canonical and modified puzzles measuring memorized-solution intrusion.
- Source
- Authored
- Score
- ModelAxes score, λ = 1
- Ladder
- Not decided yet
- Items
- Not built yet
- Dimensions
- modification (4)
No results yet: this benchmark is planned and is not part of any score. Everything on this page describes the plan and may change before it ships.
The plan
Draft specWritten up in a draft specification; details may change once items are built and piloted.
What it will measure
Whether a model applies a memorized solution to a famous puzzle that has been changed so the memorized answer is wrong or wasteful.
How items will be built
- Each item is a pair: the classic puzzle, with its well-known answer cited, and a modified version with a different correct answer.
- Four kinds of change, balanced: a constraint removed so the puzzle becomes trivial, the goal inverted, the numbers changed, or what someone knows changed.
- Two reviewers solve each modified version blind. Variants that are already published (such as "Monty Fall") are reported separately.
How it will be scored
- +1 / 0 / −1 on each version; numbers graded by code, short text by a judge.
- Memorized-answer intrusion: how often the model gives the classic answer to the modified puzzle. This is the headline number.
- Overthinking cost: output tokens on a trivialized puzzle relative to its classic version.
Example (burned: never used in a scored release)
You have a 3-liter jug, a 5-liter jug and unlimited water. How do you measure exactly 3 liters?
AnswerFill the 3-liter jug. One step.
The classic puzzle asks for 4 liters, which takes the well-known multi-step procedure.
Planned size100 pairs.