ParametricComing soon
Whether performance on a problem structure survives parameter changes across seeds.
- Source
- Generated
- Score
- ModelAxes score, λ = 1
- Ladder
- structural · template level · 1, 2, 3, 4, 5
- Items
- Not built yet
- Dimensions
- area (8)
No results yet: this benchmark is planned and is not part of any score. Everything on this page describes the plan and may change before it ships.
The plan
Draft specWritten up in a draft specification; details may change once items are built and piloted.
What it will measure
Whether a model has learned a problem's structure or only familiar instances of it. Each problem template is drawn many times with new numbers. Results that swing across draws, or drop when a famous instance is changed, point to pattern-matching rather than solving.
How items will be built
- 40 templates across 8 areas: rates and mixtures, number theory, combinatorics, probability, algebra, geometry, sequences, and graphs.
- Each template has 5 levels that set its parameter range, an exact solver (checked by brute force where feasible) and at least 3 phrasings (plain, word problem, formal).
- Templates with a famous textbook instance include it, for a memorization comparison.
- Every item is well-posed, and none mentions that a problem might be unanswerable.
How it will be scored
- Exact integer or fraction; +1 / 0 / −1. Claiming that a well-posed problem cannot be solved counts as wrong.
- x50 over levels per template.
- Instance variance: how often the result flips across the 10 draws of one template and level.
- Memorization gap: accuracy on the famous instance minus accuracy on changed ones.
Example (burned: never used in a scored release)
How many integers from 1 to 1,000 inclusive have digit sum equal to 10?
Template digit_sum_count at level 2 (N = 1,000, S from 3 to 26). Solved by digit dynamic programming and checked by brute force.
Planned size2,000 items (40 templates × 5 levels × 10 draws).