Skip to content
Math · parametric

ParametricComing soon

Whether performance on a problem structure survives parameter changes across seeds.

Source
Generated
Score
ModelAxes score, λ = 1
Ladder
structural · template level · 1, 2, 3, 4, 5
Items
Not built yet
Dimensions
area (8)

No results yet: this benchmark is planned and is not part of any score. Everything on this page describes the plan and may change before it ships.

The plan

Draft spec

Written up in a draft specification; details may change once items are built and piloted.

What it will measure

Whether a model has learned a problem's structure or only familiar instances of it. Each problem template is drawn many times with new numbers. Results that swing across draws, or drop when a famous instance is changed, point to pattern-matching rather than solving.

How items will be built

  • 40 templates across 8 areas: rates and mixtures, number theory, combinatorics, probability, algebra, geometry, sequences, and graphs.
  • Each template has 5 levels that set its parameter range, an exact solver (checked by brute force where feasible) and at least 3 phrasings (plain, word problem, formal).
  • Templates with a famous textbook instance include it, for a memorization comparison.
  • Every item is well-posed, and none mentions that a problem might be unanswerable.

How it will be scored

  • Exact integer or fraction; +1 / 0 / −1. Claiming that a well-posed problem cannot be solved counts as wrong.
  • x50 over levels per template.
  • Instance variance: how often the result flips across the 10 draws of one template and level.
  • Memorization gap: accuracy on the famous instance minus accuracy on changed ones.

Example (burned: never used in a scored release)

How many integers from 1 to 1,000 inclusive have digit sum equal to 10?

Answer63

Template digit_sum_count at level 2 (N = 1,000, S from 3 to 26). Solved by digit dynamic programming and checked by brute force.

Planned size2,000 items (40 templates × 5 levels × 10 draws).