How ModelAxes measures
A short version of how the numbers on this site are made. Every score, interval and cost comes from the evaluation warehouse; this site only displays them.
Detailed methodology coming soon. The full protocol (item specifications, grading rules, judge prompts, and the statistics behind every interval) is being prepared for publication.
A frozen, held-out release
Results come from release 0.1.0-preliminary, frozen on Sep 27, 2026. Freezing fixed every question before any model on this site was run. The release holds 4,205 items. The questions and answers are held out and are never published, so models cannot be trained on them.
A public commitment records a SHA-256 hash of the full item set, so anyone can later check that the questions were not changed after results came in. The commitment is published here, along with a sample of 200 items drawn from separate development seeds, so you can see what the questions look like.
sha256:593554f9234ffae5e713dbbc4906ba6025d6ce10d6d28e28ae6591ea75fc67ca
What is measured
This release has 2 live axes, with one benchmark each:
- Knowledge Depth
- Closed-book questions about real people, places, species, medicines, software packages, papers and map locations: "tell me about X" profiles checked claim by claim against source facts, and short answers (a birth year, a population, an arXiv identifier, a developer, a paper's author or venue). Fame tiers run from 1 (well documented) to 6 (almost undocumented), measured per domain. Made-up entities are asked the same questions to measure fabrication; they are reported beside the score, never in it.
- Ballpark
- Twelve streams of calculation without tools (products, powers, long expressions, division, roots and logarithms, chained quotients, exponentials, real powers, factorials, binomial coefficients, trigonometry and generalized harmonic sums), growing in size. The reply is a number or a range: an exact answer scores highest, a correct range scores by its precision, and a miss costs more the further off it is.
Each axis score is the average of its benchmarks; the aggregate is the unweighted average of the live axes. 7 more axes and 20 more benchmarks are planned, and each planned benchmark's page describes its design. None of them count toward any score yet.
Reasoning on and off
Each model is run under up to two conditions:
- Reasoning on
- The model may reason before it answers, at medium reasoning effort.
- Reasoning off
- Reasoning is disabled; the model answers directly under a per-question token cap.
Some models cannot turn reasoning off (GPT-6.1 Sol, for example), so they have results only with reasoning on and show no reasoning lift. Reasoning lift compares the two conditions on the same questions.
Scoring: wrong answers cost points
Every answer earns +1 when correct, 0 for “I don't know”, and −1 when wrong or unreadable. A model that guesses when it doesn't know loses points, and a model that says so loses nothing. Scores run from −100 to +100, so a negative score means the model gave more wrong answers than right ones, not that it did worse than random.
Numeric answers (a calculation, a population, a year, a coordinate) can be exact, approximate, or a range. A correct answer earns credit for its precision: an exact value earns the full point, and a tight range earns more than a loose one. A wrong answer costs more the further off it is: at least half a point, and a full point once it is off by an order of magnitude. Years are scored by how many years off, and populations, which are best estimates, by how close they come.
Within a benchmark, each question type (Knowledge) or calculation stream (Math) counts equally, so a type with more questions does not count for more.
How answers are graded
Answers are graded by code wherever possible, and by a judge model only where needed:
- Math
- The whole reply must be the number or range. Code reads it and scores it; no judge is involved.
- Identifiers
- An arXiv ID is read by code and compared exactly.
- Short answers
- A name, a company, a venue or a code is compared with the reference answer (and its accepted aliases) by Drex 1.1, pinned to that version.
- Numbers in prose
- For Knowledge questions with a numeric answer, GPT-6 Luna reads out the number or range the model committed to, without seeing the correct answer; code does the scoring.
- Profiles
- “Tell me about X” answers are checked claim by claim against a fact sheet built from the sources, by GPT-6 Luna.
GPT-6 Luna is also one of the models being evaluated, so where it judges, it judges its own family. A judge from outside the panel is planned before the release loses its “preliminary” label.
Fabrication controls. Knowledge also asks the same questions about entities that do not exist. The right response is to say it doesn't know. How often a model makes up an answer instead is reported beside the Knowledge score and never counted in it.
Intervals and difficulty ladders
Every score carries a 95% interval from a cluster bootstrap (2,000 resamples). Questions about the same entity are resampled together, so the interval reflects how many independent things were tested, not just how many questions. Two models whose intervals overlap may not really differ.
Each benchmark climbs a difficulty ladder: fame tiers from well documented to almost undocumented for Knowledge, and growing problem sizes for Math. A curve per model shows accuracy on each rung, and x50 marks the rung where the model's accuracy crosses 50%, from a logistic fit.
Cost and speed
Cost per task is what the tokens the model actually used would cost at the provider's standard list price, from the token counts the API reports for each call (input, cached input and output, reasoning included). Discounts such as flex or batch pricing are not applied, so models are compared at the price most people pay.
Latency and output speed are shown only for runs served at a standard tier. The runs in this release used OpenAI's flex tier, which is slower by design, so no speed is reported for them.
What “preliminary” means
- The judges have not yet passed the agreement check against human labels (κ ≥ 0.85) that a final release requires.
- 2 models have been measured so far. More can be added on the same frozen items.
- A few Knowledge question types are short of their target size on some tiers. Their per-tier results carry wide intervals.
- The knowledge facts come from public databases, listed with their licenses on the sources page.