Axis
Knowledge
Closed-book knowledge of real entities as they get more obscure, and what a model does when it doesn't know.
ScoreAccuracy95% interval
| Rank# | Model | Score | ±95% | $/task |
|---|---|---|---|---|
| 1 | GPT-6.1 Sol | Score 27.6, interval ±4.3, accuracy 54.6% | 27.6 | |
| 2 | GPT-6 Luna | Score −2.7, interval [−7.0, 1.7], accuracy 38.0% | −2.7 |
knowledge_depth
Knowledge Depth
Closed-book questions about real people, places, species, medicines, software packages, papers and map locations: "tell me about X" profiles checked claim by claim against source facts, and short answers (a birth year, a population, an arXiv identifier, a developer, a paper's author or venue). Fame tiers run from 1 (well documented) to 6 (almost undocumented), measured per domain. Made-up entities are asked the same questions to measure fabrication; they are reported beside the score, never in it.
ScoreAccuracy95% interval
| Rank# | Model | Score | ±95% | $/task |
|---|---|---|---|---|
| 1 | GPT-6.1 Sol | Score 27.6, interval ±4.3, accuracy 54.6% | 27.6 | |
| 2 | GPT-6 Luna | Score −2.7, interval [−7.0, 1.7], accuracy 38.0% | −2.7 |
Coming soontemporal_knowledge· coming soon