Skip to content
Axis

Knowledge

Closed-book knowledge of real entities as they get more obscure, and what a model does when it doesn't know.

ScoreAccuracy95% interval
Rank#ModelScore
1GPT-6.1 SolScore 27.6, interval ±4.3, accuracy 54.6%27.6
2GPT-6 LunaScore −2.7, interval [−7.0, 1.7], accuracy 38.0%−2.7
Score ↑$/task · log ←↖ Better: higher score, lower cost$0.0010$0.0030−20−10010203040GPT-6 LunaGPT-6.1 Sol
knowledge_depth

Knowledge Depth

Closed-book questions about real people, places, species, medicines, software packages, papers and map locations: "tell me about X" profiles checked claim by claim against source facts, and short answers (a birth year, a population, an arXiv identifier, a developer, a paper's author or venue). Fame tiers run from 1 (well documented) to 6 (almost undocumented), measured per domain. Made-up entities are asked the same questions to measure fabrication; they are reported beside the score, never in it.

ModelAxes score, λ = 12,005 itemsDetails →
ScoreAccuracy95% interval
Rank#ModelScore
1GPT-6.1 SolScore 27.6, interval ±4.3, accuracy 54.6%27.6
2GPT-6 LunaScore −2.7, interval [−7.0, 1.7], accuracy 38.0%−2.7
Score ↑$/task · log ←↖ Better: higher score, lower cost$0.0010$0.0030−20−10010203040GPT-6 LunaGPT-6.1 Sol