Skip to content
Knowledge · knowledge_depth

Knowledge Depth

Closed-book questions about real people, places, species, medicines, software packages, papers and map locations: "tell me about X" profiles checked claim by claim against source facts, and short answers (a birth year, a population, an arXiv identifier, a developer, a paper's author or venue). Fame tiers run from 1 (well documented) to 6 (almost undocumented), measured per domain. Made-up entities are asked the same questions to measure fabrication; they are reported beside the score, never in it.

Source
Generated
Score
ModelAxes score, λ = 1
Ladder
structural · fame tier · 1, 2, 3, 4, 5, 6
Items
2,005 items · 90 public
Dimensions
domain (8), kind (15)
ScoreAccuracy95% interval
Rank#ModelScore
1GPT-6.1 SolScore 27.6, interval ±4.3, accuracy 54.6%27.6
2GPT-6 LunaScore −2.7, interval [−7.0, 1.7], accuracy 38.0%−2.7
Score ↑$/task · log ←↖ Better: higher score, lower cost$0.0010$0.0030−20−10010203040GPT-6 LunaGPT-6.1 Sol

Ladder

Accuracy by fame tier rung for every model; the focused model shows its Wilson band and x50.

0%25%50%75%100%123456fame tierGPT-6.1 SolGPT-6 Luna
Modelx50
GPT-6.1 Sol3.8 tier
GPT-6 Luna2.1 tier

Outcomes

Correct 1,712 (45%)Incorrect 1,195 (32%)Abstained 883 (23%)Format failure 0 (0%)
Model
Outcomes
Correct
1
54%
GPT-6.1 Sol: 1,018 correct, 439 incorrect, 436 abstained, 0 format failures
2
37%
GPT-6 Luna: 694 correct, 756 incorrect, 447 abstained, 0 format failures

Slices · domain

Score per domain value for every model.

GPT-6.1 Sol
67.4
23.9
35.6
22.1
14.6
15.6
39.9
16.4
GPT-6 Luna
20.6
−5.8
−2.4
−30.7
−0.8
−2.4
19.0
−15.8

Slices · kind

Score per kind value for every model.

GPT-6.1 Sol
67.4
22.3
23.9
31.0
42.9
14.9
14.6
29.7
26.4
46.2
6.9
5.4
37.0
16.4
29.6
GPT-6 Luna
20.6
−0.9
−5.8
−9.7
14.3
0.0
−46.9
−14.5
6.5
19.3
−0.7
−13.7
23.6
−15.8
−16.8