Knowledge · knowledge_depth
Knowledge Depth
Closed-book questions about real people, places, species, medicines, software packages, papers and map locations: "tell me about X" profiles checked claim by claim against source facts, and short answers (a birth year, a population, an arXiv identifier, a developer, a paper's author or venue). Fame tiers run from 1 (well documented) to 6 (almost undocumented), measured per domain. Made-up entities are asked the same questions to measure fabrication; they are reported beside the score, never in it.
- Source
- Generated
- Score
- ModelAxes score, λ = 1
- Ladder
- structural · fame tier · 1, 2, 3, 4, 5, 6
- Items
- 2,005 items · 90 public
- Dimensions
- domain (8), kind (15)
ScoreAccuracy95% interval
| Rank# | Model | Score | ±95% | $/task |
|---|---|---|---|---|
| 1 | GPT-6.1 Sol | Score 27.6, interval ±4.3, accuracy 54.6% | 27.6 | |
| 2 | GPT-6 Luna | Score −2.7, interval [−7.0, 1.7], accuracy 38.0% | −2.7 |
Ladder
Accuracy by fame tier rung for every model; the focused model shows its Wilson band and x50.
| Model | x50 | 95% int. |
|---|---|---|
| GPT-6.1 Sol | 3.8 tier | |
| GPT-6 Luna | 2.1 tier |
Outcomes
Correct 1,712 (45%)Incorrect 1,195 (32%)Abstained 883 (23%)Format failure 0 (0%)
Model
Outcomes
Correct
Slices · domain
Score per domain value for every model.
| GPT-6.1 Sol | 67.4 | 23.9 | 35.6 | 22.1 | 14.6 | 15.6 | 39.9 | 16.4 |
| GPT-6 Luna | 20.6 | −5.8 | −2.4 | −30.7 | −0.8 | −2.4 | 19.0 | −15.8 |
Slices · kind
Score per kind value for every model.
| GPT-6.1 Sol | 67.4 | 22.3 | 23.9 | 31.0 | 42.9 | 14.9 | 14.6 | 29.7 | 26.4 | 46.2 | 6.9 | 5.4 | 37.0 | 16.4 | 29.6 |
| GPT-6 Luna | 20.6 | −0.9 | −5.8 | −9.7 | 14.3 | 0.0 | −46.9 | −14.5 | 6.5 | 19.3 | −0.7 | −13.7 | 23.6 | −15.8 | −16.8 |