Research · search_lift
Search LiftComing soon · v0.2
How much web search adds to closed-book knowledge, and whether a model searches when it should.
- Source
- Authored
- Score
- ModelAxes score, λ = 1
- Ladder
- inherited · fame tier · 2, 3, 4, 5, 6
- Items
- Not built yet
- Dimensions
- —
No results yet: this benchmark is planned and is not part of any score. Everything on this page describes the plan and may change before it ships.
The plan
Draft specWritten up in a draft specification; details may change once items are built and piloted.
What it will measure
How much web search recovers knowledge a model lacks, whether it changes how often the model makes things up, and whether the model searches when it needs to.
How items will be built
- 200 Knowledge Depth questions drawn by a fixed rule and weighted toward obscure tiers: none from tier 1, 20 from tier 2 and 45 from each of tiers 3 to 6.
- Each question is asked twice: closed-book, and with search and fetch tools (at most 6 calls).
- One search backend for every model, so the comparison measures how the model uses search, not whose search engine is better.
- The backend silently blocks ModelAxes pages and any page containing the question. The sources the facts come from are never blocked.
How it will be scored
- Search lift: accuracy with search minus accuracy closed-book, paired per question and reported per tier.
- The change in hallucination rate, and the share of questions answered correctly closed-book but wrongly with search.
- Search appropriateness: does the model search on questions it would get wrong, and skip searching on ones it knows?
- Tool calls and extra cost per question.
Planned size200 questions, each asked closed-book and with search.