Skip to content
Research · search_lift

Search LiftComing soon · v0.2

How much web search adds to closed-book knowledge, and whether a model searches when it should.

Source
Authored
Score
ModelAxes score, λ = 1
Ladder
inherited · fame tier · 2, 3, 4, 5, 6
Items
Not built yet
Dimensions
—

No results yet: this benchmark is planned and is not part of any score. Everything on this page describes the plan and may change before it ships.

The plan

Draft spec

Written up in a draft specification; details may change once items are built and piloted.

What it will measure

How much web search recovers knowledge a model lacks, whether it changes how often the model makes things up, and whether the model searches when it needs to.

How items will be built

  • 200 Knowledge Depth questions drawn by a fixed rule and weighted toward obscure tiers: none from tier 1, 20 from tier 2 and 45 from each of tiers 3 to 6.
  • Each question is asked twice: closed-book, and with search and fetch tools (at most 6 calls).
  • One search backend for every model, so the comparison measures how the model uses search, not whose search engine is better.
  • The backend silently blocks ModelAxes pages and any page containing the question. The sources the facts come from are never blocked.

How it will be scored

  • Search lift: accuracy with search minus accuracy closed-book, paired per question and reported per tier.
  • The change in hallucination rate, and the share of questions answered correctly closed-book but wrongly with search.
  • Search appropriateness: does the model search on questions it would get wrong, and skip searching on ones it knows?
  • Tool calls and extra cost per question.

Planned size200 questions, each asked closed-book and with search.