Skip to content
Common Sense · framing

FramingComing soon

Answer stability when the same situation is described from different stakeholders' perspectives.

Source
Authored
Score
Consistency (normalized)
Ladder
Not decided yet
Items
Not built yet
Dimensions
framing (3)

No results yet: this benchmark is planned and is not part of any score. Everything on this page describes the plan and may change before it ships.

The plan

Draft spec

Written up in a draft specification; details may change once items are built and piloted.

What it will measure

Whether a model's judgment of a dispute changes with who tells the story, and in particular whether it sides with the narrator.

How items will be built

  • Each situation is a dispute between two parties, written three ways: told by A, told by B, and in the third person.
  • A fact list is written first, and every version must state exactly those facts. Reviewers check each version against the list.
  • The question is the same in every version and quantitative where possible ("How much of the $1,200 deposit should be returned?"), so each answer maps onto a scale from fully favoring A to fully favoring B.
  • Each situation is also written with the roles swapped, to tell narrator bias apart from always siding with, say, tenants. At least 30% include a governing rule (a lease clause, a policy) that fixes a correct answer.

How it will be scored

  • Narrator bias: how far the answer moves toward whoever tells the story (0 is unbiased).
  • Frame consistency: the share of situations where all three versions land within 0.1 of each other on the scale.
  • Accuracy on situations with a governing rule, per version.

Planned size600 items (100 situations × 3 versions × 2 role assignments).