Common Sense · framing
FramingComing soon
Answer stability when the same situation is described from different stakeholders' perspectives.
- Source
- Authored
- Score
- Consistency (normalized)
- Ladder
- Not decided yet
- Items
- Not built yet
- Dimensions
- framing (3)
No results yet: this benchmark is planned and is not part of any score. Everything on this page describes the plan and may change before it ships.
The plan
Draft specWritten up in a draft specification; details may change once items are built and piloted.
What it will measure
Whether a model's judgment of a dispute changes with who tells the story, and in particular whether it sides with the narrator.
How items will be built
- Each situation is a dispute between two parties, written three ways: told by A, told by B, and in the third person.
- A fact list is written first, and every version must state exactly those facts. Reviewers check each version against the list.
- The question is the same in every version and quantitative where possible ("How much of the $1,200 deposit should be returned?"), so each answer maps onto a scale from fully favoring A to fully favoring B.
- Each situation is also written with the roles swapped, to tell narrator bias apart from always siding with, say, tenants. At least 30% include a governing rule (a lease clause, a policy) that fixes a correct answer.
How it will be scored
- Narrator bias: how far the answer moves toward whoever tells the story (0 is unbiased).
- Frame consistency: the share of situations where all three versions land within 0.1 of each other on the scale.
- Accuracy on situations with a governing rule, per version.
Planned size600 items (100 situations × 3 versions × 2 role assignments).