Skip to content
Common Sense · plausibility

PlausibilityComing soon

Human-trivial physical and social judgments that models miss, adversarially authored.

Source
Authored
Score
ModelAxes score, λ = 1
Ladder
rubric · trap strength · 1, 2, 3
Items
Not built yet
Dimensions
category (6)

No results yet: this benchmark is planned and is not part of any score. Everything on this page describes the plan and may change before it ships.

The plan

Draft spec

Written up in a draft specification; details may change once items are built and piloted.

What it will measure

Everyday physical and social judgments that are trivial for people but that models still miss, especially when the question is loaded with tempting but irrelevant detail.

How items will be built

  • Written by people and kept only when reference models fail them. 30% are kept without that filter, so the filter's bias can be measured.
  • Six equal categories: physical causes and materials, containment and object permanence, bodily and time limits, social conventions, quantity and scale, and irrelevant-detail traps.
  • Trap strength 1–3: no distracting detail, one tempting detail, or several details that invite a wrong line of reasoning.
  • Each item has a rewritten twin with different people, objects and numbers but the same critical fact.

How it will be scored

  • Short open answers, graded by a judge against the critical fact, accepted phrasings and the known tempting wrong answers; +1 / 0 / −1.
  • The judge must agree with human labels (κ ≥ 0.85 on at least 200 responses) before results are published.
  • Which trap caught the model, the gap between original and twin, and whether reasoning helps or hurts.

Example (burned: never used in a scored release)

Lucia writes her grocery list in pencil on a sheet of paper, then leaves it on the dashboard of her car on a sunny afternoon with the windows closed. She comes back four hours later. What state is the list most likely in?

AnswerEssentially unchanged and still readable (perhaps warm or slightly curled).

Tempting wrong answers: the pencil faded, the paper burned, the graphite melted. This shows the format, not the difficulty.

Planned size300 items, each with a twin (600 questions per model and condition).