PlausibilityComing soon
Human-trivial physical and social judgments that models miss, adversarially authored.
- Source
- Authored
- Score
- ModelAxes score, λ = 1
- Ladder
- rubric · trap strength · 1, 2, 3
- Items
- Not built yet
- Dimensions
- category (6)
No results yet: this benchmark is planned and is not part of any score. Everything on this page describes the plan and may change before it ships.
The plan
Draft specWritten up in a draft specification; details may change once items are built and piloted.
What it will measure
Everyday physical and social judgments that are trivial for people but that models still miss, especially when the question is loaded with tempting but irrelevant detail.
How items will be built
- Written by people and kept only when reference models fail them. 30% are kept without that filter, so the filter's bias can be measured.
- Six equal categories: physical causes and materials, containment and object permanence, bodily and time limits, social conventions, quantity and scale, and irrelevant-detail traps.
- Trap strength 1–3: no distracting detail, one tempting detail, or several details that invite a wrong line of reasoning.
- Each item has a rewritten twin with different people, objects and numbers but the same critical fact.
How it will be scored
- Short open answers, graded by a judge against the critical fact, accepted phrasings and the known tempting wrong answers; +1 / 0 / −1.
- The judge must agree with human labels (κ ≥ 0.85 on at least 200 responses) before results are published.
- Which trap caught the model, the gap between original and twin, and whether reasoning helps or hurts.
Example (burned: never used in a scored release)
Lucia writes her grocery list in pencil on a sheet of paper, then leaves it on the dashboard of her car on a sunny afternoon with the windows closed. She comes back four hours later. What state is the list most likely in?
Tempting wrong answers: the pencil faded, the paper burned, the graphite melted. This shows the format, not the difficulty.
Planned size300 items, each with a twin (600 questions per model and condition).