AnticipationComing soon
Chained consequence reasoning over generated multi-agent systems with stated policies.
- Source
- Generated
- Score
- ModelAxes score, λ = 1
- Ladder
- structural · agents x rounds · 4, 9, 15, 32, 60, 120
- Items
- Not built yet
- Dimensions
- —
No results yet: this benchmark is planned and is not part of any score. Everything on this page describes the plan and may change before it ships.
The plan
Draft specWritten up in a draft specification; details may change once items are built and piloted.
What it will measure
Following the knock-on effects in a system of agents that each follow a stated rule: when something changes, where does the system end up?
How items will be built
- 2–6 agents (shops, drivers, teams) each follow a policy stated in a sentence or two, drawn from about 30 templates such as undercut, match, threshold switch, tit-for-tat and cooldown.
- The update order is stated. Higher rungs add agents, rounds and one or two changes partway through ("from day 3, shop B never prices below 7").
- The ground truth is exact simulation. The answer must differ from the starting state, from the state after one round, and from what you would get if no one reacted.
- 30% of items come in pairs, with and without the change, to test whether the model tracks its effect.
How it will be scored
- Exact answers graded by code; +1 / 0 / −1.
- x50 over agents × rounds, accuracy on the pairs, and tagged error types (reacting to the current round instead of the previous one, ignoring a floor, applying a change on the wrong round).
Example (burned: never used in a scored release)
Three shops sell bread. Each morning, every shop sets its price to 1 less than the lowest price the other two shops charged the previous day, but never below 3. On day 1 the prices are A = 10, B = 8, C = 9. What are the prices on day 4?
Day 2: 7, 8, 7. Day 3: 6, 6, 6.
Planned size300 items (6 rungs × 50).