Skip to content
Strategy · operations

OperationsComing soon

Resource allocation and scheduling graded by objective gap against a computed optimum.

Source
Generated
Score
Optimality gap (normalized)
Ladder
structural · instance size · 1, 2, 3, 4, 5, 6
Items
Not built yet
Dimensions
stream (4)

No results yet: this benchmark is planned and is not part of any score. Everything on this page describes the plan and may change before it ships.

The plan

Draft spec

Written up in a draft specification; details may change once items are built and piloted.

What it will measure

Finding good plans under constraints (allocation, scheduling, routing), graded by how close each plan comes to the best possible one, and whether the model knows what its own plan achieves.

How items will be built

  • Four streams: 0/1 knapsack (4 to 24 items), single-machine scheduling (3 to 10 jobs), shortest tour (4 to 10 cities) and assigning workers to tasks (3 to 10 workers).
  • Every instance has a unique optimum, computed by an exact solver.
  • On at least 70% of each rung the obvious greedy rule is not optimal, so the benchmark doesn't just test whether the model knows the greedy heuristic.
  • The model gives its plan and states the plan's total.

How it will be scored

  • Optimality gap: how far the plan's true total is from the optimum. An infeasible plan counts as a gap of 1.
  • Optimal rate, feasible rate and mean gap, with x50 on the optimal rate per stream.
  • Claim mismatch: how often the total the model states differs from what its own plan actually achieves.

Example (burned: never used in a scored release)

A knapsack holds weight 10. Items (weight, value): A (5, 10), B (4, 40), C (6, 30), D (3, 50). Which items give the most value, and what is the total?

AnswerB and D, weight 7, value 90.

Format only: taking items greedily by value per weight also finds this, so the greedy rule above would reject this instance.

Planned size960 items (4 streams × 6 rungs × 40).