YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Benchmark v1.1 for decision models
Version 1.1 (8 October 2026) is version 1.0 with the items of 10 of the 13 core families replaced by those of the corrected generators (../gen_v4, after the audit of 8 October 2026). The reference track and the three other core families have the same items as 1.0. CHANGES.md lists what changed, RUNBOOK.md the commands that bring in the runs and rebuild every table.
A decision model reads a state (a text) and one typed question (yes/no, one of N options, or a level) and returns a probability for every option in one forward pass, without writing text. Such a model is used to decide automatically when it is sure and to hand the case on (to a second look, to more context, to a person) when it is not. Version 1 measures that use. Evaluation only: the items must never enter a training set or a prompt for a model that is being trained.
One item format (SCHEMA.md), one scorer (score.py), one set of result tables (tables/, report.md). Everything runs with Python 3.9 and the
standard library.
The two tracks
- Core β generated families. Every item is written by a script from a seed, the answer is computed by a program, all entities are invented, so no system can have seen an item. A family has three levels of difficulty and, in most cases, contrast groups: the same case twice or three times with ONE planted difference that turns the answer. Level 3 needs at least four computation steps or three facts from different places of the state. The headline of v1.1 is the core track, because on it the strongest systems are 10 to 40 points apart where they are tied on public test sets.
- Reference β 35 public test sets converted to the item format, restricted to items that were not found in the authors' training data. It is a check
against outside data, with two cautions: the top systems are close together on it, and a system that trained on a related data set has an advantage
there (
EXPOSURE.mdsays who).
Contents: tasks/ (the items, one folder per task), TASKS.md (the registry, written by make_tasks_table.py), GAPS.md (what each core family lacks:
levels, contrast groups), EXPOSURE.md / exposure.json, scores.json and report.md (all scores of all systems), tables/ (T1 to T4 as CSV and LaTeX),
MISSING_RUNS.md (the runs still to be made), runs/ (the answer files), systems.json, runs_index.json, run/ (how to produce an answer file),
CHANGES.md, RUNBOOK.md, cache/ (copies of the cluster runs, written by collect_gen_v4.py), analysis/ and derived/ (failure analysis and the tables and figures of the paper), general_llm/ (the two general chat models of v1.0 asked again on the 10 changed families with the same code, prompt and parameters: raw replies, logs, run_log.txt; the method is described in ../v1/general_llm/README.md).
What the scores mean
All scores are computed from the system's own top probability and nothing is fitted to the benchmark. Percentages are shown as points (x 100) with one decimal. The top option of an item is the option with the largest probability (the first listed wins a tie); the top probability is its probability; an item is right when its top option is the gold option; an item without an answer is wrong with the uniform distribution.
- accuracy β the share of items that are right. Example: 7 of 10 items right: 70.0.
- pair accuracy β for tasks with contrast groups: the share of groups in which EVERY member is right. Example: two pairs, in the first both members are right, in the second one: accuracy 75.0, pair accuracy 50.0. A system that has learned a surface cue is right on one member of a pair and wrong on the other.
- pair direction β for yes/no pairs: the share of pairs in which the member whose truth is yes gets a higher P(yes) than the member whose truth is no (a tie does not count). It asks whether the probabilities rank the two cases correctly, whatever threshold is used. Example: P(yes) 0.8 against 0.3 counts, 0.4 against 0.6 does not: 50.0 for these two pairs.
- accuracy and pair accuracy per level β the same two numbers on the items of level 1, 2 and 3 (in v1.1 every core family has
meta.level, 200 items per level;GAPS.md). - auto_at_5, auto_at_1 β the largest share of the items that can be decided automatically while the error rate among the decided items stays at or below 5% (1%). Items are sorted by top probability, highest first; items with the same probability enter together; the answer is the longest prefix that meets the target. Example: ten items, the eight most confident hold no error and the ninth is wrong, so the prefix of nine has 11% errors: auto_at_5 is 80.0. This is the score that says how much work a system can take over at a stated error rate.
- acc_at_090, share_at_090 β the accuracy among the items whose top probability is at least 0.90, and the share of all items that reach it. Example: 5 of 10 items reach 0.90 and 4 of those are right: acc_at_090 = 80.0 with share 50.0. A system that says "sure" should be right about nine times in ten.
- aurc β area under the risk-coverage curve: the mean, over the items, of the error rate of the prefix (sorted as above) in which the item enters. 0 is perfect and lower is better. Example: right, right, wrong, right in the order of confidence: error rates 0, 0, 1/3, 1/4, aurc = 14.6.
- trigger_recall_20 β the share of the system's errors that lie among the 20% of items with the lowest top probability: what the rule "think again, fetch more context or hand over to a person when the confidence is low" would catch if it looked at one item in five. Example: 10 items, 2 errors, one of them among the two least confident items: 50.0.
- calibration error β 15 equal-width bins of the top probability; the sum over the bins of |sum of top probabilities - number right|, divided by the number of items. Example: every item at 0.90 and 70% right: 20.0.
- Brier score, log score, AUROC, ndcg@10, skill β as in version 0 (
SCHEMA.md); the reference tasks keep their own metric. - tokens β mean length of the thought of a system that thinks first (macro mean over families).
- Intervals β 95% percentile bootstrap, seed 20261007, 1,000 resamples (
--resamples); where a task has groups the groups are resampled, not the items.
The tables
- T1 core headline β one row per system: accuracy, pair accuracy, level-3 pair accuracy, auto_at_5, acc_at_090 (with its share), calibration error. Macro means over
the headline families: the core families with a valid run for at least half of the systems in the tables (the list is printed under the table; a new family
enters when enough systems have run on it). A system is in the table only if it has a valid run on every headline family. Pair accuracy is the mean over the
headline families that have contrast groups, level-3 pair accuracy over those of them that also have a level ladder (the note under the table lists them).
tables/T1_core_headline_intervals.csvhas the intervals. - T2 core per family β pair accuracy where the family has contrast groups, accuracy where it has not.
- T3 reference per layer and total β the chance-corrected score (0 = chance, 100 = perfect) per layer and the mean of the layers, as in version 0.
- T4 one pass against with reasoning β the systems that have both kinds of run: one-pass read-out and the read-out after the model's own thought, same items.
report.md has every score of every system, per family and per task.
How to produce an answer file and score it
run/README.md: the answer-file format, the interface for typed-decision systems, the wording for a model without one, the case of a model that thinks first.
python3 score.py NAME=runs/NAME [NAME=... ] [--tasks a,b] [--out DIR] [--resamples 1000] [--subsample FILE] # a few runs
python3 score.py --all --jobs 8 # everything: scores.json, report.md, tables/, MISSING_RUNS.md
python3 test_score.py # unit tests of the scores, hand-worked cases
A run counts for a task only if at least 99% of the items have a line and, for a generated family, it was made on the items of this version. Otherwise the cell is
"not run on this version" and MISSING_RUNS.md lists it.
Rebuilding and extending
python3 build_v1_1.py assembles tasks/ from its sources and can be run again at any time: the 13 core families from gen_v4/tasks (each one only if it passes
validate.py), the 35 reference tasks from v0/tasks, restricted to the items of v0/overlap_ours_clean_ids.json (checked byte for byte against v1.0). It never deletes
anything. carry_v1_runs.py copies the answer files of v1.0 that stay valid (every reference task; the three core families with unchanged items when v1.0 recorded
their items_sha256), collect_gen_v4.py fetches the runs on the gen_v4 items from the cluster and takes each one that passes its checks (docstrings), and
general_llm/collect.py turns the replies of the chat models into answer files. Then python3 score.py --all (a run made on other items of a family does not
count and shows up in MISSING_RUNS.md). build_exposure.py writes the exposure files (copied from v1.0, valid because the reference items are the same).
validate.py checks task folders; audit.py is the shortcut audit of version 0 (rules that never read the content). bash refresh.sh runs build, carry-over, scoring, registry and tests in one go.
- Downloads last month
- 7