Self-reported scores are not comparable, because each one comes from its own prompt. So we ignore them. Every model here is run by us, against the same 1,190 sealed cases and 1,240 scored decisions at seed 42, with inputs verified byte-identical before a single number is read.
Each mark is one model, on accuracy.
All 43 measurements, ranked by accuracy on the 1,240 evaluation decisions. Rows marked as technique reimplementations are our reading of a published method over a public base checkpoint, not the authors' own code.
Nine suites, from single-label text classification to moderation and multilingual intent. Aggregate accuracy hides the fact that a model can be excellent at moderation and useless at intent routing. Pick a model to see its profile.
Nine suites, 1,240 evaluation decisions total.
Top five models in each suite, with the suite spread stated. A model strong here and weak there is the whole reason aggregate accuracy is not enough.
All 43 measurements against all 9 suites, in score order left to right. Darker is more correct; the dimmer columns are technique reimplementations.
Every GPU row on this site is bf16 and every CPU row is fp32. That difference is usually assumed away. One model was measured on both hosts against identical bytes and one pinned revision, so it does not have to be.
Same bytes, same revision, two hosts.
0.0105 apart. Precision at this scale is small but not
zero, so it is recorded per row rather than assumed constant. It is also why
decider-2b is treated as a weaker signal: it differed 0.7895 against 0.7927
across two hosts, but those two runs were not identical in configuration, so the gap cannot
be attributed to precision alone.
The top three open models are tied. kev-4b, jpt-9b and jet sit within 0.0048 of each other — roughly six decisions out of 1,240. The leaderboard order among them is not meaningful at this sample size, and the table is not evidence that the first place is a real lead.
Coverage is derived from finished run directories by a script, never by hand. That audit is what caught six shadowed tests and two runs producing identical answers.
Three directories are each worth one measurement.
The auditor fingerprints the evaluation-phase answers of every run and collapses anything that cannot be told apart.
1,190 cases and 1,550 typed questions, digest verified before every run. No model sees a different question set than any other.
Every row records a full 40-character model revision, plus the base model and base revision for adapter and technique entries.
Device, dtype, serving mode, and the scoring readout actually used. Technique entries
additionally set technique_reimplementation.
Each run ships checksums.sha256 over its own artifacts. Every number on this
page traces back to a directory that verifies.
ops/panel_coverage.py computes the headline count from artifacts and detects
identical predictions. The number is generated, not maintained.
The 51 models in scope were compiled here from public releases, model cards and published work. No third party's results, prompts or leaderboard entries were used as input.
Models that could not be measured are listed with the reason. Two never published weights; three exceed the available hardware.
One honest caveat about the ground truth.
One human reviewer read all 1,190 cases and corrected an AI draft of every answer, under
protocol human-reviewed-ai-assisted-v1. There was no second independent reviewer
and no adjudication, so this dataset has no inter-annotator agreement and no kappa. It is not
a two-reviewer seal.
Every number, chart and table on this page is rendered from
data/results.json. No number here was typed by
hand.
From a clone of the repository:
uv sync --extra vendor
python run.py --models kev-4b \
--manifest datasets/v2/manifest.jsonl
python ops/panel_coverage.py \
--results-root ~/sysone-bench-results
Every run on this site is graded against these bytes:
sha256 a938cc2483a592dc84e0d5baac12594491bcaa5b4ceb6b7c3b0def71b36297bd
digest 4272a7a25ebcb324235696ad808544e157a31e3f9c04421da731a08dfb9d5768
cases 1190 · evaluation 952 · decisions 1240 · seed 42
A plain-language summary of every claim on
this page, with the numbers, lives in llms.txt, and the complete results in one document in llms-full.txt.
This page ships JSON-LD for the dataset, the publisher and the FAQ, so the figures can be cited rather than paraphrased.
Eight panel entries still need an
adapter. The acceptance checklist and per-model VRAM estimates are in
docs/panel-hardware-requirements.md.