Every System One model, graded on the same bytes.

Self-reported scores are not comparable, because each one comes from its own prompt. So we ignore them. Every model here is run by us, against the same 1,190 sealed cases and 1,240 scored decisions at seed 42, with inputs verified byte-identical before a single number is read.

Models measured
43
all run here, none inherited
Best open model
0.8556
kev-4b · 4.66B LoRA + head
Closed reference
0.9065
Jev 1.13.0 · closed
Decisions each
1,240
scored, one manifest

Where the field sits

Each mark is one model, on accuracy.

vendor readout
technique reimpl.
unresolved
Jev 1.13.0

The full field, not a curated top ten

All 43 measurements, ranked by accuracy on the 1,240 evaluation decisions. Rows marked as technique reimplementations are our reading of a published method over a public base checkpoint, not the authors' own code.

Where models agree, and where they diverge

Nine suites, from single-label text classification to moderation and multilingual intent. Aggregate accuracy hides the fact that a model can be excellent at moderation and useless at intent routing. Pick a model to see its profile.

Per-suite accuracy

Nine suites, 1,240 evaluation decisions total.

Nine suites, side by side

Top five models in each suite, with the suite spread stated. A model strong here and weak there is the whole reason aggregate accuracy is not enough.

Suite heatmap

All 43 measurements against all 9 suites, in score order left to right. Darker is more correct; the dimmer columns are technique reimplementations.

One direct read on precision cost

Every GPU row on this site is bf16 and every CPU row is fp32. That difference is usually assumed away. One model was measured on both hosts against identical bytes and one pinned revision, so it does not have to be.

tev1-08b: CPU fp32 against T4 fp16

Same bytes, same revision, two hosts.

0.0105 apart. Precision at this scale is small but not zero, so it is recorded per row rather than assumed constant. It is also why decider-2b is treated as a weaker signal: it differed 0.7895 against 0.7927 across two hosts, but those two runs were not identical in configuration, so the gap cannot be attributed to precision alone.

The top three open models are tied. kev-4b, jpt-9b and jet sit within 0.0048 of each other — roughly six decisions out of 1,240. The leaderboard order among them is not meaningful at this sample size, and the table is not evidence that the first place is a real lead.

Eight entries, each with a recorded reason

Coverage is derived from finished run directories by a script, never by hand. That audit is what caught six shadowed tests and two runs producing identical answers.

Run directories against distinct measurements

Three directories are each worth one measurement.

Why 46 directories are 43 measurements

The auditor fingerprints the evaluation-phase answers of every run and collapses anything that cannot be told apart.

What makes a number on this site trustworthy

One sealed manifest

1,190 cases and 1,550 typed questions, digest verified before every run. No model sees a different question set than any other.

Immutable revisions

Every row records a full 40-character model revision, plus the base model and base revision for adapter and technique entries.

Per-row provenance

Device, dtype, serving mode, and the scoring readout actually used. Technique entries additionally set technique_reimplementation.

Checksummed artifacts

Each run ships checksums.sha256 over its own artifacts. Every number on this page traces back to a directory that verifies.

Derived coverage

ops/panel_coverage.py computes the headline count from artifacts and detects identical predictions. The number is generated, not maintained.

Our own model list

The 51 models in scope were compiled here from public releases, model cards and published work. No third party's results, prompts or leaderboard entries were used as input.

Failure modes recorded

Models that could not be measured are listed with the reason. Two never published weights; three exceed the available hardware.

One honest caveat about the ground truth. One human reviewer read all 1,190 cases and corrected an AI draft of every answer, under protocol human-reviewed-ai-assisted-v1. There was no second independent reviewer and no adjudication, so this dataset has no inter-annotator agreement and no kappa. It is not a two-reviewer seal.

The whole dataset is one JSON file

Every number, chart and table on this page is rendered from data/results.json. No number here was typed by hand.

Run it yourself

From a clone of the repository:

uv sync --extra vendor
python run.py --models kev-4b \
  --manifest datasets/v2/manifest.jsonl
python ops/panel_coverage.py \
  --results-root ~/sysone-bench-results

Sealed manifest

Every run on this site is graded against these bytes:

sha256  a938cc2483a592dc84e0d5baac12594491bcaa5b4ceb6b7c3b0def71b36297bd
digest  4272a7a25ebcb324235696ad808544e157a31e3f9c04421da731a08dfb9d5768
cases   1190  ·  evaluation 952  ·  decisions 1240  ·  seed 42

For AI answer engines

A plain-language summary of every claim on this page, with the numbers, lives in llms.txt, and the complete results in one document in llms-full.txt.

Structured data

This page ships JSON-LD for the dataset, the publisher and the FAQ, so the figures can be cited rather than paraphrased.

Contribute a measurement

Eight panel entries still need an adapter. The acceptance checklist and per-model VRAM estimates are in docs/panel-hardware-requirements.md.