Same model. Different harness. Different outcome.

The runtime around an agent (tools, memory, retries, failure handling) moves the score. Read the labeled preview table, then open the benchmark for a live same-task compare. Numbers below are labeled preview fixtures, not audited rankings (preview-seed).

Scored harnesses
2
+9 catalog-only · Preview
Tasks represented
13
Preview
Mean success
68.8%
Preview
Maximum mean duration
1.24s
Preview

Fixed model, fixed suite. The harness is the variable.

+38.2 pts success delta, preview

Structured ReAct

preview-model-snapshot

87.8%

Preview

6 seed runs in the process-local preview store

Baseline Loop

preview-model-snapshot

49.7%

Preview

6 seed runs in the process-local preview store

One table. Same work. Comparable harnesses.

Terminal Core seed fixtures on a fixed model snapshot. Preview only; mock-eval playground runs are not included in these ranks. Data class: preview-seed. Live harness comparison is on the benchmark. Catalog structure is on the field.

Preview harness leaderboard with success rate, latency, cost, and reliability. Sample and process-local data only. Catalog-only external harnesses without scores are omitted.
#HarnessSuccess rateMean latencyCost / taskReliability
01
Structured ReAct
0.9.4 · 6 preview runs
87.8%
1.24s$0.006100.0%
02
Baseline Loop
0.3.1 · 6 preview runs
49.7%
1.16s$0.00616.7%

Preview seed data (preview-seed) · not audited Index results · process-local store · mock-eval excluded from ranks

The benchmark is live comparison.

Same prompt, real harness CLIs on Herdr. Herdr is offline here, so this is not a live board. The Index ranks above stay seed-only.

Herdr offline · 0 admitted

Open benchmark anyway

Nothing to compare until Herdr is connected and two kinds are admitted. Use the Index until then.

The playground is the instrument.

Pixel chrome, dither, a prompt, a labeled mock. It writes Ops. It never enters Index ranks. That design stays on this door.

Trust is the product. Scores are the interface.

An index is only as valuable as it is defensible. These four commitments will govern every measurement we publish. Today's Index tables remain labeled preview fixtures while the suite and runner land.

  1. 01

    Frozen suites

    Every published score will identify a versioned, immutable task suite so comparisons stay reproducible against the same work.

  2. 02

    Independent evaluation

    Neutral infrastructure and sealed prompts. Vendors will submit runtimes, not self-reported results.

  3. 03

    Full cost model

    Success will sit beside latency, token cost, and failure modes. Accuracy alone will not decide the ranking.

  4. 04

    Open audit trail

    Each published aggregate will link to underlying traces so a reader can follow any number back to the run that produced it.