Same model. Different harness. Different outcome.
The runtime around an agent (tools, memory, retries, failure handling) moves the score. Read the labeled preview table, then open the benchmark for a live same-task compare. Numbers below are labeled preview fixtures, not audited rankings (preview-seed).
Fixed model, fixed suite. The harness is the variable.
+38.2 pts success delta, preview
Structured ReAct
preview-model-snapshot
87.8%
Preview
6 seed runs in the process-local preview store
Baseline Loop
preview-model-snapshot
49.7%
Preview
6 seed runs in the process-local preview store
One table. Same work. Comparable harnesses.
Terminal Core seed fixtures on a fixed model snapshot. Preview only; mock-eval playground runs are not included in these ranks. Data class: preview-seed. Live harness comparison is on the benchmark. Catalog structure is on the field.
| # | Harness | Success rate | Mean latency | Cost / task | Reliability |
|---|---|---|---|---|---|
| 01 | Structured ReAct 0.9.4 · 6 preview runs | 87.8% | 1.24s | $0.006 | 100.0% |
| 02 | Baseline Loop 0.3.1 · 6 preview runs | 49.7% | 1.16s | $0.006 | 16.7% |
Preview seed data (preview-seed) · not audited Index results · process-local store · mock-eval excluded from ranks
The benchmark is live comparison.
Same prompt, real harness CLIs on Herdr. Herdr is offline here, so this is not a live board. The Index ranks above stay seed-only.
Herdr offline · 0 admitted
Open benchmark anywayNothing to compare until Herdr is connected and two kinds are admitted. Use the Index until then.
The playground is the instrument.
Pixel chrome, dither, a prompt, a labeled mock. It writes Ops. It never enters Index ranks. That design stays on this door.
Trust is the product. Scores are the interface.
An index is only as valuable as it is defensible. These four commitments will govern every measurement we publish. Today's Index tables remain labeled preview fixtures while the suite and runner land.
- 01
Frozen suites
Every published score will identify a versioned, immutable task suite so comparisons stay reproducible against the same work.
- 02
Independent evaluation
Neutral infrastructure and sealed prompts. Vendors will submit runtimes, not self-reported results.
- 03
Full cost model
Success will sit beside latency, token cost, and failure modes. Accuracy alone will not decide the ranking.
- 04
Open audit trail
Each published aggregate will link to underlying traces so a reader can follow any number back to the run that produced it.