first, the word

What is an agent harness?

If it isn’t model weights, it’s harness.

The model writes tokens. Everything else is the harness around it: what it’s allowed to see, what it can reach, what it remembers, what happens when a call fails. You already picked one. Most people pick it by picking a tool and never think about it again.

  outside:  a parent agent · someone asking · a deadline                                       
  none of it touches the model. it dies at the wall; the harness re-tells it.                  
                                                                                               
+-- the harness ----------------------------------------------------+                          
|  I  SYSTEM PROMPT  ·  re-injected every lap                       |                          
|         :                                                         |                          
|         v                                                         |                          
|  +----------------------+           +------------------+          |                          
|  |      THE MODEL       |   call    |  II  TOOLS       |          |                          
|  |    a sealed box      |  ------>  |      read a file |          |                          
|  |    no eyes           |           |      search      | - - - -  |- >  the web              
|  |    no hands          |  curated  |      send        | - - - -  |- >  an inbox             
|  |    no clock of its   |  <- - -   +------------------+          |                          
|  |    own               |                                         |  only a tool reaches out.
|  |                      |                                         |  the model never does.   
|  |   THE ONLY DOORS IN  |                                         |                          
|  |    the system prompt |                                         |                          
|  |    the ask           |                                         |                          
|  |    last lap's result |                                         |                          
|  +----------------------+                                         |                          
|       ^                                                           |                          
|  IV   |  THE SOCKET  ·  models drop in, the harness stays         |                          
|       |  [ Anthropic    ]  [ OpenAI       ]  [ open weights ]     |                          
|       |  drawn equal on purpose · this sheet will not rank them   |                          
|       |                                                           |                          
|       +--- no: go again ---+                                      |                          
|  III  enough yet?          |                                      |                          
|       yes: hand it over  --+---                                   |- >  the answer           
+-------------------------------------------------------------------+                          

swipe or scroll for outside notes →

The model is sealed. The harness is everything else.

part III opens
no part covers
I
The system prompt is written by the harness and goes in again on every lap, not just the first.
II
The harness offers the tools and describes them. Which one gets called, and when, is the model's decision.
III
The ask goes in, the search comes back thin, and the model decides on its own to go again. That decision is the loop.
IV
One loop, any model. The socket is why the choice of model stays yours rather than the lab's.
—
The call is emphatic and the return is thin and dotted, because the harness chose what came back and what it left out.

I–IV are the four things a harness does, after Earendil, “What is a Harness?”, 20 Aug 2026. The five compartments are ours: three of them open part III, and two belong to none of the four.

The model reads nothing else. Everything it sees, the harness chose.

The shorthand for this is “Agent = Model + Harness”, and Earendil call that framing simplistic. They’re right. A two-part split says nothing about the four different jobs a harness actually does: the instructions it injects, the tools it hands over, the loop it runs, and the translation layer that lets you point it at a different model tomorrow.

Our line isn’t that equation. It’s a test you can apply to any piece of the stack: hold it up and ask whether it’s model weights. If it isn’t, it’s harness, and it was somebody’s decision.

If you want the long version, read What is a Harness by Earendil. It is the best writing on harness engineering we know of, and it’s by the people who build pi, one of the three harnesses we have measured. Read it knowing they make one.

Same model. Different harness. Different outcome.

That gap is what we measure here. It’s also why a new harness landing every other week is genuinely annoying: nobody can tell you what changed, so everyone argues from vibes.

why the choice matters

Same model, different harness, different rank.

This is measurable, and somebody already measured it. On the Terminal Bench 2.0 leaderboard, the same model, Opus 4.6, scores far below itself depending on which harness it is running in. Nothing about the model changed. The runtime around it did.

LangChain report taking their own coding agent from Top 30 to Top 5 on that leaderboard by changing only the harness.

Their own product, so read it as a vendor reporting on themselves. LangChain also make Deep Agents, one of the three harnesses on this page. The leaderboard itself is third-party and you can go and look at it.

and what it costs

Same model, same job, twice. The only thing that changed was the runtime around it.

runtimecostresult
first20 min$9broken
second6 h$200playable

Twenty-two times the cost, and the difference between something that runs and something that doesn’t.

One reported run from Anthropic. An anecdote, not a study. We are building the measurements that would replace it.

the harnesses we have opened up

Pi

live

TypeScript on Bun · 1.86 MiB · MIT · v0.84.2

A small kernel on three libraries. It extends itself at runtime, and it ships no permission gate on purpose.

open pi →

fx

live

Zig · one binary · Apache-2.0 · v0.0.3

One binary, nothing else to install. Permissions live inside the harness.

open fx →

Deep Agents

live

Python on LangGraph · MIT · v0.7.7

A middleware stack you fill with files.

open deep agents →

Playground

See how each harness is built. Every block is a concept from its own docs, sized by how much source sits behind it.

one turnrecorded 2026-08-19 · 6.81s19 steps
  turn_start          ->  one model response plus its tool calls
  tool_start   164ms  ->  bash
  tool_end     165ms  ->  bash
  tool_start  1347ms  ->  read
  file_read   1348ms  ->  read package.json  (package.json)
  tool_end    1349ms  ->  read
  tool_start  2205ms  ->  read
  file_read   2206ms  ->  read cli.js  (cli.js)
  tool_end    2207ms  ->  read
  tool_start  3336ms  ->  read
  ...             9 more steps

Times are wall-clock from the start of the turn, unsmoothed. The gap before an edit is the model thinking, and it is shown as it happened.

the index

Ranks come after measurement, not before it.

Here’s the awkward part: the scoreboard this site is named after isn’t open. What’s there today is a preview seed on a frozen pack. Fine for checking that the machinery works, useless for picking a harness. It says so everywhere it appears, and it’ll keep saying so until real runs replace it.

[see all eleven in the directory]

hibench · coming

One task, one model, many harnesses, graded the same way. It hasn’t run yet, so there’s nothing to publish and nothing here but this sentence.

[join the waitlist]

hi-harness

The harness for harnesses. It drives several of them at once on one model, watches the runs and grades them. hibench gets built on top of it.

status: in development