Pi
liveTypeScript on Bun · 1.86 MiB · MIT · v0.84.2
A small kernel on three libraries. It extends itself at runtime, and it ships no permission gate on purpose.
open pi →If it isn’t model weights, it’s harness.
The model writes tokens. Everything else is the harness around it: what it’s allowed to see, what it can reach, what it remembers, what happens when a call fails. You already picked one. Most people pick it by picking a tool and never think about it again.
outside: a parent agent · someone asking · a deadline none of it touches the model. it dies at the wall; the harness re-tells it. +-- the harness ----------------------------------------------------+ | I SYSTEM PROMPT · re-injected every lap | | : | | v | | +----------------------+ +------------------+ | | | THE MODEL | call | II TOOLS | | | | a sealed box | ------> | read a file | | | | no eyes | | search | - - - - |- > the web | | no hands | curated | send | - - - - |- > an inbox | | no clock of its | <- - - +------------------+ | | | own | | only a tool reaches out. | | | | the model never does. | | THE ONLY DOORS IN | | | | the system prompt | | | | the ask | | | | last lap's result | | | +----------------------+ | | ^ | | IV | THE SOCKET · models drop in, the harness stays | | | [ Anthropic ] [ OpenAI ] [ open weights ] | | | drawn equal on purpose · this sheet will not rank them | | | | | +--- no: go again ---+ | | III enough yet? | | | yes: hand it over --+--- |- > the answer +-------------------------------------------------------------------+
swipe or scroll for outside notes →
The model is sealed. The harness is everything else.
I–IV are the four things a harness does, after Earendil, “What is a Harness?”, 20 Aug 2026. The five compartments are ours: three of them open part III, and two belong to none of the four.
The model reads nothing else. Everything it sees, the harness chose.
The shorthand for this is “Agent = Model + Harness”, and Earendil call that framing simplistic. They’re right. A two-part split says nothing about the four different jobs a harness actually does: the instructions it injects, the tools it hands over, the loop it runs, and the translation layer that lets you point it at a different model tomorrow.
Our line isn’t that equation. It’s a test you can apply to any piece of the stack: hold it up and ask whether it’s model weights. If it isn’t, it’s harness, and it was somebody’s decision.
If you want the long version, read What is a Harness by Earendil. It is the best writing on harness engineering we know of, and it’s by the people who build pi, one of the three harnesses we have measured. Read it knowing they make one.
Same model. Different harness. Different outcome.
That gap is what we measure here. It’s also why a new harness landing every other week is genuinely annoying: nobody can tell you what changed, so everyone argues from vibes.
This is measurable, and somebody already measured it. On the Terminal Bench 2.0 leaderboard, the same model, Opus 4.6, scores far below itself depending on which harness it is running in. Nothing about the model changed. The runtime around it did.
LangChain report taking their own coding agent from Top 30 to Top 5 on that leaderboard by changing only the harness.
Their own product, so read it as a vendor reporting on themselves. LangChain also make Deep Agents, one of the three harnesses on this page. The leaderboard itself is third-party and you can go and look at it.
Same model, same job, twice. The only thing that changed was the runtime around it.
| run | time | cost | result |
|---|---|---|---|
| first | 20 min | $9 | broken |
| second | 6 h | $200 | playable |
Twenty-two times the cost, and the difference between something that runs and something that doesn’t.
One reported run from Anthropic. An anecdote, not a study. We are building the measurements that would replace it.
TypeScript on Bun · 1.86 MiB · MIT · v0.84.2
A small kernel on three libraries. It extends itself at runtime, and it ships no permission gate on purpose.
open pi →Zig · one binary · Apache-2.0 · v0.0.3
One binary, nothing else to install. Permissions live inside the harness.
open fx →Python on LangGraph · MIT · v0.7.7
A middleware stack you fill with files.
open deep agents →See how each harness is built. Every block is a concept from its own docs, sized by how much source sits behind it.
turn_start -> one model response plus its tool calls tool_start 164ms -> bash tool_end 165ms -> bash tool_start 1347ms -> read file_read 1348ms -> read package.json (package.json) tool_end 1349ms -> read tool_start 2205ms -> read file_read 2206ms -> read cli.js (cli.js) tool_end 2207ms -> read tool_start 3336ms -> read ... 9 more steps
Times are wall-clock from the start of the turn, unsmoothed. The gap before an edit is the model thinking, and it is shown as it happened.
Here’s the awkward part: the scoreboard this site is named after isn’t open. What’s there today is a preview seed on a frozen pack. Fine for checking that the machinery works, useless for picking a harness. It says so everywhere it appears, and it’ll keep saying so until real runs replace it.
[see all eleven in the directory]One task, one model, many harnesses, graded the same way. It hasn’t run yet, so there’s nothing to publish and nothing here but this sentence.
[join the waitlist]The harness for harnesses. It drives several of them at once on one model, watches the runs and grades them. hibench gets built on top of it.
status: in development