hibench · not open
One task. One model. Every harness.
Give the same job to the same model three times and change only the runtime around it. Whatever moves is the harness. That’s the whole design, and it’s the only way we know to say one harness beats another without hand-waving.
the shape of a run
one task
|
+--------+--------+--------+
| | | |
v v v v
pi fx deep agents ... <- the only thing that changes
| | | |
v v v v
graded graded graded graded <- same grader, same criteria
the method: ablation
Take a piece out and see what breaks.
A score on its own teaches you nothing. Ablation does: hold the model still, pull one subsystem out, and measure what it cost you. Those subsystems are the same districts the playground already draws, so a result points at somewhere you can go and look.
hold the model. remove one part. measure the drop. full harness ############### baseline - memory ########## how much did forgetting cost - tool surface ###### how much did the tools carry - permissions gate ############# did the guardrail cost anything
Sometimes pulling a part out changes nothing. That’s a result too, and it goes up with the same weight as a dramatic one.
two columns, no winner
Outcome and cost are different questions.
One harness solves it for two dollars, another for twenty. Which one you want depends on what you’re doing, so there’s no honest way to squash that into a single number. We report both and leave the choice with you.
| harness | model | outcome | cost | origin |
|---|---|---|---|---|
| — | held fixed | n/7 tasks | $ · tokens · wall clock | runner |
| — | held fixed | n/7 tasks | $ · tokens · wall clock | runner |
status
It’s blocked on hi-harness. Running one task across several harnesses on one model needs the runtime that drives them all, and we’re still building it. Until it runs there’s nothing to publish, and we’d rather show you nothing than a number we can’t stand behind.
And when it does run, some harnesses still won’t carry a score. Refusing to score one is a result we publish, not a gap we’re waiting to fill.
status: not open · blocked on hi-harness · no runs recorded
hibench runs one task across harnesses on the same model and grades the results. It is not open yet. One email when it is.