hibench · not open

One task. One model. Every harness.

Give the same job to the same model three times and change only the runtime around it. Whatever moves is the harness. That’s the whole design, and it’s the only way we know to say one harness beats another without hand-waving.

the shape of a run


        one task
            |
   +--------+--------+--------+
   |        |        |        |
   v        v        v        v
  pi       fx    deep agents  ...        <- the only thing that changes
   |        |        |        |
   v        v        v        v
 graded   graded   graded   graded       <- same grader, same criteria

the method: ablation

Take a piece out and see what breaks.

A score on its own teaches you nothing. Ablation does: hold the model still, pull one subsystem out, and measure what it cost you. Those subsystems are the same districts the playground already draws, so a result points at somewhere you can go and look.


   hold the model.  remove one part.  measure the drop.

   full harness          ###############   baseline
   - memory              ##########        how much did forgetting cost
   - tool surface        ######            how much did the tools carry
   - permissions gate    #############     did the guardrail cost anything

Sometimes pulling a part out changes nothing. That’s a result too, and it goes up with the same weight as a dramatic one.

two columns, no winner

Outcome and cost are different questions.

One harness solves it for two dollars, another for twenty. Which one you want depends on what you’re doing, so there’s no honest way to squash that into a single number. We report both and leave the choice with you.

mock · what a result row will look like · no run has happened
harnessmodeloutcomecostorigin
—held fixedn/7 tasks$ · tokens · wall clockrunner
—held fixedn/7 tasks$ · tokens · wall clockrunner

status

It’s blocked on hi-harness. Running one task across several harnesses on one model needs the runtime that drives them all, and we’re still building it. Until it runs there’s nothing to publish, and we’d rather show you nothing than a number we can’t stand behind.

And when it does run, some harnesses still won’t carry a score. Refusing to score one is a result we publish, not a gap we’re waiting to fill.

status: not open · blocked on hi-harness · no runs recorded

hibench runs one task across harnesses on the same model and grades the results. It is not open yet. One email when it is.