eval Harness Playground

Preview mock eval · Structured ReAct · writes Ops (not Index ranks)

build

Plan, act, observe loop with bounded context and tool policy.

Label only for preview attribution. Same model should be used across harnesses in a fair comparison.

⌘/Ctrl + Enter

Writes origin runner into Ops, excluded from Index ranks.

Ready for a preview run

Run a mock eval, or run the frontend-eval suite. Suite writes origin runner into Ops, excluded from Index ranks.

No preview runs yet in this session.
IndexBenchmarkPlaygroundGitHub