eval Harness Playground
Preview mock eval · Structured ReAct · writes Ops (not Index ranks)
Plan, act, observe loop with bounded context and tool policy.
Label only for preview attribution. Same model should be used across harnesses in a fair comparison.
⌘/Ctrl + Enter
Writes origin runner into Ops, excluded from Index ranks.
Ready for a preview run
Run a mock eval, or run the frontend-eval suite. Suite writes origin runner into Ops, excluded from Index ranks.
No preview runs yet in this session.