Skip to content

All pages

FrontierHarness Eval

Nine coding harnesses on one model: pass rates cluster, cost spreads 17.5x

AI15d ago
Category
AI
Language
English
Listed
15d ago

FrontierHarness Eval compares nine coding harnesses across 12 configurations on identical software engineering tasks, holding the model and runtime fixed. All 360 trials begin from the same fresh checkpoint restore, so no run benefits from a warm cache. Pass rates land in a narrow band between 50.0% and 66.7%, while median cost per task ranges from $1.05 to $18.34, a 17.5x spread. The published chart plots each harness by cost, speed, and pass rate.

Like this product?