Halv gets 57.1% lower model cost across 42 paired coding task runs
Disclosure: I am building Halv, a paid tool for coordinating coding agents. I compared its routed workflow with a vanilla Codex workflow on the first 14 SWE-rebench tasks, with three repetitions per task. I developed this Halv workflow through a lot if iterating and testing. An investigation agent assessed each task, and JEV selected the worker model and reasoning effort based on that assessment Across 42 paired runs: Estimated model cost: $144.67 with HALV versus $337.50 with vanilla. Successful runs: 25/42 for both workflows. HALV used approximately 3.49× more tokens, while shifting work toward cheaper models. This resulted in 57.1% lower cost This is an interim result from 14 of 111 planned tasks. It does not establish equivalent performance or guarantee savings on other workloads, but gives clear insight on how Halv helps reduce cost in complex tasks. JEV agent selection is included in that subscription, with no separate routing fee. The technical report includes per-task results, methodology, excluded attempts, and downloadable evidence.