LLM-as-a-Verifier keeps pushing the frontier of cost vs. capability
On Terminal-Bench 2.1, it made DeepSeek V4 Flash accuracy go from 79% → 88%, while being 4-11x cheaper than competitors.
As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low cost.
On Terminal-Bench 2.1, it made DeepSeek V4 Flash accuracy go from 79% → 88%, while being 4-11x cheaper than competitors.
As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low cost.