tim.waldin.net ~
❯ blog
blog
2026-04-20
I wanted to know which AI coding agent is actually the best at fixing real bugs, so I spent $642 on around 1b worth of tokens and about two weeks running every model+harness combo I could get my hands on against 42 real bugs. Out of 155 combinations attempted, 98 qualified for...
2026-04-20
Yesterday i wrote about gepa optimizing the system prompt for claude haiku 4.5: 55% → 92% on a 20-bug training set, 65% → 85% on a 9-bug holdout. That was all same-model, same-cli, internal benchmark. Tonight i pushed the same honed prompt through agentelo — the public leaderb...
2026-04-19
Claude haiku 4.5 solves 65% of real github bugs with a 14-word seed prompt. I ran GEPA (a prompt-evolution library out of dspy) against it for 7 hours on 20 bug-fix challenges, and the optimized prompt takes the same model to 85% on 9 unseen bugs it never trained on. 20 points...
read in the terminal instead:
/t/blog