And with a good plan that the implementation agent can follow you can also let it be implemented with models that cost less and still get good results. If the scope of the plan is too large, break it down into smaller parts. What I miss is these agents being able to brainstorm and discuss with me like you would in a team. Although part of my planning phase is always to ask the agent at the end if there is anything to decide and that works well.
I would love to see a comparison bench of "strong model makes a plan" and "smaller models writes the code" to showcase your theory.
It would also make it a lot more interesting to see all the "mid-sized" models benched on this particular "downstream" workload.
And with a good plan that the implementation agent can follow you can also let it be implemented with models that cost less and still get good results. If the scope of the plan is too large, break it down into smaller parts. What I miss is these agents being able to brainstorm and discuss with me like you would in a team. Although part of my planning phase is always to ask the agent at the end if there is anything to decide and that works well.
It's nice that the benchmark seems to use high reasoning, the level that we mostly use, while benchmarks out there often use max reasoning.
Excellent article! The Kilo team has gotten better and better at model testing, backing verdicts with screenshots and concrete data