Hot take: The engineering process matters more than which coding model you're using
Been debating with myself on posting and asking this, but decided to YOLO it anyhow. I've used Opus, Sol 5.6/6.1, and Astra across some pretty substantial software projects, and I've gotten to the point where I honestly don't see a huge difference in the end results. They all generally get the job done. They all miss things. And after a round or two of code reviews, they usually end up in about the same place. I think a lot of that comes down to how I'm using them, though. I'm not just giving them a prompt and hoping they figure everything out. On larger projects, my system instructions can be 10k–20k tokens, backed by hundreds of pages of documentation covering architecture, requirements, coding standards, testing, design decisions, etc. They have strict rules about following that documentation, keeping it updated, and reviewing their own work. I also have agents take on different review roles to look at things from different perspectives. Architecture, security, code quality, missed requirements, that sort of thing. When something gets missed, it gets caught in review, fixed, and ideally the documentation or instructions get improved so it doesn't happen again. It's a lot of setup upfront, but once everything is established, I can pretty much hand off entire projects instead of individual coding tasks. I've spent years managing engineering and data science teams, and I've found myself approaching agents pretty much the same way. I'm setting priorities, defining requirements, assigning work, reviewing results, and making sure the overall system is working. Just like on my teams, I still have to review, discuss and suggest improvements to what is produced. The actual coding has become a pretty small part of my involvement. And frankly, the consistency and quality of the output have been better than what I've personally experienced with traditional development processes. That's not a knock on developers. It's just a very different way of organizing the work. There are definitely differences between the models in speed, cost, context handling, and how often they get things right on the first try. I'm not arguing otherwise. But I wonder if we're spending too much time comparing which model writes the best code and not enough time talking about how to build engineering processes around them. At least for me, once the documentation, workflows, and review processes are solid, switching models doesn't seem to change the outcome all that much... outside of the speed dips. Curious if anyone else working on larger, long-running projects has gotten to this point, or if you're still seeing big differences between models even with structured workflows and reviews.