A new pattern is emerging in organizations. Someone (increasingly a non-coder) vibe-codes a tool which cosmetically accomplishes a task but which is nowhere near production readiness. That POC receives much fanfare, creating the perception that the tool is "almost done" (ok, maybe this isn't such a new pattern). Then, some poor software engineer or team is tasked with fixing all of the security holes, architecture & data model problems, code quality issues, and non-existent deployment strategy to get it production ready (aka, the hard parts).
That work is nowhere near as flashy or likely to be paraded at an all-hands meeting, but will certainly invite comments like "why is it taking so long?" That feeds the idea that production quality tools should be ready in a days or weeks, which causes leadership to second guess effort & headcount requirements for future projects, hampering org output.
But...AI!
"But wait!", you say, "AI has drastically reduced development time...." Has it? I had a very interesting experience in planning the architecture for a recent project.
I tried an AI-first approach to the architecture & effort estimation as a BDUF and full time estimate were required. The requirements forced an inherently complex deployment, security, and data redundancy story. I fed Claude the direct requirements documents, went through many iterations with both Opus & Fable, and had it create a spreadsheet which broke the work into granular chunks no longer than 1 person-week each. Then, I manually verified/fixed all of those chunks and each formula in the spreadsheet.
After it had generated the timeline, I asked Claude (arguable the most biased entity I could have asked), "how much time will we save using an AI-first coding approach? Use only peer-reviewed sources in your answer." It found only the METR 2025 study and its 2026 follow up, so I instructed it to expand the search to studies from reputable sources. It aggregated studies from Microsoft, Google, and others (list below). Claude's conclusion from those sources? On a complex project like this one, spec'd with senior developers, less than a 10% improvement.
Let that sink in. On average, those studies showed a ~25% improvement, which is incredible, but hardly the multiple the hype would have us believe. Back in 2025, using older models, the original METR study actually showed a loss in velocity, even though the participants felt they were moving faster.
Why isn't the gain more substantial? In my opinion, because AI is currently very good at writing code. It is not great (yet) at producing a robust, secure, scalable working product. It will give you a soapbox derby car in no time, but in a software development org, what you need is probably closer to a Formula 1 car.
And my subjective experience from this exercise? After estimating the person-week chunks, I had it cross-reference to the skillsets on our team, then assign developers to tasks to create a per-person timeline. I then added the facet that we were considering a third party vendor for a portion of the scope and asked it to give an alternative timeline for that. At that point, it fell over completely, and I had to go through the spreadsheet with a fine tooth comb. Overall, it felt like the AI-first approach took longer.
Are We There Yet?
I am not an AI doubter; I use it daily, coding and otherwise (though not for writing blog posts 😄). In fact, I consider myself an AI optimist. We are reaching the point in human history where we need something like AI. Software projects are reaching a level of complexity which makes them difficult to reason about. Things like operating systems have been there for a while, but other areas are catching up.
I see computer science going the way of medicine in which there doctorates available in numerous sub-fields. In any field, a doctorate is already the culmination of 20-ish years of education and the accumulation of human knowledge is only speeding up. We can't expect a human to undertake much more education before entering the workforce. We need AI. But, for the moment at least, I am not convinced it will get a non-trivial software project to market all that much faster.
Claude's Sources
Primary studies
-
METR early-2025 RCT — Becker, Rush, Barnes, Rein,
Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (the 19% slowdown result).
Paper: https://arxiv.org/abs/2507.09089 · Summary: https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ -
METR February 2026 update — Becker, Rush, Cunningham, Rein, Mahamud,
We are Changing our Developer Productivity Experiment Design (the walkback: raw −18%/−4% speedup estimates, selection-effect caveats, 57 devs / 800+ tasks).
https://metr.org/blog/2026-02-24-uplift-update/ -
Paradis et al. (Google), ICSE-SEIP 2025 —
How much does AI impact development speed? An enterprise-based randomized controlled trial (N=96, ~21% speedup with wide CI, not significant after controlling for proficiency/familiarity). Peer-reviewed.
arXiv: https://arxiv.org/abs/2410.12944 · Published: DOI 10.1109/ICSE-SEIP66354.2025.00060 -
Cui et al., Management Science —
The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers (N=4,867 at Microsoft/Accenture/Fortune 100; +26.08% completed tasks, SE 10.3%; larger gains for juniors). Peer-reviewed.
https://pubsonline.informs.org/doi/fpi/10.1287/mnsc.2025.00535 · Preprint: https://papers.ssrn.com/abstract=4945566 -
Microsoft agentic-CLI rollout study (July 2026) — Murphy-Hill, Butler, Savelieva,
Adoption and Impact of Command-Line AI Coding Agents (tens of thousands of engineers, telemetry-identified usage of Claude Code / Copilot CLI; ~+24% merged PRs, +14.5% to +33.7%, persistent over four months). Preprint.
https://arxiv.org/abs/2607.01418 -
Peng et al. 2023 —
The Impact of AI on Developer Productivity: Evidence from GitHub Copilot (55.8% faster on greenfield toy task; weakest external validity, included for completeness since it anchors the optimistic end).
https://arxiv.org/abs/2302.06590
Synthesis / secondary (peer-reviewed)
-
Mohamed, Assi, Guizani,
ACM TOSEM — The Impact of LLM-Assistants on Software Developer Productivity: A Systematic Review and Mapping Study (39 peer-reviewed studies, 2014–2024; majority report gains on routine work, code-quality effects unresolved).
https://doi.org/10.1145/3809494 · Preprint: https://arxiv.org/abs/2507.03156 -
Empirical Software Engineering (Springer, 2026) —
Echoes of AI: Investigating the downstream effects of AI assistants on software maintainability — the peer-reviewed source for the study-by-study comparisons I cited, including the He et al. Cursor difference-in-differences result whose velocity lift
faded by month three.
https://link.springer.com/article/10.1007/s10664-026-10889-1
Comments
Post a Comment