Peng et al. - controlled greenfield task
95 freelance developers built a JavaScript HTTP server. The AI group completed it in 71.2 minutes versus 160.9 minutes.
Single bounded task; vendor-affiliated authors; wide confidence interval.
Evidence, not impressions
The findings are strong in some areas, contested in others, and absent in a few. Each result below keeps its source tier and its most important limitation.
AI consistently increases output volume. Whether it shortens time to a correct outcome depends on task familiarity, codebase maturity, and downstream verification.
95 freelance developers built a JavaScript HTTP server. The AI group completed it in 71.2 minutes versus 160.9 minutes.
Single bounded task; vendor-affiliated authors; wide confidence interval.
Three randomized trials at Microsoft, Accenture, and a Fortune 100 manufacturer followed 4,867 developers over two to eight months.
Measured throughput; researchers had no access to produced code quality.
16 experienced open-source maintainers completed 246 tasks in repositories they had worked in for an average of five years.
Small, early-2025 snapshot. METR's February 2026 follow-up was judged an unreliable signal, because developers increasingly declined to work without AI. Its raw estimates point the other way, toward a speedup, and METR now believes developers are likely faster in 2026. Treat 19% as a finding about mature codebases in early 2025, not a current measure.
Different methods - generated scenarios, real GitHub snippets, exploit-based benchmarks, and model trend tests - converge on the same conclusion.
1,689 programs generated across 89 security-relevant scenarios.
Early Copilot generation, but later research has not displaced the central result.
733 AI-generated snippets were analyzed across 43 CWE categories.
JavaScript measured 24.2%; snippet-level analysis cannot capture every architectural flaw.
392 backend tasks across 14 frameworks tested whether generated implementations were correct and exploitable.
Incorrect or insecure; approximately half of correct solutions were still exploitable.
More than 150 models tested from 2023 through Spring 2026.
Industry research, though the method and longitudinal comparison are published.
10,517 confirmed applications identified; 200 deployed web apps audited with agents, exploit probes, and two human experts.
Major 2026 preprint, not yet peer-reviewed; unusually rigorous audit design.
The human problem is not just trust. AI changes the work from producing to verifying - and verification requires skills that delegation can weaken.
AI-assisted learners versus control in Anthropic's Trio experiment.
for generation followed by active interrogation of the code.
for repeatedly asking AI to solve the errors it introduced.
cite AI answers that are "almost right, but not quite."
Package hallucination is repeatable. Attackers can register names that models reliably invent - a supply-chain technique called slopsquatting. The rate has fallen sharply on 2026 frontier models, but the attack surface has not closed.
of recommended packages did not exist across 16 models tested in 2024.
still recommended nonexistent packages in that same study.
of hallucinated names recurred in all ten reruns.
across five frontier models, a large drop. But 53 names invented by all five remained registrable after disclosure.
GitClear's observational analysis of 623 million changes reports less refactoring, more copy-and-paste, more short-term churn, and fewer cross-file calls. It is a strong directional signal - not causal proof - and should be read with that limitation intact.
A credible account must show the gaps as clearly as the findings.