Trend Merch · free brief

What the 2026 research actually says about AI coding

Six findings from 2026 studies, in plain English, with an honest note on how strong each one is. No email required, nothing to sign up to.

Read this bit first

Most of the studies below are arXiv preprints. That means they are published but have not been through peer review. They are still worth reading, but a preprint is weaker evidence than a peer-reviewed paper, and anyone who tells you otherwise is selling something. Each entry says which it is.

Where a number is self-reported rather than measured, it says so. That distinction matters more than most headlines admit.

1. The job changed shape

Preprint · self-reported
82%

of engineers reported spending less time writing code. The researchers found a broader shift in focus from creation to verification, and proposed a name for the new work: supervisory engineering work, meaning directing, evaluating and correcting AI output.

Vella, A. and Blincoe, K. (2026). The Impact of AI Coding Assistants on Software Engineering: A Longitudinal Study. Two surveys six months apart, 158 then 101 participants, matched cohort of 95. arXiv:2605.23135

2. More productive, worse experience

Preprint · self-reported
14% → 27%

In the same study, 84% reported their productivity had improved, and that held steady across both time points. But among the matched participants, the share reporting a worsened developer experience in at least one dimension nearly doubled. What eroded was flow state and cognitive load. Feedback loops improved.

The authors call this the productivity-experience paradox. If you have felt faster and flatter at the same time, that is the finding.

Vella, A. and Blincoe, K. (2026). arXiv:2605.23135

3. AI code gets touched less

Preprint · repository mining

Across over 1,000 files and about 3,200 changes in 100 popular repositories, AI-generated files received less frequent maintenance than human-authored code, with updates affecting only a small fraction of file size. The most common changes to AI code were feature extensions, whereas human updates focused on bug fixes. Human developers did the large majority of the maintenance either way.

Worth sitting with: code nobody touches is sometimes code nobody understands. The study measures frequency, not quality.

Sawada, S. et al. (2026). To What Extent Does Agent-generated Code Require Maintenance? An Empirical Study. arXiv:2605.06464

4. A moderate productivity effect; no clear learning effect

Preprint · meta-analysis
g = 0.33

A meta-analysis pooled 23 studies reporting 27 effect sizes (ACM, arXiv, Scopus, Web of Science; studies published 2019 to 2025). It found a statistically significant, moderate positive effect on developer productivity, Hedges' g = 0.33, 95% CI [0.09, 0.58], with substantial variation between settings.

On learning it found no statistically significant effect: g = 0.14, 95% CI [-0.18, 0.47], a confidence interval that crosses zero.

The detail that should interest you most: gains were larger in controlled experiments and smaller in open-source and enterprise contexts. The further you get from a lab, the smaller the effect.

Maier, S., Gunzenhäuser, M., Schweisthal, J., Schneider, M. and Feuerriegel, S. (2026). A meta-analysis of the effect of generative AI on productivity and learning in programming. arXiv:2605.04779

5. Somebody measured what it does to your brain

Preprint · biometric · students, not professionals

A within-subjects crossover study at two universities (Bari and Copenhagen) recorded EEG, eye-tracking, electrodermal activity and heart rate variability while people programmed with and without AI assistance.

Under AI assistance the EEG theta/alpha ratio was lower in the first task and blink rate was higher in the second, both of which the authors read as reduced cognitive engagement when developers offload the generative effort to the model. What they take from it is the interesting part, and they put it carefully: the findings suggest that AI-assisted programming is not a faster version of solo coding but a cognitively distinct activity.

Read this with the caveat. The participants were undergraduate and graduate students, not professional developers, and the abstract does not state how many took part. Treat it as a signal about the activity, not a measurement of your working day.

Burelli, P., Calefato, F., Grassi, D., Hristova, M.Y., Novielli, N., Romano, A. and Tell, P. (2026). Using Biometrics to Understand AI-Assisted Coding Performance and its Perception. arXiv:2606.20598

6. The number you will see quoted everywhere, and why to be careful

You will see figures circulating like "98% more pull requests but 91% longer review times" and "a 19% slowdown for experienced developers". Those numbers appear in a 2026 paper, but that paper is a literature review of 67 sources. It is quoting other people's studies, not reporting its own data.

If you repeat those figures, cite the underlying studies, not the review. This is the single most common way a real finding turns into a misattributed one.

The Productivity-Reliability Paradox (2026), a multivocal literature review. arXiv:2605.01160

What to take from all this

Building got faster. The work moved rather than disappeared, toward reviewing, steering and verifying. The evidence that this makes you more effective is mixed, and at least one study suggests the experience of the job is quietly getting worse even as output rises.

None of that means stop. It means the bottleneck moved, and it is worth knowing where it went.

single until series b

The embroidered dad hat for people who are still locked in. £28, five colours, embroidered in the UK.

Have a look