Work
Productivity and why it is hard to measure
The best experiments on AI and productivity disagree about the sign, and the aggregate statistics cannot see the effect at all.
Individual productivity is only measured cleanly where output is countable and identical from unit to unit, which is why every rigorous result in this field comes from call centres, writing tasks or coding exercises. Within those settings, generative AI has produced large measured gains: 14% more support issues resolved per hour, 40% less time on writing tasks, 55.8% faster completion of a scoped programming task. It has also produced the opposite. A randomised trial of experienced open-source developers working in codebases they knew well found AI access made them 19% slower, while those same developers believed it had made them 20% faster. Aggregate national productivity statistics cannot currently attribute any part of their movement to digital tools or AI. For a practical commercial comparison with this kind of workplace measurement, see this overview.
The paradox is older than the technology being blamed for it
The gap between visible computing and invisible productivity gains was named in 1993, formalising Robert Solow's 1987 remark that the computer age was everywhere except in the productivity statistics (Brynjolfsson, Communications of the ACM, 1993). Solow's line was a throwaway sentence in a book review, not a finding, and it is routinely cited as though it were one. The modern restatement argues that general-purpose technologies need complementary intangible investment — reorganisation, skills, process — before output responds, so measured gains lag adoption by years (Brynjolfsson, Rock & Syverson, NBER, 2017). That framework is plausible and close to unfalsifiable over short horizons, because any absence of gains can be attributed to lag.
The aggregate series cannot see it
National productivity accounts cover the whole economy, and they are too noisy and too aggregated to identify a technology.
US labour productivity rose 1.4% in Q2 2026 at an annual rate and 2.2% year over year, against a long-term average of 2.1% since 1947.
Direction: No detectable effect. Strength of evidence: Strong.
Caveat Quarterly productivity is volatile and heavily revised, and the aggregate cannot attribute any part of the change to digital tools or AI.
The UK series points the other way, and illustrates why residual measures should be handled carefully.
UK multi-factor productivity is estimated to have fallen 0.6% in 2024 compared with a year earlier.
Direction: Decrease. Strength of evidence: Mixed.
Caveat Multi-factor productivity is a residual that absorbs all measurement error in capital and labour inputs, and ONS labour inputs have had known quality problems since the Labour Force Survey response-rate collapse.
Where AI raised measured output, and by how much
The largest field study puts AI assistance into a job with countable output and staggered rollout. Its most quoted result is not the average.
Access to a generative AI assistant raised issues resolved per hour by 14% on average, with about 34% gains for novices and minimal gains for the most experienced agents.
Direction: Increase. Strength of evidence: Strong.
Caveat One firm, non-random rollout timing and a scripted, high-volume task, and the skill-levelling result is the part most often over-generalised.
The cleanest randomised design uses short writing tasks with recruited professionals, which buys internal validity at the cost of resembling an exam rather than a job.
ChatGPT access cut average time taken by 40% and raised graded output quality by 18%.
Direction: Increase. Strength of evidence: Strong.
Caveat Short, self-contained, incentivised tasks completed by online-recruited participants, with quality graded by other online raters.
The coding result most often quoted in support of AI tooling comes with an interest to declare.
Developers given GitHub Copilot completed the task 55.8% faster than controls.
Direction: Increase. Strength of evidence: Mixed.
Caveat The authors are employed by Microsoft and GitHub, which sell Copilot, and the study covers a single greenfield toy task with no measurement of code quality, review burden or maintenance cost.
Experienced developers were slower and were certain they were faster
The most useful single result in this literature breaks the link between how productive people feel and how productive they are. It took developers working on real issues in repositories they already maintained — the opposite of a toy task — and randomised AI tool access at the issue level.
AI tool access made developers 19% slower, although they had expected a 24% speed-up beforehand and still believed afterwards that they had been 20% faster.
Direction: Decrease. Strength of evidence: Mixed.
Caveat Only 16 developers, all working in large codebases they knew intimately with early-2025 tooling and limited prior experience of it, and the authors state explicitly that this is not evidence AI fails to speed up developers generally.
Two things follow. The sign of the effect depends on who is working and on what: novices on scripted, countable tasks gain most, and experts on complex work they already understand can lose. And self-reported productivity gains cannot be trusted. Developers who were 19% slower believed they had been 20% faster, which is why survey questions asking people whether AI made them more productive are not measurement — including in vendor reports about workplace tools.
Scaling task-level gains to an economy is a calibration, not a finding
Attempts to convert these experiments into a macroeconomic number inherit every weakness in them. One widely cited exercise puts AI's cumulative total factor productivity gain at under about 0.53% to 0.66% over ten years (Acemoglu, NBER, 2024). It is a calibration: it takes task-level gains from a handful of experiments, several of them the ones above, and scales them by the task-exposure shares used in automation estimates. Other economists using the same framework with different assumptions reach far larger numbers, which is a fact about the assumptions.
The short version
- Generative AI raised issues resolved per hour by 14% across 5,179 support agents, with the gains concentrated among novices (Brynjolfsson and colleagues, 2023).
- A randomised trial of 453 professionals found ChatGPT cut writing time by 40% and raised graded quality by 18%.
- A randomised trial of 16 experienced developers on their own repositories found AI tools made them 19% slower, while they believed they had been 20% faster.
- The 55.8% Copilot speed-up was measured on one greenfield toy task by authors employed by the companies selling the tool.
- The evidence has a hard limit: quarterly national productivity figures are volatile and heavily revised, and cannot attribute any of their movement to AI or to any other digital tool.