Where did all the productivity in AI go?

The experiments say AI works. The accounting says it hasn’t. Both are right, and the reasons matter.

In 2025, METR paid 16 experienced open-source developers $150 an hour to work on 246 real issues in their own mature repositories—codebases averaging over a million lines. Half the tasks allowed AI; the other half didn’t1. The developers believed that the AI would make them 24% faster. After doing the work, they believed they were 20% faster. The test measured 19% slower2.

That is a 39% difference between what they believed and what the test measured. When I read this report last year, I was shocked. While I don’t believe AI will take away huge swaths of jobs as some do, I still believe AI will offer significant productivity gains. Trinsic provides AI hosting and applications, so we have a vested interest in AI working.

Two types of evidence on AI productivity disagree. At the Micro level (one worker, one task), you can find abundant evidence of productivity gains. The Quarterly Journal of Economics reported3 in 2025: +15% more issues resolved per hour and a +30% productivity increase for the least experienced customer support agents. These are real productivity gains.

Yet when around 6,000 executives across four countries were surveyed about revenue gains from AI, they reported only a cumulative gain of 0.29%4. Worse, utilization-adjusted US total factor productivity grew only 0.07% over the four quarters ending Q1 2026.

 

Level What the best evidence says The number
Micro — one worker, one task Real, replicated, sometimes large +15% issues resolved per hour, ~+30% for the least experienced — 5,172 customer-support agents, Quarterly Journal of Economics5
Firm — does it reach the P&L? Mostly not yet Nine in ten of ~6,000 executives across four countries report no productivity impact over three years; measured cumulative gain 0.29%6
Macro — does it reach the statistics? Not detectably Utilization-adjusted US total factor productivity grew 0.07% over the four quarters ending Q1 20267

 

The data isn’t asking whether AI increases productivity; it is asking why it disappears as you move up the ladder.

Before we answer this question, we need to take a step back and define what we mean by a ‘productivity increase.’ They generally measure productivity in two ways (there are others, but these are the most commonly cited). The first is Labor Productivity (LP), and the Second Is Total Factor Productivity (TFP).

The distinction matters. LP rises for reasons unrelated to anyone getting better at their job. If you buy your employees better equipment and each worker produces more per hour, great. But the equipment is a capital expense and could therefore be a wash for the bottom line. TFP is what’s left after you count what it cost you for the productivity gain. If AI is truly providing productivity gains, then we should see it in the TFP. Here is the problem: we don’t.

 

 

From 2024 through 2025, amid an AI-accelerated implementation, LP and TFP dropped, which isn’t what we should see if AI is providing large productivity gains. TFP growth halved in the heaviest year of AI deployment yet. The data center buildout held LP up, contributing nearly 39% of US GDP growth in the first nine months of 2025. This measures investment, not productivity gains, which is why it doesn’t show up in TFP8.

Let me give you an example. If one of our clients buys forty new laptops for their employees, and they are much faster than the old ones, there will be some productivity gains (LP). Still, when you add the capital expense, it will wash out at the bottom line (TFP). This seems to be what’s happening with AI. Making someone’s job easier (LP) doesn’t always translate to a revenue increase (TFP). GDP growth and a productivity boom don’t always go hand in hand.

What makes the data more complex is that it’s all over the place.

 

I don’t think anyone on this list is lying. Every entry appears to be solid data and real. The harder question is why they’re getting such different answers. The devil is in the details, and there are some things to note. METR measured a -19% reduction in productivity, but then, ten months later, surveyed technical workers and found a median 3X speedup9,10. This seems to be a pattern: surveyed employees claim AI production gains, but when measured, they disappear or shrink significantly.

The Danish study is the strongest null result in the literature because it uses administrative tax records, not surveys: 25K workers across around 7K workplaces in 11 occupations exposed to AI. Workers report real-time savings. Earnings and hours didn’t move11.

This is why you should evaluate vendor reports carefully. This includes even the author’s company (Trinsic Technologies Inc.). To see true ROI and productivity gains with AI, you need to measure work before and after AI implementation. This also includes understanding what AI is good at and what it isn’t.

In a pre-registered randomized controlled trial, 758 BCG consultants found that when AI was used within its capability range12 it delivered +12.2% more tasks completed, +25% faster completion, and +29.9- 33.9% higher quality—real, replicated results.

However, on a task deliberately built to sit just past the edge of what AI can handle, consultants using GPT-4 were 19 percentage points less likely to be correct. Worse, those who received prompt-engineering training did better than those who did not. To add insult to injury, the wrong answers were rated more coherent and more persuasive than the unassisted wrong answers. AI hallucinations remain a real problem.

AI is jagged intelligence, doing well in some tasks and not well in others. While LLMs have improved, this remains a problem. The following costs don’t show up in the ROI models:

METR’s developers accepted fewer than 44% of AI suggestions and spent 9% of their total task time reviewing and cleaning AI output13.

In the Danish data, roughly 17% of workers in AI initiative workplaces report new AI-created tasks — and about 35% of those new tasks are AI quality review and AI ethics compliance14. The tool created a job whose purpose is to check the tool.

The clearest published measurement of skill erosion: across 1,443 unassisted colonoscopies at four Polish centers, adenoma detection fell from 28.4% before AI was introduced to 22.4% after15. Observational, not randomized — say so in the sentence.

 

 

What does this tell us? AI should be used carefully, particularly in regulated industries such as Law Firms, TAS, call centers, medical clinics, etc. Regulatory agencies have been slow to enforce privacy concerns related to AI. This will not last. In many of these industries, wrong answers can have serious consequences.

However, the skeptics don’t get a free pass either. METR walked back its own headline, saying it had to abandon the study after 30-50% of developers admitted they withheld exactly the tasks they thought AI would speed up. METR believes developers are working faster now, but they can’t measure it16.

Many people cite the MIT study that shows 95% of AI implementations fail (I have been guilty of this). However, that is not what the MIT study showed. It showed that only 5% of AI projects actually reach the implementation phase. This doesn’t mean the other 95% necessarily failed; they didn’t reach production for various reasons17,18 .

The AI implementation is real. Some evidence suggests productivity gains at the micro level, but these gains aren’t translating into TFP or the bottom line. Why? Three potential answers fit the data.

The J-Curve – gains arrive only after organizations reorganize around the tool. We’ve seen this curve with other technology adaptations, including the Internet and manufacturing19.

Intensity – Firms haven’t actually bought much. 37% of organizations report any EBIT impact from AI and only 6% attribute more than 5% — both unchanged from 2025, while the share scaling AI across the enterprise rose from 38% to 44%20. We are still early in this wave.

Ceiling – The gains are real but narrow — fewer than 3% of workplace tasks have AI adoption above 50%, and none exceed 70% — and verification, oversight, and rework eat a share of what’s left21.

 

 

It is 2026, and the question is what level of AI productivity gains will be realized. It is still too early to tell. Evidence suggests there will be gains, but they may not live up to the hyperbole. One thing we do know is that only time and the consumer will tell.

 


 

1. Joel Becker, Nate Rush, Beth Barnes & David Rein, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” METR, 10 July 2025; arXiv:2507.09089. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ · https://arxiv.org/abs/2507.09089
2. Joel Becker, Nate Rush, Beth Barnes & David Rein, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” METR, 10 July 2025; arXiv:2507.09089. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ · https://arxiv.org/abs/2507.09089
3. Erik Brynjolfsson, Danielle Li & Lindsey Raymond, “Generative AI at Work,” Quarterly Journal of Economics 140(2): 889–942, 8 April 2025. https://academic.oup.com/qje/article/140/2/889/7990658 — Note: a staggered-rollout difference-in-differences design, not a randomized trial.
4. Ivan Yotzov, Jose Maria Barrero, Nicholas Bloom et al., “Firm Data on AI,” NBER Working Paper 34836, February 2026 (revised March 2026). https://www.nber.org/papers/w34836
5. Erik Brynjolfsson, Danielle Li & Lindsey Raymond, “Generative AI at Work,” Quarterly Journal of Economics 140(2): 889–942, 8 April 2025. https://academic.oup.com/qje/article/140/2/889/7990658 — Note: a staggered-rollout difference-in-differences design, not a randomized trial.
6. Ivan Yotzov, Jose Maria Barrero, Nicholas Bloom et al., “Firm Data on AI,” NBER Working Paper 34836, February 2026 (revised March 2026). https://www.nber.org/papers/w34836
7. Serdar Ozkan, Aakash Kalyani & Nicholas Sullivan, “AI and Productivity: What Firms Are Saying on Earnings Calls,” Federal Reserve Bank of St. Louis On the Economy, 31 July 2026. 8. https://www.stlouisfed.org/on-the-economy/2026/jul/ai-productivity-what-firms-say-earnings-calls — The 0.07% figure is utilization-adjusted TFP over the four quarters ending Q1 2026, sourced to the San Francisco Fed’s 4 June 2026 update.

8. US Bureau of Labor Statistics, Total Factor Productivity — 2025, 19 March 2026. https://www.bls.gov/news.release/prod3.nr0.htm
9. Joel Becker, Nate Rush, Beth Barnes & David Rein, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” METR, 10 July 2025; arXiv:2507.09089. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ · https://arxiv.org/abs/2507.09089
10. Joel Becker, “Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity,” METR, 11 May 2026; 349 technical workers. https://metr.org/blog/2026-05-11-ai-usage-survey/
11. Anders Humlum & Emilie Vestergaard, “Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI,” NBER Working Paper 33777, May 2025 (revised March 2026). https://www.nber.org/papers/w33777 · https://www.rfberlin.com/wp-content/uploads/2026/03/26078.pdf
12. Fabrizio Dell’Acqua, Edward McFowland III, Ethan Mollick et al., “Navigating the Jagged Technological Frontier,” Organization Science 37(2), 11 March 2026. https://pubsonline.informs.org/doi/10.1287/orsc.2025.21838 — Four co-authors were BCG employees, and BCG ran the study on its own consultants. Note also that the widely quoted “over 40% higher quality” comes from the 2023 working paper; the peer-reviewed figure is 29.9–33.9%.
13. Joel Becker, Nate Rush, Beth Barnes & David Rein, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” METR, 10 July 2025; arXiv:2507.09089. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ · https://arxiv.org/abs/2507.09089
14. Anders Humlum & Emilie Vestergaard, “Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI,” NBER Working Paper 33777, May 2025 (revised March 2026). https://www.nber.org/papers/w33777 · https://www.rfberlin.com/wp-content/uploads/2026/03/26078.pdf
15. Krzysztof Budzyń, Marcin Romańczyk, Diana Kitala et al., “Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study,” The Lancet Gastroenterology & Hepatology 10(10): 896–903, 12 August 2025. https://www.thelancet.com/journals/langas/article/PIIS2468-1253(25)00133-5/abstract
16. METR, “We are Changing our Developer Productivity Experiment Design,” 24 February 2026. https://metr.org/blog/2026-02-24-uplift-update/
17. Aditya Challapally, Chris Pease, Ramesh Raskar & Pradyumna Chari, The GenAI Divide: State of AI in Business 2025, MIT NANDA, July 2025. https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf (non-peer-reviewed preliminary findings)
18. Ray Poynter, “Myth Number 2: MIT Showed That 95% of AI Pilots Fail,” NewMR, 31 May 2026. https://newmr.org/blog/myth-number-2-mit-showed-that-95-of-ai-pilots-fail/
19. Kristina McElheran, Mu-Jeung Yang, Zachary Kroff & Erik Brynjolfsson, The Rise of Industrial AI in America: Microfoundations of the Productivity J-curve(s), US Census Bureau CES-WP-25-27, April 2025. https://www.census.gov/library/working-papers/2025/adrm/CES-WP-25-27.html
20. McKinsey & Company, The State of AI: Global Survey 2026, published 25 August 2026; fielded 4 May – 8 June 2026; 1,719 respondents across 97 countries. https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai (consultancy research — McKinsey sells AI transformation services)
21. Federal Reserve Bank of St. Louis, “What Work Does Generative AI Do?”, September 2026; ~14,000 workers across four quarterly waves, August 2025 – May 2026. https://www.stlouisfed.org/on-the-economy/2026/sep/what-work-does-generative-ai-do

 

Whether you’re looking for a dynamic partner on your next tech project, managed IT service providers, or are interested in joining our team of seriously awesome technicians — submit a contact form and we’ll be in touch!

Other blogs