Interactive model
The Quiet Compounding Gap of AI Automation
Follow one role over time. People get faster as they learn, then level off. When they leave, the seat goes empty and the next hire starts over. AI changes the slope, the ceiling, and whether the knowledge stays. Small weekly differences add up to a gap most companies never see on a report.
By CoSystem. Published .
Key takeaways
- Over 5 years, one role staffed by people alone produces 238 units of output. Turnover costs 13 of them (5%), and those weeks never come back.
- AI tools that help a person lift the total by +40%. The gain has a ceiling: AI can only speed up the part of the job it touches, and the knowledge still leaves with the person.
- Agents that hold the work, with one accountable person, keep what they learn. At the research defaults they produce +3,122% more than people alone over 5 years. At the most cautious settings (conservative evidence, one agent) the lift is +46%.
- The difference compounds. A weekly gap that looks small in month one is the largest number on the page by year 5.
1 unit is one week of output from a fully ramped expert in the role.
Graph 1 · People only
Expertise rises, levels off, and walks out the door
Each hire ramps along a learning curve that is steep early and flat later. Individual output has a ceiling. When the person leaves, output drops to zero until the seat is filled, and the new hire starts below their own normal level because the company is new to them.
Graph 2 · AI-assisted · Levels 1 to 4
AI helps the person do the job
People use AI tools and approved apps with AI built in. The plateau rises, new hires start higher and ramp faster, and better models keep raising the plateau. But AI can only speed up the part of the job it touches, so the gain has a ceiling. The person still holds the job, so when they leave, output still drops to zero.
Graph 3 · AI-employed · Levels 5 to 9
Agents hold the job, and the job keeps improving
Agents do the work and a person is accountable for their output. Agents keep everything they learn and do not resign. Because the agents do the work, better models flow straight into output with no ceiling set by a person's hours. At level 9 an improvement loop reviews every run and makes the agents better each month, and those gains compound.
Aggregate output for one role over 5 years
Each line adds up every week of output so far. Small weekly differences turn into large totals, and losses from turnover never come back.
| End of | People only | AI-assisted | AI-employed | AI-employed vs people |
|---|---|---|---|---|
| Year 1 | 47 | 58 | 188 | 4.0× |
| Year 2 | 98 | 125 | 612 | 6.3× |
| Year 3 | 136 | 182 | 1,516 | 11× |
| Year 4 | 187 | 256 | 3,472 | 19× |
| Year 5 | 238 | 334 | 7,672 | 32× |
The nine levels
Model progress works differently for each approach. For AI-assisted levels it speeds up only the share of the job AI touches, so gains approach a ceiling (Amdahl's law). For AI-employed levels the realized rate compounds on the agents' whole output. Levels 1 to 4 rest on controlled studies and field experiments. Levels 5 to 9 are scenario assumptions anchored to the cases cited, because rigorous measurement stops around level 5. Gartner expects over 40% of agentic AI projects to be canceled by the end of 2027, so treat the higher levels as what a well-run system can reach, not what any system will reach.
- Level 1, AI built into tools. Usage data only (smart compose, smart reply). Humlum and Vestergaard 2025 found about 3% time saved.
- Level 2, General AI tools. St. Louis Fed: users save 5.4% of hours. BCG field experiment, on tasks AI handles well: 12.2% more tasks, 25.1% faster; on tasks outside that frontier, 19 points less likely to be correct. Noy and Zhang: 40% less time on short writing tasks.
- Level 3, Scripted AI responder. Company-reported: AI handles up to about two thirds of support chats. Klarna added human agents back in 2025 in a hybrid model.
- Level 4, AI pair for the role. Brynjolfsson et al.: +14% on average, +34% for novices; two months with AI matched six months without. Cui et al.: +26% tasks.
- Level 5, Single agent. METR's early-2025 trial found experienced developers 19% slower with AI. METR now calls that result out of date and expects later tools to help more. Wide uncertainty.
- Level 6, Agent on the team. Scenario. The Cybernetic Teammate study: one person with AI matched a two-person team.
- Level 7, Self-directed agent. Scenario. Vendor reports of 4× to 20× on scoped work. Agents complete about 30% of realistic office tasks (TheAgentCompany) and about 66% of computer tasks (OSWorld).
- Level 8, Team of agents. Scenario. No rigorous measurement yet.
- Level 9, Agents with an improvement loop. Scenario. Measured capability has compounded about 10% a month since 2019 and about 27% a month since 2024 (METR); realized output lags.
Frequently asked questions
- What is the compounding gap in AI automation?
- It is the growing difference in total output between a role done by people alone and the same role done by agents that keep what they learn. Each week the difference is small. Added up over 5 years, and with turnover losses that never come back, it becomes the largest number in the model.
- How much output does employee turnover cost one role?
- At the research defaults (30 months tenure, 8 weeks to refill the seat, 6 months to ramp), turnover costs 13 units over 5 years, or 5% of what the same person would produce if they never left. 1 unit is one week of output from a fully ramped expert in the role.
- What is the difference between AI-assisted and AI-employed work?
- AI-assisted work (levels 1 to 4) keeps a person in the job and gives them AI tools. Gains are real but capped by the share of the job AI can speed up, and they leave with the person. AI-employed work (levels 5 to 9) has agents do the job with one accountable person reviewing it. The agents keep what they learn and do not resign.
- How much more output can agents produce than people alone?
- In this model, between +46% (conservative evidence, a single agent) and +3,122% (base evidence, agents with an improvement loop) over 5 years. Levels 5 to 9 are scenarios, not measured results, so treat the upper end as what a well-run system can reach.
- Is this a forecast for my company?
- No. It is a model of a typical path for one role, built on published research for levels 1 to 4 and labeled scenarios for levels 5 to 9. Use Adjust for my business to set your own tenure, hiring time, ramp time and adoption, and read the result as a direction, not a promise.
Method and sources
A weekly simulation of a single role. Departures happen at the average tenure, so results show a typical path rather than a forecast. Model progress uses the same realized monthly rate for graphs 2 and 3. Levels 5 to 9 replace the person doing the work with agents and keep one accountable person, whose own turnover stalls part of the agents' output.
- Newell and Rosenbloom (1981), Mechanisms of skill acquisition and the law of practice. Heathcote, Brown and Mewhort (2000), The power law repealed, Psychonomic Bulletin and Review.
- Ericsson (2006), The influence of experience and deliberate practice on the development of superior expert performance. Link
- Shaw and Lazear (2008), Tenure and output, Labour Economics. Link
- Groysberg, Lee and Nanda (2008), Can they take it with them? Management Science. Link
- Oxford Economics for Unum (2014), The Cost of Brain Drain (UK). Link
- The Bridge Group (2024), Account Executive Metrics and Compensation report.
- Gardner, Van Iddekinge and Hom (2018), If you've got leavin' on your mind, Journal of Management 44(8).
- SHRM (2025), Recruiting Benchmarking Report. Link
- BLS (2026), Employee Tenure in 2026. Link
- Carta (2018), Employment tenure at startups. Link
- BLS CPS Table 47 (2025), Absences from work. Link
- Humlum and Vestergaard (2025), Large language models, small labor market effects, BFI Working Paper 2025-56.
- Noy and Zhang (2023), Experimental evidence on the productivity effects of generative artificial intelligence, Science.
- Brynjolfsson, Li and Raymond (2023), Generative AI at Work, NBER w31161. Link
- Dell'Acqua et al. (2023), Navigating the Jagged Technological Frontier, HBS WP 24-013. Link
- Cui, Demirer et al. (2025), The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers. Link
- Bick, Blandin and Deming (2025), The impact of generative AI on work productivity, St. Louis Fed. Link
- METR (2025), Measuring AI Ability to Complete Long Tasks. Link
- METR (2026), Time Horizon 1.1. Link
- METR (2025), Early-2025 AI impact on experienced open-source developers. Link
- Dell'Acqua et al. (2025), The Cybernetic Teammate, NBER w33641. Link
- Xu et al. (2024), TheAgentCompany: benchmarking LLM agents on consequential real-world tasks. Xie et al. (2024), OSWorld.
- Gartner (2025), Over 40% of agentic AI projects will be canceled by end of 2027. Link