The Human and Machine Company Free → Read it
← Writing

Jamming the Funnel

Producing work got nearly free while checking it still runs at human speed, through people whose judgment cannot be bought in bulk. That is where AI programs back up, at every company size. Here is the math on the jam and what clears it.

BetterUp Labs and Stanford’s Social Media Lab surveyed 1,004 full-time desk workers in September 2025 and found that 40 percent had been handed AI-generated work that looked finished and wasn’t. Sorting out one instance took the person who received it about an hour and fifty minutes. That cleanup ran roughly 20 minutes longer than it would have taken the sender to do the work themselves.

The task never got faster. It moved downstream and grew on the way.

Those same workers estimate that about 15 percent of everything they now receive arrives in that condition, and 53 percent admit to sending some of it themselves. That is most of a workforce, standing on both ends of the same problem.

Call it a funnel jam. Every company has a step where work gets produced and a step where someone checks it before it counts. AI dropped the cost of the first step close to zero, across every function, inside about two years. The second step held still, because checking runs on the judgment of senior people with full calendars. For any one workflow, that is a handful of people, and the same handful it was two years ago. Production scaled with the software. Review scaled with your headcount.

90%of technology professionals surveyed report using AI at workDORA, Google Cloud, 2025
24%express a lot of trust in the code it producesDORA, Google Cloud, 2025

Both numbers come from the same report and describe the same working day. Adoption is close to universal, and about a quarter of users will say they trust the output much. Everyone is producing with a tool they do not fully believe, so somebody has to look at all of it.

Every function has a review step

The researchers called the unfinished work “workslop.” The plain version travels better in a board conversation: work that arrives looking finished and isn’t, so the person receiving it does the job over.

The load lands on whoever does the checking. In the same research, 54 percent of managers reported receiving it against 38.5 percent of individual contributors. Managers are more exposed because reviewing is what the job is.

Who receives unfinished AI work
The load lands on the reviewers.
SHARE RECEIVING UNFINISHED AI WORK IN THE PAST MONTH54%38.5%MANAGERSINDIVIDUAL CONTRIBUTORS
BetterUp Labs and Stanford Social Media Lab, survey of 1,004 full-time U.S. desk workers, September 2025.

It costs something beyond time. Recipients rated the senders less capable and less reliable, and 42 percent rated them less trustworthy. A tool sold to make teams faster is making colleagues think less of each other. That is a management problem wearing a technology badge.

2 hrsaverage time to clean up one piece of AI work that arrived unfinished, about 20 minutes longer than if the sender had done itBetterUp Labs and Stanford Social Media Lab, 2025

BetterUp put the cost at more than $9 million a year for an organization of 10,000 people. Apply their per-employee estimate of $186 a month to a 400-person company and the bill comes to roughly $890,000. The smaller figure is my math on their per-employee number, and the larger one is theirs. Either way it is the kind of cost that never appears in a business case, because no line item carries its name. It shows up as slower approvals, later closes, and senior people who seem oddly underwater in a year when everything was supposed to get easier.

The pattern showed up before anyone had a name for it. Upwork’s Research Institute surveyed 2,500 executives, employees, and freelancers across four countries and published in July 2024, more than a year ahead of the BetterUp study. It found that 77 percent of employees using AI said the tools had added to their workload. Thirty-nine percent reported spending more time reviewing or moderating AI-generated content. The BetterUp researchers put the mechanism in a single sentence on their research page: “Rather than saving time, it leaves colleagues to do the real thinking and clean-up.”

Run down the functional list and the same shape appears everywhere. Finance produces reconciliations and a controller signs them. In legal, the contract language goes to whoever has authority to approve it. Marketing’s claims stop at a compliance reviewer. Clinical, underwriting, quality, and safety all work this way, and in most of them the reviewer exists because a regulator or an insurer requires a named human to be accountable. That accountability is the part AI cannot absorb. The queue forms there.

Legal is the clearest case, because the production side moved so fast. FTI Consulting and Relativity found that 87 percent of general counsel reported generative AI use inside their teams in 2026, against 44 percent the year before. Adoption doubled in twelve months, and the number of people authorized to approve a contract stayed where it was.

A capability you can buy for thirty dollars a seat now feeds a bottleneck you cannot buy at any price.

The takeaway: production capacity became purchasable and review capacity did not, so the constraint moved onto your senior people. Find out which people, by name, before the next expansion of the program.

Engineering is the only function with instruments

Software teams are further along here because they measure. Twenty years of delivery telemetry means the jam shows up in a chart. Everywhere else, you find it by noticing that a reliable manager has started missing deadlines.

Google’s DORA program surveyed close to 5,000 technology professionals and published in September 2025. Adoption reached 90 percent of respondents, with a median of two hours a day spent working alongside AI. More than 80 percent said it improved their productivity, and 59 percent said it improved code quality. Throughput rose.

For the first time, DORA found a positive relationship between AI adoption and delivery throughput. In the same report: “AI adoption does continue to have a negative relationship with software delivery stability.” Google’s explanation is that acceleration exposes whatever was already weak further down the line, so a rise in change volume turns into instability. The report’s central conclusion is that AI amplifies whatever an organization already is.

Faros AI, an engineering analytics company, measured the jam directly. Its telemetry covered more than 10,000 developers across 1,255 teams, and teams with high AI adoption merged 98 percent more pull requests. Review time on those pull requests rose 91 percent. Bugs per developer rose 9 percent. At the company level, the report found no significant correlation between AI adoption and improvement, and its one-line explanation for where the gain went is the sentence I would put on the first slide: “downstream bottlenecks are absorbing the value created by AI tools.”

98%more pull requests merged by teams with high AI adoptionFaros AI, 10,000+ developers, 2025
91%longer review time on those same pull requestsFaros AI, 10,000+ developers, 2025

METR tested the people you would expect to be immune. In a randomized trial published in July 2025, 16 experienced open-source developers worked on 246 real issues in repositories they had contributed to for years. Going in, they expected AI to make them 24 percent faster. With the tools, they took 19 percent longer. Afterward, they still believed AI had sped them up by 20 percent. METR’s candidate explanations include mature codebases with “many implicit requirements” that the AI output does not meet, so a person has to supply them. That person is, by definition, the one who knows the codebase best.

Code quality research says the same thing from a different angle. GitClear analyzed 211 million lines of code and found that blocks with five or more duplicated lines grew eightfold during 2024, while moved lines, the fingerprint of someone consolidating and cleaning up, fell about 40 percent. 2024 was the first year on record when copied lines outnumbered moved ones.

8xgrowth in duplicated code blocks in a single year, while refactoring activity fell about 40 percentGitClear, 211 million lines analyzed, 2025

Put those together and you get more output, arriving faster and carrying more defects per unit, into a review step staffed the same as last year. Every function in your company is running that experiment. Engineering is the one that can prove it.

Company size changes the shape of the failure, and every size gets one. A $90M distributor has a controller, a general counsel who also runs HR, and one operations manager everyone trusts. When production capacity goes up tenfold, those three people are the aperture for the entire company, and the jam is loud. Approvals slow where everyone can see it, and everyone knows whose desk the work is sitting on.

Put the same tenfold increase inside a 4,000-person enterprise and it disappears into the system. There is a quality function, a model risk team, and a compliance group with a budget line. Two hundred reviewers can absorb an enormous rise in volume by each shaving a few seconds off their per-item time. Nobody is underwater and no alarm fires. The defect rate drifts for a year or more before it surfaces as a regulatory finding or a product recall. Bigger review capacity converts an acute problem into a chronic one. Chronic problems generate no pressure, so they go unfixed.

The takeaway: any company with a review step has this problem, and size decides whether it arrives as a visible queue or a silent drift. Work out which one you have before funding the next tool.

The best argument against all of this

The strongest objection is that the jam is temporary. Reviewers get AI too. A model can check another model’s work, catch the obvious failures, and hand the human a shorter queue of genuine exceptions. Some of that is working, and waving it away would be sloppy.

Two things limit how far it goes.

Accountability does not delegate. When a controller signs a reconciliation or a general counsel approves an indemnity clause, the signature is the product. A tool can narrow what they look at. It cannot be the party answerable for the result, and in regulated work it is not permitted to try.

Second-order review is still review. A screening model that flags 15 percent of outputs for human attention has shrunk the queue, and a tenfold rise in volume still puts more work in front of that human than existed before the program started. The percentage improved while the absolute number got worse. I have sat in meetings where a team celebrated the first and was buried by the second inside two quarters.

The pilot proved that one person’s attention works at twenty documents a week, and nothing beyond that.

Why the pilot worked

Here is the napkin math behind a failure pattern my clients describe to me constantly without having a name for it.

A pilot runs for eight weeks. One senior reviewer handles twenty outputs a week at fifteen minutes each, so five hours a week, absorbed alongside the day job. Accuracy looks excellent, the demo goes well, and leadership approves the scale-up.

Production moves the same workflow to 2,000 outputs a week. At fifteen minutes each that is 500 hours of senior review, something like twelve and a half full-time people who do not exist and were never in the budget. What happens next is never a decision anyone makes in a meeting. Review time per item compresses from fifteen minutes to ninety seconds, and approval becomes a formality performed by a tired person. The defect rate that looked excellent at pilot scale was a property of the attention, and the attention is gone.

PilotProduction
Outputs per week202,000
Minutes of review each1515
Senior hours required weekly5500
People that representspart of oneabout 12.5
What got testedthe reviewernothing yet

The scale-up decision treated review as a fixed cost. It is a variable cost, and it is denominated in the most expensive hours in the building.

The takeaway: a pilot validates the reviewer until review hours are measured at production volume. Rerun every approved business case with review time priced in.

Finding the jam in your own company

This is diagnosable in an afternoon, with no tool and no committee.

The Review Audit

  1. For each AI initiative, name the human who approves the output before it counts. A first name. If nobody can produce one, the work is going out unchecked and you have a different problem.

  2. Count how many initiatives name the same person. Three or more pointing at one controller, one general counsel, or one review team is a queue with a date on it.

  3. Ask that person how many minutes they spent per item before the program and how many they spend now. The gap between those two numbers is your real quality change.

  4. Multiply projected volume by current minutes per item. Convert to full-time people. Compare against the people you have.

  5. Check whether a single business case in the portfolio carries review hours as a cost. In the ones I have reviewed, close to none do.

Then make the constraint someone’s job. In the companies I work with, review capacity has no owner. That is how it became the binding constraint without ever appearing in a plan. Give it to whoever owns the workflow, along with three available moves.

Sequencing matters more than the choice itself. Improving the quality of what arrives costs the least and takes the longest. Buying more senior judgment works fastest and costs the most, and hiring or promoting runs two quarters minimum. The third move, throttling how much enters review at all, is the one I almost never see proposed, because it reads as retreating from AI. For the eighteen months it takes the other two to land, it is frequently the right call.

One more move, and it is the cheapest available. Stop measuring AI programs by output produced. Every dashboard I get handed counts drafts generated, tickets closed, documents summarized. None of them count items that cleared review on the first pass. That number tells you whether the program is creating value or creating homework for your controller, and any team can start counting it next week.

The takeaway: first-pass approval rate is the number that separates an AI workflow producing value from one producing homework. Start counting it next week and put it on the same page as the volume figures.

The funnel keeps filling either way. Production cost is heading toward zero for most of the work an organization does, and none of that capacity is the constraint anymore. The constraint is the people whose judgment the business runs on, however many of them you have. Go look at their calendars.


How RLK Can Help

The Operating Teardown maps where work queues in your company, including the review steps that never appear on an org chart. The AI Business Case engagement prices review hours into the numbers before you scale a pilot. That line is missing from most of the business cases I see. And the free AI Diagnostic takes about eight minutes and will tell you whether your program has a production problem or a review problem. Start the conversation.


Sources

Ryan King

About the author

Ryan King

Fifteen years in technology strategy at McKinsey and Deloitte. Now running RLK Consulting: enterprise-caliber tech strategy, one strategist doing every hour of the work. Over $10B in documented value capture across 50+ engagements and 12 industries.