Sipahi Demir
Enterprise AI6 min read

The Missing Layer Is Inside the Enterprise

Everyone wants agents. Then you tell one to handle a real process and the question appears. How? The answer lives in people, in judgment, in exceptions, and it was never written down anywhere.

AI made software cheaper to build and capable of doing the work itself. Put those two unlocks together and you get something larger than better applications: an operating system for the business. That is exactly where the problem starts.

Everyone wants agents. Agents that sell. Agents that procure. Agents that handle support. Agents that analyze. Agents that operate back office processes. Agents that coordinate other agents.

And technically, we are getting there. But watch what happens when agents meet real work. Carnegie Mellon built a simulated software company and ran agents against 175 realistic professional tasks. The best model completed 30.3% of them autonomously. The instructive part is not the number. It is the failure modes. The agents did not mainly fail at reasoning. They failed at the workplace. They lost track of multi-step communication. One was defeated by an ordinary closable popup.

Tell an agent to handle a process

And immediately the real question appears: How?

Which decision should it make? What information matters? Which source should it trust? What is an exception? When should it stop? Who should approve it? What would an experienced employee do? What would make that employee reject the obvious answer? What signals does that person notice that never appear in the database?

These questions expose something that enterprise software has never completely solved. Companies have documented their data. They have documented their systems. They have documented policies. They have documented parts of their workflows. But they have never fully documented how the organization actually thinks and operates.

The most valuable operational knowledge often lives somewhere else. It lives in people. In judgment. In experience. In exceptions. In habits. In decisions that never become database entries. In the reason someone says: "No. Don't do it that way." And when you ask why, the answer is often: "Because I've seen this before."

That is operational intelligence. And it is precisely the layer that conventional software struggles to capture.

An old bottleneck with a name

This problem is older than it looks. Michael Polanyi named it in 1966: "we can know more than we can tell." Edward Feigenbaum met the same wall while building the first expert systems, and by 1982 he was calling knowledge acquisition the critical bottleneck of artificial intelligence. The expert-systems industry of the 1980s died largely of that bottleneck. The knowledge could not be economically extracted, encoded and maintained.

David Autor points out that there have only ever been two ways around it. Simplify the work until tacit skill is not needed. Or let a machine learn from a large record of correct performance. An organization that never documented how it decides has neither option. There are no explicit rules to program, and there is no record to learn from.

Where the measured gains actually come from

This explains something we are already seeing in AI deployment. AI can produce extraordinary results when the work is clearly specified. But when work depends on unwritten context, performance becomes much harder.

One large study of 5,179 customer-support agents found an average productivity improvement of about 14%, with much larger gains among less experienced workers: 34% for the bottom fifth by skill, and no measurable gain for agents past their first year. The researchers' evidence pointed at the mechanism. The AI was diffusing the practices of the strongest agents through everyone else. And the study carries a condition that is easy to miss. It worked because that contact center recorded every conversation and its outcome, so the model had the firm's own expertise to learn from. Where the work leaves no record, there is nothing to diffuse.

That finding is profound. Because it tells us something about where AI's value comes from. The model creates leverage. But the organization still has to provide the intelligence that gets leveraged.

And another body of evidence makes the problem even clearer. The measured effect of AI changes dramatically depending on how describable the work is. In a field experiment with 758 BCG consultants, tasks inside the model's capability envelope were finished 25% faster at markedly higher rated quality. On a task chosen from just outside that envelope, correctness fell from 84% to the low sixties. The consultants who had been trained in prompting did worse, not better, and the AI made the wrong recommendations more persuasive. In another study, experienced developers predicted they were 24% faster with AI, still believed they were 20% faster after the study, and were measured 19% slower.

So controlled experiments on well-specified tasks can show large productivity improvements, while the same randomized methods applied to real work in real firms show much smaller effects. When Danish researchers followed roughly 25,000 workers in the occupations most exposed to AI, the average time saved was 2.8% of work hours, and the effect on earnings and hours was indistinguishable from zero. The difference is not the intelligence of the model. It is the complementary work required to redesign the operation around it.

So the bottleneck is not simply: "Which model are you using?" The deeper question is: "How well does the organization understand and specify the work?" That is why better models alone do not automatically produce better enterprise outcomes.

The organisations already know this

They know it in their own way. Across seven years of the same survey, the share of senior Fortune 1000 data and AI executives naming culture, people and process as the principal barrier has never fallen below 78%, and in most years it has been above 90%. Not technology. The technology underneath that answer was replaced twice in those years. The answer did not move.

And the employees have not been waiting. In a study of 32,000 employees across 47 countries, 70% of those using AI at work reach it through free public tools, only 42% through anything their employer provides, and 57% conceal their use. The adoption already happened. The organization simply cannot see it.

The missing profession

Not another person who knows how to prompt a model. Not another consultant who produces an AI strategy deck. Someone who can go into an enterprise and answer: How does this company actually work? Then: Why does it work this way? Then: Which parts should remain human? Which parts should become software? Which parts should become agents? Which parts should be redesigned entirely?

That is a fundamentally different discipline. And it leads to the third great transition.

Sources

  • TheAgentCompany. Xu et al., Carnegie Mellon, NeurIPS 2025. Best agent: 30.3% of 175 realistic tasks completed autonomously. The failures are workplace-competence failures, not reasoning failures.
  • Knowledge Engineering for the 1980s. Edward Feigenbaum, Stanford, 1982. Knowledge acquisition named the critical bottleneck of AI. The concept dates to his 1977 IJCAI paper.
  • Polanyi's Paradox and the Shape of Employment Growth. David Autor, NBER, 2014. The two engineering routes around tacit knowledge: simplify the environment, or learn from recorded ground truth.
  • Generative AI at Work. Brynjolfsson, Li & Raymond, NBER 2023 / QJE 2025. 5,179 support agents, about 14% average productivity gain in the preferred specification (the published journal abstract reports 15%), concentrated almost entirely in novices. The model diffused top performers' practices.
  • Navigating the Jagged Technological Frontier. Dell'Acqua et al., Harvard/BCG, 2023. 758 consultants: large gains inside the capability envelope, a 19-point correctness drop outside it, and prompt training made the outside-envelope failure worse.
  • Early-2025 AI on Experienced Developer Productivity. METR, 2025. Predicted 24% faster, believed 20% faster afterwards, measured 19% slower. n=16 developers, mature codebases. An existence proof, not a law.
  • Large Language Models, Small Labor Market Effects. Humlum & Vestergaard, Chicago BFI/NBER, 2025. ~25,000 Danish workers linked to payroll: 2.8% self-reported time savings, precisely-estimated zero effects on earnings and hours.
  • 2025 AI & Data Leadership Executive Benchmark Survey. Bean & Davenport, formerly NewVantage/Wavestone. Culture, people and process named the principal barrier by 91.2% in 2025. Never below 77.6% across seven measured years.
  • Trust, Attitudes and Use of AI. University of Melbourne & KPMG, 2025, n=32,352 employees. 70% of AI-using employees rely on free public tools. 57% conceal their use.
Comment
ShareXLinkedIn

Get the next essay by email.

Roughly one every two weeks. No other email, ever. Unsubscribe in one click.

Related

Comments

I would love to hear your thoughts.

Signed with your name, and shown here once I have read it. Your email address is never published — it is only so I can reply.

Sipahi Demir

Written by Sipahi Demir. contact@sipahidemir.com