Shadow AI Is a Symptom, Not a Threat: The Operating-Model Gap Behind the 95%
The first thing I do when a new client organisation brings us in is ask a quiet question. Not “what is your AI roadmap,” which produces a slide, but “what are your people already using that nobody signed off on.” The answers arrive with a small, guilty smile. A finance lead pasting board material into a chatbot she pays for herself. A team that quietly runs its customer replies through a model the IT department has never heard of. In the roughly thirty organisations we have surveyed at Astu Labs between April 2025 and June 2026 — several hundred respondents in all — every single one had shadow AI running inside it. Not most. Every one.
That is worth sitting with for a moment. We did not go looking for the enthusiasts. This was ordinary onboarding across ordinary Nordic organisations, and the tools were already everywhere before any roadmap arrived.
The pilot didn’t fail on the technology
Here is the claim I want to defend. The reason most enterprise AI never leaves the pilot is not that the technology is immature, the data is messy, or the talent is missing. It is that the operating model never changed. And shadow AI is not the security problem it is usually filed under. It is the clearest demand signal you will ever get — free, unsolicited, and almost universally ignored.
I have watched this pattern for three decades from both sides. I started as a coder; I have spent the years since leading people and cultures through technology change. The failure is rarely where the post-mortem points. It points at the model, the vendor, the integration. It should be pointing at how the work is organised.
What the evidence actually says about shadow AI
Start with the number everyone quotes. MIT’s Project NANDA, in The GenAI Divide: State of AI in Business 2025, found that around 95% of generative-AI pilots produce no measurable business return — only about 5% extract real value. The path narrows brutally: 60% of organisations evaluate enterprise-grade systems, 20% reach a pilot, 5% reach production. The headline is easy to weaponise as AI-scepticism. But read the researchers’ own conclusion, because it is the part that matters: the core barrier is learning, not infrastructure, regulation, or talent. The systems, and the organisations around them, do not retain feedback, adapt to context, or improve over time.
Now hold that against the adoption data from the same report: only about 40% of companies have official large-language-model subscriptions, while in over 90% of companies employees use AI tools on their own. Our own client data puts a floor under how far this goes at the individual level. When we asked people directly — “do you use a generative-AI tool you were not given by your employer, and that your employer does not know about” — 24% said yes. That is a deliberately self-incriminating question. The true figure is almost certainly higher than a quarter; 24% is the number people will admit to.
So the picture is not an adoption problem. Adoption already happened, from the bottom up, in the shadows. The pilots are failing on top of an organisation where the appetite is proven and the operating model has not moved an inch to meet it.
Why does the appetite not convert into results? Because general-purpose technologies do not pay off on installation. Brynjolfsson, Rock and Syverson call this the productivity J-curve: technologies like AI require complementary, mostly intangible investments — “business process redesign, co-invention of new products and business models” — before the gains appear. Productivity dips or flattens first, and rises only once those intangibles are harvested. This is not a new or AI-specific finding. Ichniowski, Shaw and Prennushi showed it across 36 steel production lines: it was systems of complementary work practices, not any single technology, that lifted productivity. Bresnahan, Brynjolfsson and Hitt showed the same at firm level — IT delivered its largest effect only when paired with organisational redesign.
And if you want the AI-era proof point, it exists. In a randomised controlled trial across 66 companies and over 7,000 knowledge workers, giving people Microsoft 365 Copilot saved them around two hours of email a week — a real, individual time saving — but the structure of their work did not change. The number and composition of their tasks stayed the same. The authors are explicit: deeper change would require redesigning processes and reallocating responsibilities. The tool alone moves the individual. It does not move the organisation.
“But the models keep getting better”
The honest objection is that this is a timing argument, not a structural one. Wait two model generations, the reasoning goes, and the capability gap closes on its own — no reorganisation required.
There are two mistakes in that. The first is conceptual: the binding constraint the MIT data identifies is learning, and learning is an organisational property, not a model property. A more capable model dropped into a workflow that cannot retain feedback or reassign work produces a faster version of the same stalled pilot. The second is strategic. If advantage came from the model, it would be available to every competitor on the same subscription the following quarter. What does not copy is the operating model you build around it — how decisions, feedback and responsibility are wired. That is precisely why it is worth the hard work, and precisely why waiting is the expensive option.
What to do instead — start here
- Treat shadow AI as market research, not misconduct. Before you write a policy that drives it further underground, map it. Which tasks are people quietly automating, and why those? That map is the highest-signal, lowest-cost prioritisation input you will get. Amnesty first, governance second.
- Fund the redesign, not just the licence. A Copilot seat is the cheap part. Budget explicitly for the process redesign and role changes the J-curve says are the actual source of return — and expect the curve to dip before it lifts. A pilot with no redesign line in its budget is a pilot you have already scoped to fail.
- Measure learning, not logins. Seat counts and prompt volumes flatter you. Track whether the organisation is getting better at the work over time: cycle times on redesigned processes, feedback that actually re-enters the system, capability that compounds. If nothing improves month over month, you have adoption without transformation — the exact definition of the divide.
- Put one redesigned workflow into production before you run the next ten pilots. The MIT path loses almost everyone between pilot and production. One workflow taken all the way through — process redesigned, responsibilities reallocated, learning loop closed — teaches you more than a portfolio of experiments that never cross the line.
The gap is human, and that is the good news
The GenAI Divide is not, at root, a divide between companies with better technology and companies with worse. It is a divide between organisations willing to change how they work and organisations waiting for a tool to do it for them. The 95% are not unlucky. They bought the licence and skipped the redesign.
Shadow AI has already told you your people are ready. The question the pilot was really testing was never about the model. It was whether the organisation around it could learn. Technology keeps you in the game. People win it.
Lenni Laukkanen is the founder of Astu Labs and works with leadership teams on turning AI from experimentation into an operating model. For speaking and advisory enquiries: lenni@lennilaukkanen.fi — I reply within one business day.
Sources
- MIT Project NANDA. The GenAI Divide: State of AI in Business 2025. ~95% of GenAI pilots show no measurable P&L return; 60% → 20% → 5% evaluation-to-production path; ~40% official LLM subscriptions vs. 90%+ employee self-use; core barrier identified as learning.
- Astu Labs client onboarding surveys, April 2025 – June 2026 (~30 Nordic organisations, several hundred respondents; aggregate, no client identifiers). Shadow AI present in every organisation surveyed; 24% self-report using an unapproved, unknown-to-employer GenAI tool.
- Brynjolfsson, Erik; Rock, Daniel; Syverson, Chad. “The Productivity J-Curve.” American Economic Journal: Macroeconomics 13(1), 2021.
- Ichniowski, Casey; Shaw, Kathryn; Prennushi, Giovanna. “The Effects of Human Resource Management Practices on Productivity.” American Economic Review 87(3), 1997.
- Bresnahan, Timothy; Brynjolfsson, Erik; Hitt, Lorin. “Information Technology, Workplace Organization, and the Demand for Skilled Labor.” QJE 117(1), 2002.
- Dillon, Eleanor W.; Jaffe, Sonia; Immorlica, Nicole; Stanton, Christopher T. “Shifting Work Patterns with Generative AI.” NBER Working Paper 33795, 2025 (RCT, 66 firms, n=7,137; working paper).
