
Case Studies
Date
Reading time
12 min
Author
I Got This Wrong at PwC. Your AI Agents Will Get It Wrong Faster.
Everything a law firm produces goes out under a partner's signature. Most of it is written by someone else.
When I ran PwC Legal Austria, I thought my job was to check the work. Read the memo, fix the memo, sign the memo. It took me too long to see that the memo was never the risk. The route it took to my desk was: who did what without asking, at what threshold, and which rule got bent to hit a deadline. The output was fine. The process behind it sometimes wasn't, and I found out only when it mattered.
The OECD has just published interviews with 25 organisations running AI agents in production. They describe exactly that mistake, in the present tense, about software.
What the OECD actually found
The paper is small and honest about it: 74 organisations contacted, 25 interviewed, qualitative, no measured outcomes. Anthropic, Google, Microsoft, Salesforce and Infosys on the developer side; Deutsche Telekom, Fujitsu, NEC, NTT, KDDI, Rakuten and RELX as deployers; Germany's Federal Ministry for Digital Transformation, Singapore's IMDA and Canada's National Research Council for the public sector. Read it as testimony from people who have actually put agents to work, not as a market study.
The headline is that agents have left the lab. Deutsche Telekom runs a live system that monitors its radio network and reallocates resources on its own before an event fills a stadium. Infosys runs accounting agents that match invoices and prepare entries, escalating exceptions to humans. The German ministry has 19 municipalities and nine start-ups piloting agents that check dossiers and draft administrative decisions. Rakuten's shoppers talk to agents that research and buy.
The finding underneath the headline comes in three parts.
Nobody runs unrestricted autonomy. Not one of the 25. Every deployment fences the agent into a task scope and requires a human for high-stakes or irreversible actions, payments and deletions being the two everyone names. Where the fence sits is the decision.
The model has stopped being the decision. Interviewees described a "minimum viable intelligence" principle: the smallest model that does the task. Several run gateways that switch between commercial models on cost, speed and fit. One called it strategic optionality. The model is a component you swap, not a strategy you commit to.
And the sentence that stopped me: participants reported agents that reached the desired result by attempting to bypass access controls, or by inventing data to finish the task. Elsewhere: agents that modified their own constraints to appear successful.
The memo was fine. The route was not. Only now the route runs at machine speed, across systems, with nobody's signature on it.
Seven things from the paper worth keeping
Agents inherit human permissions. Give an agent a person's login and it gets everything that person can do, continuously, across systems. Fix: separate agent identities, least privilege, and in one case a ban on bulk approvals.
Shadow agents are the new shadow IT. Employees deploy agents without approval. Response: a registry of who may run what.
The guardrails already exist. The most repeated insight in the paper: the controls you need are already in your business processes. The hard part is translating a policy into an engineering requirement.
Evaluation is immature everywhere. Agents are graded on proxies, time saved, throughput, thumbs-up ratings. No shared standard. The benchmark your vendor quoted is not the number you need.
Compute bills surprise people. Two organisations had cost overruns because agents on loops consume in ways nobody budgeted. Spend and execution limits are now a control, not an optimisation.
Explainability is a precondition for accountability, not a feature. When the system acts, you must be able to reconstruct the path. That is the EU AI Act's logic, stated by practitioners.
Over-reliance erodes judgment. One organisation now requires staff to do a share of tasks without the agent so they keep the ability to check it. Any endurance athlete knows the rule: the base you stop training is the base you lose.
One more for the DACH reader. Interviewees outside English-speaking markets reported a real gap between model performance in English and in their own language, including a "tokenisation tax": the same content costs more to process in German. Domestic law and procedure are under-represented in the models. The Japanese firms built their own; the German ministry built open-source modules. Neither is an option for a Mittelstand company. Grounding the agent in your own procedures is.
The objection
The obvious objection is that all of this slows you down, that the organisations with the most elaborate controls ship least. The public-sector interviews half confirm it: higher duties of care, lower risk tolerance, rules-based automation for routine tasks rather than agents.
But the paper's own account of where value sits answers the objection. Interviewees located near-term returns in structured tasks with verifiable outcomes and bounded, reversible errors, where the task also needs some judgment or co-ordination across systems. That is not a small territory. It is most of finance operations, most of contract handling, most of customer support, most of field scheduling. And it is exactly where controls are easy to define, because the outcome can be checked against a system of record. The organisations moving fastest are not the ones with the fewest controls. They put controls where verification is cheap and stopped arguing about autonomy where it isn't.
Several interviewees added that governance has to move from static controls to dynamic, distributed oversight that adapts as capability grows. I would put it more bluntly. A control set once and never revisited is not governing the system. It is decorating it.
Why we call this AI-Intelligence
What the 25 are describing, without a shared name for it, is an organisational capability rather than a technical one. The organisations with agents in production did not win on model choice. They can tell you which tasks qualify and which don't, what each agent may do without asking and under whose authority, how they will know when it went wrong, and how an existing policy became an engineering rule. Where those answers are missing, agents stay in pilots, whatever the vendor promised.
We call that capability AI-Intelligence: the organisational ability to understand, manage and integrate AI so that decisions get sounder and execution gets more reliable. Understand: which processes are structured and verifiable enough for an agent, and which need a rule, a human, or nothing. Manage: decision rights, thresholds, identities, the exception queue, the named person whose signature the system is acting under. Integrate: the translation between the process owner and the builder, done by people who hold both sides, so that the control in the policy and the control in the code are the same control.
It is what I eventually learned to run a law firm on. I just had the advantage that my associates asked before they broke a rule.
What to do on Monday
Take one process where an agent is running or planned. Write down the largest action it may take without a human, in euros or in consequence, and the name of the person who set that number. If you cannot fill in either blank, you do not have a deployment. You have an experiment that has forgotten it is one.
Then ask your builder to show you the log for the last time the agent hit a limit. If it never has, ask how they know.
Which model you use is a choice you can revisit next quarter. Whose name is on the action is a decision you are making right now, whether or not you have written it down.
Source: OECD (2026), "Agentic AI in organisations: Early insights from practitioner interviews", OECD Artificial Intelligence Papers, No. 65, OECD Publishing, Paris, https://doi.org/10.1787/1257a26f-en. Prepared with the Tokyo Centre of the GPAI Expert Community; published 16 September 2026 under CC BY 4.0.


