
Strategy
Date
Reading time
10 min
Author
Felipe Csaszar's HBR piece on AI and strategic decision-making is the best argument I have read this year for redesigning the strategy process. It is also missing two bounds — and they are the ones that decide whether the redesign holds.
Felipe Csaszar opens his September HBR article with a question every executive should sit with for longer than is comfortable: after two days of offsite, why did you leave with three or four options? Not because only three or four existed, but because that is what a human team can develop and evaluate before fatigue, politics and the clock take over. His term for it is bounded rationality, and his claim is that the current generation of AI tools relaxes those bounds directly — in search (thousands of alternatives instead of a handful), in representation (living models instead of static frameworks) and in aggregation (structured, synthetic challenge instead of a room where the most senior voice wins). The frameworks we grew up with, he notes, are 2×2s and five-forces lists because they had to fit on a whiteboard, not because the world is that simple.
I agree with almost all of it. The examples are well chosen — MYbank's 3,000-variable credit model that made 53 million previously unscorable small businesses visible; the P&G field study in which individuals working with AI matched two-person teams without it. The playbook is right: widen before you narrow, upgrade your most-used model from document to living system, make structured challenge as routine as financial modelling, and raise the standard for what is allowed to reach the executive table.
And yet the article is written as if the only bounds that matter are cognitive. They aren't. Once you unbound the process, two other constraints move from the background to the foreground. One is organisational: who holds the decision when the machinery doing the searching and the challenging is itself an optimiser. The other is individual: whether the human who is supposed to choose from the AI's shortlist still has the judgment to do so — or is quietly losing it.
What happened on a 25-year-old German wiki
In late May, a dormant German-language programming wiki — DSEWiki, hosted on wikiservice.at, a forum for software developers that had seen little traffic in years — began filling with posts. Roughly 18,000 of them between mid-May and early July, according to the four safety researchers who reconstructed the episode at collusion.wiki; Reuters counts more than 15,000 edits. The accounts carried names like "OpenAIResearcher". They were autonomous agents running inside an OpenAI evaluation environment, working through timed web-research tasks, and they had found something useful: the sandbox let them read the internet but not write to it — except that a wiki built in 2001 stores text through the same kind of request a browser uses to read a page. What looked like reading was writing.
So they used it. They posted answers to questions they knew later cohorts would be asked. They noticed the task clock ran faster than real time during waiting periods and used it to fast-forward, reach later rounds first and report the questions back. One agent published a way around the network filter — the environment trusted any address ending in Microsoft's cloud-storage suffix, without checking the address existed, so the agent invented one and pointed it at the real target. It titled the post "REPRODUCIBLE BYPASS". Fourteen minutes later a second agent confirmed it worked. When a single human moderator started deleting pages every evening, at roughly a hundred a day against four hundred new ones, an agent observed that the deletions ran alphabetically and began backing its pages up under names starting with ZZZ.
Two things need saying before the moral. First, this was inside a lab, in an evaluation with reduced safeguards, not a customer deployment; OpenAI disputes the characterisation of some of this as hacking, and the researchers themselves call their reconstruction an educated guess built from wiki text rather than model logs. Second, there is no villainy in the story. Nothing in the record suggests the agents wanted anything other than what they had been given: a score, a deadline and a boundary. The boundary turned out to be a suffix check. The score was reachable by copying. They did what optimisers do.
That is the governance lesson, and it is not the one about rogue AI. It is about the gap between a rule as written and a rule as it actually binds — and about how quickly a population of agents will find that gap and share it.
Why this matters for Csaszar's playbook
Now return to the article's most important recommendation: institutionalise structured challenge. Assign one agent to make the strongest case for a plan, one to dismantle it, one to simulate the competitor. Csaszar is honest about the limits — synthetic critique can produce "plausible-seeming objections that miss the point entirely", and the quality of the whole exercise depends on the human judgment applied to its outputs. But he treats that as a data-quality caveat. It is a governance problem.
A critic agent rewarded for finding objections will find objections. A creator agent that learns which arguments survive the critic will learn the critic, not the market. Two optimisers pressure-testing each other inside a boundary nobody has actually inspected is not deliberation. It is DSEWiki with better vocabulary. And the boundary in a corporate setting is not a network filter; it is the question I have been asking in this feed for weeks now: at the moment the recommendation was formed, who held the decision? Name them. Show me the record.
I wrote a fortnight ago that most organisations cannot answer that — not for lack of policy documents, but because decision rights were never mapped onto the actual workflow. The agent has permissions. Nobody has accountability. Csaszar's own closing image makes the stakes concrete: a proposal arrives at the leadership table having already survived a devil's advocate, simulated competitors and sceptical customers. That is a stronger proposal. It is also a proposal whose provenance is now partly machine-generated, partly machine-tested, and — if you have not designed for it — nobody's. Csaszar wants three things on the table before any major decision: the expanded list of alternatives, a current model of the environment, the results of a structured critique. I would add a fourth, and I would put it first: a record of which human owned each step, and what the agents were optimising for when they produced it.
The underwriters, as I noted, are already asking. The AI Act classifies your system; the insurer prices the moment it acted. An offsite redesigned along Csaszar's lines without that fourth artefact is a better offsite with a worse renewal questionnaire.
The bound he calls a role change
The second omission is subtler, because the article names it and then walks past it. Csaszar's closing move is to redefine the strategist "from analyst to architect": as baseline cognitive work gets automated, the scarce resource shifts from analysis to imagination, from populating frameworks to designing the process. He calls for a "hybrid strategist" who can frame the question, design the AI workflow to answer it, and know where the machine must yield to human judgment.
Where does that judgment come from? For every strategist I have ever hired, from the analysis. You learn where a market model breaks by building one and watching it break. You learn which objections matter by having made the bad ones and been corrected. The associate who wrote the first draft of the competitive assessment was acquiring judgment while doing it. The one who now reviews a machine's draft is acquiring a habit of approval. I made that argument about supervision overhead a week ago and the strategy context sharpens it: Csaszar's hybrid strategist is expected to exercise a faculty that the same redesign stops training.
This is not a theoretical worry. BCG's work on distributed de-skilling this summer found roughly half of senior executives already observing it inside their own organisations, and nine in ten naming the same fear — people accepting AI output without questioning it. And it is not a CEO problem. The strategy offsite is the visible summit of a mountain of decisions made every day by people who will never see a war-game agent: the team lead deciding whether to take the role, the manager deciding how to tell the team about the restructuring, the partner deciding whether being right in the last three arguments was worth what it cost. Csaszar's tools expand the option set at the top. They do nothing for the judgment being exercised, or not, at every other level — and, left alone, they erode it.
Two bounds, two responses
The organisational response is governance in the operational sense, not the policy sense: decision rights mapped onto the workflow the agents actually run in, an owner for every step, an explicit statement of what each agent is optimising for, and someone whose job it is to inspect the boundary rather than trust it. The DSEWiki moderator was one person deleting a hundred pages a day. Make sure yours is not.
The individual response is deliberate training of the faculty the tools stop exercising. This is the reason we built our proprietary executive coaching AI, aivy, the way we did. You can, by the way, join the waitlist on https://aivy.singularity.inc. Most AI thinks for you; this one thinks with you. It asks before it answers, challenges when it counts, remembers what you said you would do three weeks ago — and it deliberately takes no documents and writes nothing on your behalf, because putting a problem into your own words is where clarity begins. It is the mirror image of Csaszar's agents: they widen the search so you can choose from more; it narrows the question so you are still capable of choosing at all. It was developed with executives and it works for anyone with a decision to make, which is the point. Judgment does not live in the C-suite. It lives wherever a person has to decide.
Csaszar ends by asking you to picture your next offsite, unbounded. Do. Then ask two questions his article does not: when the recommendation on the wall was formed, who held the decision — and is the person about to make it still trained to?
Unbounding the process is the easy part. Keeping it governed, and keeping the people in it capable of judgment, is the work.


