A surface of thousands of individual elements

Do you stop addictive behaviour when it has a one-in-ten chance of killing you?

Do you stop addictive behaviour when it has a one-in-ten chance of killing you?

Do you stop addictive behaviour when it has a one-in-ten chance of killing you?

Frontier labs put AI's odds of ending humanity at one in ten. History says we won't stop. On what executives can still govern — and why it starts with a name.

Frontier labs put AI's odds of ending humanity at one in ten. History says we won't stop. On what executives can still govern — and why it starts with a name.

Frontier labs put AI's odds of ending humanity at one in ten. History says we won't stop. On what executives can still govern — and why it starts with a name.

Frontier labs put AI's odds of ending humanity at one in ten. History says we won't stop. On what executives can still govern — and why it starts with a name.

Category

Strategy

Date

Reading time

10 min

Author

Dr. Christian Oehner

Dr. Christian Oehner

When I was younger, my lifestyle was, in many ways, reckless. A lot of partying, a lot of drinking, a motorcycle. I loved the thrill of it. It was fun, and I did not think much about the downside.

When my first child was born, I stopped riding the motorcycle. Most of the other behaviour did not change. Eventually I changed it all, because the ones I love motivated me to. But it was a long journey, and it was anything but obvious that I would succeed.

The situation we face with AI is, in my view, similar in many ways.

There is a huge upside. The business case is real. I am not talking about the short-term view, and not about the naysayers who say it does not land in business, the synergies are not there, the economics do not work. All of that may be true today. It will very soon be a thing of the past. AI makes sense, and it opens possibilities that a couple of years ago we could only have dreamed of — for individuals and large organisations alike.

But there is also a downside. And I do not mean mass unemployment, or the loss of meaning for people whose work has essentially gone away. I mean the extinction scenario — what the public calls the Terminator scenario. The one Roman Yampolskiy has been warning about for years, and the one that, as of this week, researchers inside the major labs are warning about in public. Yampolskiy puts the probability that AI wipes out humanity within a century at 99.9 percent; he is an outlier. Geoffrey Hinton says ten to twenty percent. Dario Amodei has put the chance of things going "really, really badly" at twenty-five. This week an Anthropic researcher resigned, writing that the labs are racing to self-improving superintelligence and gambling with our lives — and Anthropic's alignment science lead endorsed the post, adding that his own estimate of AI killing all humans within a decade is above ten percent.

Not by intent, most likely. More likely as a by-product of misalignment.

And the most recent research from Anthropic itself — and from my own partner Vinicius Ferraz — points to us not being anywhere near understanding how these systems arrive at a decision. Anthropic's interpretability work is the most advanced there is, and by its own account it captures a fraction of the computation, needs hours of human effort for a prompt of a few dozen words, and cannot verify what a model reports about its own reasoning. An August paper from Anthropic's alignment team concedes that no interpretability tool currently predicts a model's behaviour better than simply reading the transcript. Vinicius's preprint, out this week, goes one level deeper. His team recorded the internal activations of four open models while they played the complete set of 144 two-by-two strategic games, and followed the payoff incentive from the prompt through the network to the choice. The incentive was readable inside every model. Whether it actually reached the decision differed from model to model — and two versions of the same model, sharing the same pretrained weights and agreeing on 96 percent of their choices, got there by different internal routes. "Choosing like a strategic agent," the abstract says, "does not mean computing like one." The output, in his closing line, is not the explanation; it is the thing that still needs explaining. We can watch these systems decide and still not know how — and two that decide alike need not decide alike.

This is not reassuring.

So the question is: will we stop walking down a path that has potentially huge rewards and potentially catastrophic failure at its end? And will we stop soon?

The empirical evidence on human behaviour does not point that way.

Smokers know that smoking may kill them. Statistically it kills about one in two of those who keep at it, and costs them ten years. Two thirds say they want to quit. Fewer than one in ten manage it in any given year.

The creators of the nuclear bomb knew that what they were building could end the world rather than dominate it. They were vocal about it — Szilard's petition, the Franck Report, Oppenheimer telling Truman he had blood on his hands, Einstein calling the letter to Roosevelt the one great mistake of his life. They built it anyway. The clock they founded stands at 85 seconds to midnight, the closest it has ever been.

A BASE jumper knows that it is the deadliest sport there is: about one death per 2,300 jumps, roughly one in sixty active jumpers dying each year. They jump.

Game theory and economics say the same. The race to AI is a prisoner's dilemma, and it has been modelled as one for a decade: the more competitors and the better they can see each other, the stronger the incentive to cut corners on safety. And the recent history of democracies and totalitarian regimes, as Yuval Noah Harari lays it out in Nexus, does not make me confident either. His argument is that the difference between the two is the presence of self-correcting mechanisms — and that AI is the first technology that makes decisions and generates ideas on its own. The people building it tell him they would like to slow down but cannot trust their competitors to do the same. They are, however, prepared to trust the AI.

There are many reasons to think humanity will keep walking straight down this path, not knowing whether it ends in Eden or Armageddon.

So if we assume all of that — what is the sensible thing to do, for any human and any organisation?

I did not stop riding the motorcycle because I had read the accident statistics. Per kilometre, a motorcycle in Austria is about twenty times as likely to kill you as a car; I knew that before and after. I stopped because there was suddenly a person for whom my decisions had consequences, and because the ride had lost its appeal. The numbers had not changed. What changed was who I was answerable to.

That is the useful part of the analogy.

Neither an individual nor a company will settle the alignment problem. That sits with a handful of labs and, if at all, with governments. What sits with us is governance — and governance is not a policy document. It is the answer to a small number of uncomfortable questions.

Make the most of what AI offers. The upside is real, and not using it carries its own risk. But understand it as well as you can — which today means understanding precisely how little anyone understands.

Be deliberate about what AI is allowed to do, and where the final decision rests with a human. Not in principle: per process, per decision, per amount.

Then walk the talk. A decision right on paper that nobody exercises has already migrated to the machine.

And know the name of the human with whom the decision rests. If you cannot name that person, the decision is not governed. It is delegated — to something that cannot tell you how it decided, and that you cannot tell how it will.

Transform how your organization operates with AI-Intelligence

Get the Singularity newsletter

Product updates, how-tos, community spotlights, and more. Delivered directly to your inbox.

Please provide your email address if you'd like to receive our monthly developer newsletter. You can unsubscribe at any time.

Transform how your organization operates with AI-Intelligence

Get the Singularity newsletter

Product updates, how-tos, community spotlights, and more. Delivered directly to your inbox.

Please provide your email address if you'd like to receive our monthly developer newsletter. You can unsubscribe at any time.

Transform how your organization operates with AI-Intelligence

Get the Singularity newsletter

Product updates, how-tos, community spotlights, and more. Delivered directly to your inbox.

Please provide your email address if you'd like to receive our monthly developer newsletter. You can unsubscribe at any time.