How saving humanity became a competitive disadvantage


· 14 min read
On Tuesday the 8th of September, a 27-year-old researcher named Jacob Coxon resigned from Anthropic and posted the reason on X. He had spent three years doing pre-training research at OpenAI and then at Anthropic, the two companies most likely to build the systems everyone is afraid of, and his verdict was that neither is acting responsibly. Both, he wrote, are racing straight toward self-improving superintelligence and "gambling with our lives." This was not a marketing stunt, he added. Executives soften their phrasing for the press and say something else in private. By Thursday the thread had passed 150 million views.
Then something stranger happened. Late on Tuesday, Evan Hubinger, an alignment science lead at Anthropic, quote-posted his departing colleague and agreed with him. Coxon was correct, Hubinger wrote: the people inside really do believe AI could kill all humans, and he personally put the odds above 10 percent within the decade. Anthropic was trying its best, he went on, but it had no plan for aligning a superintelligence and was not clearly on track to find one. Two days before that, OpenAI's chief scientist, Jakub Pachocki, had published a post saying no company had solved alignment and monitoring well enough to keep scaling at full speed for much longer, and that he expected, and hoped for, voluntary slowdowns until shared safety bars existed.
Nobody stopped anything. Both companies are moving toward what may be the largest public offerings in history. Both kept training. The senior people said it might kill everyone, the senior people kept building, and a 27-year-old walked out alone, which is the one move the system is built to make pointless. So the question is this. The people building the most powerful technology in history say, in public, under their own names, that it could kill everyone. They have the authority to slow their own companies down. They do not. The only move available to them, apparently, is to leave. Why can an industry that believes it might end the species not stop building?
The first theory. The obvious answer is that they do not believe it. The extinction talk is marketing. Nothing sells a product like the claim that it is too powerful to be safe, and a post on X costs an alignment lead nothing at all.
The dates kill this one. On 30 May 2023, Sam Altman of OpenAI, Dario Amodei of Anthropic and Demis Hassabis of Google DeepMind signed a statement that was one sentence long, twenty-two words, saying that reducing the risk of extinction from AI should be a global priority on the same footing as pandemics and nuclear war. OpenAI's charter, published in April 2018, years before there was anything much to sell, warns that late-stage development could turn into a competitive race without time for adequate safety precautions. Anthropic's own documents describe models capable of causing large-scale devastation, and Anthropic exists at all because a group of researchers walked out of OpenAI in 2021 over differences about how to build this safely. Coxon's whole point this week was that the fear is more candid in private than in the press. You do not write those sentences to move a chatbot. You write them because you mean them. Which makes the puzzle worse, not better. The more sincerely these people appear to believe in the danger, the faster they build the thing they are afraid of.
The second theory. So perhaps it is the people. Replace the wrong leaders with better ones and the race will slow. Except the replacement has already been tried, and it produced the same behaviour. The safety-minded faction left OpenAI and founded Anthropic, and Anthropic now trains frontier models. Coxon quit on Tuesday, his pre-training work will be done next week by someone else. Investors who pull out are replaced by investors who will not and governments that hesitate watch other governments that do not. When you change the players and the play stays exactly the same, the players were never the variable.
The third theory. Perhaps, then, safety research is the exit. Build the systems carefully, study them, slow down voluntarily when the studies say so, which is exactly what Pachocki said he hoped for on Sunday. Listen to how each lab actually argues. It must build more powerful AI to understand the risks of powerful AI. And humanity will be safer if it wins the race than if someone less careful does. Coxon described the version of this he heard inside Anthropic: the stakes were well understood, and the company was locked in a race to get there first, because no one else would act responsibly, so it had to. The attempt to control the danger becomes one more justification for creating it, and the door marked exit turns out to be painted on the wall.
Three theories dead, and the question is now much sharper than when we started. If sincere people, holding the right beliefs, inside institutions founded specifically to be careful, still accelerate, then the acceleration is not coming from inside the people.
Greg Elliott came to this sideways. By day he is a solutions consultant at Siemens specialising in data science and AI. On the side he wrote a book, The Psychopathic Selection Hypothesis, and the book led him to found an institute, the Institute for Symbolic Selection, to keep pulling on the thread. When Nate Hagens interviewed him for episode 231 of The Great Simplification, recorded on 13 August 2026, Elliott reached, seven minutes in, not for a formula but for an insight from the author John Steinbeck. The scene in The Grapes of Wrath where an Oklahoma tenant farmer is told that the bank taking his land is not a man. Men made it, the farmer is told, but men cannot control it.
That scene is the hypothesis in miniature. Elliott's claim is not that the people running our institutions are psychopaths. It is that large systems can select for and reproduce psychopathic behaviour even when nearly everyone inside them is decent. Human cooperation evolved in small groups under three conditions: repeated interaction, reputation, and consequences you could see. Cheat your neighbour and you eat alone. Under those conditions reciprocity wins, and the rare true psychopath, roughly one person in a hundred, is contained by ridicule, ostracism and, when it came to it, worse.
Scale inverts every one of those conditions. Anonymity replaces reputation. Feedback weakens until it barely registers. Consequences land on people you will never meet, or on people who have not been born. Elliott calls the result symbolic selection: once survival runs through markets, money, law and institutions rather than through the tribe, the strategies that win are the ones fitted to those gameboards, and the gameboards reward what he calls instrumental behaviour, treating others as means, over reciprocal behaviour, treating them as ends. He lingers, in the interview, on the open-outcry pits of the Chicago Mercantile Exchange, a room that rewards one cognitive profile and wears down the rest.
Now lay the AI race on that board. Markets reward companies for capturing investment, talent, users and infrastructure. They do not reward companies for protecting humanity from a catastrophe that is uncertain, hard to measure and spread across everyone alive. This produces what Elliott calls a tournament game. Every AI company might prefer a slower and safer transition. But each believes that if it slows while its competitors continue, someone else captures the market, builds the most powerful system and determines the future. Restraint becomes a competitive disadvantage. What starts as a choice hardens into an obligation imposed by the game.
And the game does not stop at the company. The company tells itself it cannot stop because a less responsible rival might win, investors cannot withdraw because they might miss the largest accumulation of wealth in history, governments cannot slow down because another country might gain strategic dominance, and employees cannot leave because someone else will do the work. Businesses adopt AI because their competitors are adopting it, and workers use it because their employers demand it. Nobody wants the final outcome, yet everyone is rewarded for taking the next step towards it.
The benefits are immediate and concentrated: profit, power, investment, geopolitical advantage. The possible costs are delayed and pushed onto humanity as a whole. That is precisely the environment in which, on Elliott's account, instrumental behaviour outcompetes reciprocal behaviour, and it produces what he calls an entrapment game: pro-social people locked into positions where leaving costs more than staying. Coxon told Time this week that the atmosphere inside the labs is one of near resignation, where people accept that the race is happening and put their heads down to make their own corner of it as safe as they can. That is what an entrapment game feels like from the inside. The technical term for the endpoint is a loss-dominant equilibrium. Stopping alone produces an immediate, predictable loss. Continuing contributes to a collective outcome that everyone involved may consider disastrous. Locally, acceleration is rational. Collectively, it may be suicidal.
The system does not need anyone to choose extinction. It only needs to make restraint more expensive than acceleration. That reminded me of a comment someone recently left on one of my LinkedIn posts: we are all on a plane flying straight towards a mountain, and since we can see the crash coming, surely we should turn. Someone else replied that the picture was wrong. There is no shared cockpit, no wheel we can all grab to turn the plane. What we are really doing is asking individual passengers to jump out and trust that something better will catch them on the way down.
That was meant as a rebuttal. But it was actually the answer to the question we started with. AI is not accelerating because nobody understands the danger. It is accelerating because stopping is punished. We have built a game in which choosing humanity could cost you your company, your investment, your job or your country's position, while gambling with humanity is rewarded as innovation and leadership. That is why ethical appeals to individual CEOs will never be enough. Whatever the people playing are like, the board is selecting their behaviour.
The formal name for the structure, when it plays out between rivals rather than inside one firm, is a multipolar trap. Every player would be better off if all of them stopped. No player can stop alone without being eaten by the ones who did not. Multipolar traps do not open from the inside. The only exit is a move that changes the payoff for every player at the same moment, which is a job no single player can do and which therefore has to be done by all of them together, or by someone standing outside the game with the power to rewrite its rules. Another way of putting it is that the only way out of the multipolar trap is collective action aimed at systemic change. Read Pachocki's post again with that in mind. He hopes for voluntary slowdowns, which is precisely the move the board punishes. A company that slows down alone loses its investors, its talent and its position, and the race carries on without it. Then, almost in passing, he names the actual fix: shared safety bars. Not a sacrifice one competitor makes while the others watch, but a floor all of them stand on at once.
In 2022 we created the Reduction Roadmap. A group of architects and engineers took the remaining global carbon budget for 1.5 degrees and divided it all the way down to a single square metre of Danish housing. It was an attempt to show the building industry, in plain numbers, how fast emissions have to fall for every project to stay within the 1.5 degree Paris accord budget. The result shocked the entire industry. New construction had to cut its emissions by 96 percent within seven to ten years, roughly ten percent of today's footprint removed every year for a decade. Meanwhile Denmark had just introduced its first legal carbon cap on new buildings, and it was set higher than what the average building already emitted. The law asked nothing of anyone.
So we decided to do something about it. We released the roadmap and asked people to take responsibility and follow it. Some did but then the messages started coming in. People contacted us to apologise, because they could no longer follow it. They were losing work to competitors who did not care. We had asked them to make the voluntary sacrifice the system punishes, the construction-industry version of Coxon's resignation.
That is exactly the multipolar trap. People want to act prosocially, and the gameboard turns it into a disadvantage. We racked our brains over what to do, and then it hit us: we did not need to change the individual player, we needed to change the gameboard. So instead of asking each firm to change alone, we launched a mobilisation campaign asking parliament for a carbon cap in line with the roadmap, less than half the legal limit and binding on everyone at once.
It looked like an impossible task. Why would the people who make their living pouring concrete and selling steel line up to make their own product harder to sell? But they did. Within three months the majority of workplaces in the Danish building industry had signed, and in the end more than 900 firms, municipalities, pension funds and research institutes put their names to it. Nrep, one of the largest property investors in the country, said in public that it wanted stricter rules than the ones it was legally bound by. In May 2024 a broad majority in parliament agreed to lower the cap and to keep stepping it down toward 2029. Not the full ask, but a sector responsible for around 30 percent of Denmark's climate footprint had demanded, in writing, with its competitors' signatures on the same page, that the state regulate it. It was the first time in Danish history an entire sector asked for binding climate regulation instead of fighting it.
What the campaign revealed was staggering. For some companies signing was of course a business advantage, since they were already ahead. What blew me away was what happened inside the big companies, the ones profiting from the status quo. Employees began demanding that their employers sign. In some firms more than a hundred employees threatened to quit if the company refused, and the company signed, because losing a hundred senior people overnight would bankrupt it within months. Most people want to go to work and make a positive difference. The roadmap offered them a meaningful way to regulate themselves, and they responded.
The reason it worked the second time is that we stopped asking anyone to be first. Signing cost a firm nothing. It was, if anything, an advantage: a signal to clients, to staff, to the municipalities awarding contracts. And once a few hundred had signed, not signing became the exposed position. The gameboard had been redrawn so that stopping together was safer than continuing alone.
This is how you get out of a multipolar trap, and it is the same move Elliott points to when he lists his leverage points at the end of the interview: shorten the feedback loops, raise transparency and accountability, rebuild the reciprocal conditions at whatever scale you can reach. A binding ceiling is a feedback loop with the force of law. A public signatory list is reputation, reinstalled. The sector did not turn virtuous overnight, the will was already there. What was missing was a mechanism in which virtue stopped being expensive.
This is the work we all have to start doing, turning minority insights into majority actions. The method does not appeal to the strongest player's conscience. It finds the minority who can already see the trap, gives them a demand that is collective, binding and precise, and makes joining that demand the rational move for the majority. The Roadmap was one instance. The pattern is general, and it is the only pattern that has ever opened a multipolar trap.
For AI, that means safety can no longer be a sacrifice, because as long as it is something one company gives up while its rivals keep running, no one will give it up for long. It has to become the price of entry: binding limits on capability, independent oversight, and coordination strong enough that stopping is safer than continuing. Nobody needs to appeal to Sam Altman's conscience when what he actually needs is a rule he can afford to follow.
The AI industry has signed its one sentence, and its alignment leads now say the odds out loud, but what it has not yet done is ask for the law that would let it mean a word of it. Jacob Coxon turned around alone on Tuesday, and the race did not slow by a single step, which is exactly what conscience is worth inside a trap: it costs the person who has it and changes nothing else.
We do not have a shortage of people who want to stop. We have a shortage of ways to stop without losing. The Danish building industry did not find its conscience in 2022. It found a way to use it.
Nobody can afford to be the first to turn around. So the turn has to be made together, or it will not be made at all.
This article is also published on Substack. illuminem Voices is a democratic space presenting the opinions of leading Sustainability Thought Leaders, their views do not necessarily represent those of illuminem.
The world needs sustainability knowledge. At illuminem, no interest group or shareholder can influence our work. Thank you for supporting our mission to make high-quality and independent sustainability information free for all. Every contribution helps. Thank you for donating today.
Kasper Benjamin Reimer Bjørkskov

AI · Ethical Governance
illuminem briefings

AI · Ethical Governance
illuminem briefings

AI · Public Governance
DeSmog

Oil & Gas · AI
The Guardian

AI · Public Governance
The Economist

AI · Social Responsibility