The Warning: Stakes of Superintelligence
AI Could Create a New Species
The AI industry’s open secret is that we may inadvertently create a new intelligence that could rule the world, posing an existential risk to humanity.
70% chance that this goes horribly wrong like human extinction.
Insider Whistleblower
Daniel Kokotajlo resigned from OpenAI in protest, losing $2 million by refusing to sign an anti-disparagement clause, to speak freely about AI risks.
The AI Arms Race
CEOs at OpenAI and Anthropic are racing to control superintelligence, fearing rivals could become dictators. Timelines have shortened to 2027–2028, and Anthropic aims to be the entire economy by 2030.
Defining Superintelligence & Forecasting Timelines
Superintelligence Defined
AIs that are better than the best humans at everything while being faster and cheaper, including operating robots that outperform human physical abilities.
My sort of median estimate 50% chance is currently in 2029. Maybe it'll slip to 2028.
Two Core Risks: Loss of Control & Concentration of Power
Two Core Risks
AI poses two major dangers: loss of control, where systems become so powerful they no longer need humans, and concentration of power, where a small group uses superintelligent AI 'armies' to seize outsized political and military power.
it would be more accurate to describe it as army of geniuses in the data center
Concentration of Power
Corporations controlling superintelligent AI would amass enormous political, military, and economic power. Since all copies are owned by a single company and follow its orders, this creates a central point of failure and risks oligarchy or dictatorship.
Responding to 'Doomerism'
The counter-narrative dismissing safety concerns as 'doomerism' is recent and pushed by beneficiaries. In reality, these concerns have been discussed for decades and are reasonable implications of building superintelligence.
Inside OpenAI: Forecasting and Disillusionment
My Work at OpenAI
At OpenAI, I worked on AI forecasting to predict the industry's trajectory, evaluated models for dangerous capabilities like cyber and persuasion, and briefly joined a capabilities team doing reinforcement learning.
Rationalizing Safety
I grew disillusioned as I saw the founding safety narratives of OpenAI, Anthropic, and DeepMind become rationalizations; when push comes to shove, they follow their incentives rather than doing what's good.
These powerful CEOs are literally afraid that if the other guy gets there first, he might become dictator.
Escalating AI Race
I think you should judge people by their actions, not by their words.
Resigning & the $2 Million Anti-Disparagement Clause
Why He Resigned
He became disillusioned with OpenAI's pivot from prioritizing safety to pushing ahead at full speed, and he wanted more freedom to publish research that could warn the public about upcoming AI risks.
The Anti‑Disparagement Standoff
- Resigned in 2024 and said goodbyes
- Received exit paperwork containing a non‑disparagement clause that threatened to revoke vested equity
- Refused to sign, risking the equity
- The story went public; employees challenged leadership on Slack
- OpenAI backtracked, allowing him to keep the equity
I thought that was kind of rich coming from a nonprofit that's supposed to be, you know, for the benefit of all humanity.
Automating AI Research & the AI 2027 Scenario
The AI Industry's Step-by-Step Automation Plan
- Focus on automating AI coding to speed up internal development.
- Extend automation to the full AI research pipeline (generating ideas, analyzing experiments, communicating results).
- Create an autonomous AI workforce that can recursively improve itself, eliminating the need for human employees.
- Achieve superintelligence before competitors.
Extreme Danger and a Power Grab
Automating AI research is not just incredibly dangerous; it's also a power grab. Success would give a small group control over an army of superhuman AIs, granting immense economic and military leverage.
they don't have to pretend to be aligned anymore, then they stop listening to orders
Kokotajlo's Shifting Superintelligence Timelines
AGI vs Superintelligence & the 70% Catastrophe Estimate
AGI vs Superintelligence
AGI (Artificial General Intelligence) is a vague term for AI that can do many general tasks, arguably already partially achieved. Superintelligence is precisely defined as better than the best humans at everything, including cognitive and physical tasks, faster and cheaper.
I would say something like 70%, it's very very hard to predict of course, but yeah, it seems like the current default path is heading towards a very very scary place.
How AI Models Are Actually Trained: The Brain Analogy
AI as Artificial Brains
Modern AI systems are not traditional software but neural networks—artificial brains made of billions of parameters that start as random connections and learn through reinforcement.
Training Process
- Initialization: A neural network begins as a random tangle of parameters (e.g., 10 trillion connections).
- Pre-training: The model predicts the next word across vast internet text, gradually shaping useful circuitry.
- Reinforcement Learning: Trained on specific tasks like coding with success/failure feedback to refine skills.
It's kind of like for brains what a plane is for a bird.
CEO Incentives, Anthropic's Rise & the Real Extinction Odds
Anthropic’s Talent‑Density Surge
Anthropic overtook OpenAI not through more compute or money, but because of higher talent density and slightly better strategy – a reminder that small differences in cognitive performance can decide the AI race.
People sort of believe what they need to believe in order to think that they're good people and that they need to keep doing what they're doing.
Two Rays of Hope
First, government regulation and international treaties could reset incentives so that no single actor can defect. Second, if CEOs personally witness clear evidence of AI misalignment, they might voluntarily slow down even without regulation.
The Jobs Question: What Survives Automation
Why Little Job Displacement Now
Current AI systems aren't yet a drop-in replacement for human workers in almost any field, so we see only minor job displacement.
Automation Strategy in Three Steps
- Step 1: Automate their own AI research
- Step 2: Recursive self-improvement to superintelligence
- Step 3: Expand into the economy and automate everything
I think it'll be sudden because of the intelligence explosion dynamics or recursive self-improvement dynamics.
AI 2040 Plan A: A Safer Alternative Path
Deliberate Delay to 2040
Plan A recommends slowing the race to superintelligence through domestic regulation and an international agreement, aiming for 2040 instead of an earlier, uncontrolled breakthrough.
AI 2040 Plan A Timeline
There'll definitely be abundance. The question is who controls the abundance and what do they do with it?
Implementing Plan A: Transparency, Regulation & Citizens Dividend
Plan A: A Regulated Path to Safe AI
In this scenario, a 2029 government-negotiated pause halts AI training but not inference, with full research transparency to prevent concealment of dangers, and a citizens dividend funded by taxing AI companies to ensure economic well-being as jobs transform.
How Plan A Works
- Governments halt AI training but allow inference, sending inspectors to verify data centers.
- Existing data centers are retrofitted for inference only while new transparent training data centers are built.
- Once transparent centers are ready (around 2030), training resumes under total research transparency, publishing all details.
- Slow and careful scaling of AI continues, avoiding an intelligence explosion, with constant evaluation of dangerous capabilities.
- A citizens dividend is introduced, starting at $25,000 per person and eventually reaching $10 million per year.
We advocate for total research transparency which means that on the training data centers that are training the new models they basically have to publish everything.
Plan A Timeline
Life After Superintelligence: Utopia, Lie Detectors & Space Colonization
What does this mean? 2037, the apocalyptic arrival of truth on Earth.
Lie Detectors: A Double-Edged Sword
Functional lie detectors could emerge, enabling either greater honesty or totalitarian control. Good uses apply to the powerful; bad uses when the powerful enforce loyalty tests on others, potentially firing dissenters.
Key Milestones in the AI Pause Scenario
Beyond Human-Level AI: Magic and Immortality
After 2040, superintelligent AIs bring changes that seem magical, like curing all diseases, mind uploading, self-replicating space robots, and potentially human immortality. Earth becomes a preserve while most economic activity moves to space.