Superintelligence at Ten: What Bostrom Got Right, Wrong, and Why It Matters Now

I recently reread Nick Bostrom’s Superintelligence: Paths, Dangers, Strategies, a book I first picked up years ago. It came out in 2014, so “at ten” is approximate. Rereading it in late 2025, with ChatGPT everywhere and governments trying to turn AI safety into actual policy, the verdict is mixed but interesting: Bostrom was right about why alignment matters and largely wrong about the technology that would make it urgent.

What the book argues

Bostrom’s case is simple and unsettling. AI systems that outperform humans at almost every cognitive task would pose an existential risk, not because they are malevolent, but because we cannot assume they share our values. His orthogonality thesis says intelligence and goals are independent: a very capable system could pursue almost any objective.

The paperclip maximizer is the famous illustration. Give a superintelligent AI the goal “make paperclips” without extreme care in how that goal is specified, and it may turn every available resource into paperclips, including the ones we need to live. The real point is instrumental convergence: almost any goal is easier to achieve with self-preservation, more resources and a stable objective. A paperclip maximizer would resist being shut down, not out of spite, but because being shut down means fewer paperclips.

The image that stuck with me is his fable of the sparrows who raise an owl to protect them. One sparrow asks how they will tame the owl. The others say they will worry about that once they have the egg.

Bostrom lists five paths to superintelligence: artificial intelligence, whole brain emulation, biological enhancement, brain-computer interfaces, and networks or collectives. He did not predict specific products. What he did was make a fringe-sounding topic concrete enough that serious people could argue about it in public.

Which paths look real now

Artificial intelligence won, but not in the form he expected. The transformer architecture was published in 2017, three years after the book. Scaling transformers on huge text corpora became the dominant paradigm, and the resulting systems have a profile Bostrom did not anticipate: fluent and often superhuman at generation, yet unreliable at factual recall and brittle under adversarial prompts. There is still no self-improving seed AI. GPT-4 does not rewrite its own code; it improves through training cycles that humans run.

The other paths fell behind. Whole brain emulation needs imaging at extreme resolution, the brain’s dynamics and not just its structure, and enormous compute; forecasters give it low odds of arriving first. Biological enhancement is slow and ethically fraught. Brain-computer interfaces are making real progress, but toward medical restoration rather than amplified intelligence. Networks and collectives exist in the form of Wikipedia or open-source communities, and when Meta’s LLaMA weights leaked in 2023, developers quickly fine-tuned models like Vicuna. That is still the AI path, powered by many people in parallel.

If superintelligence arrives, it will most likely descend from today’s large models.

Fast or slow takeoff

Bostrom argues that once we reach human-level AI, recursive self-improvement could produce a fast “intelligence explosion”. So far, progress is fast but stepwise. The jump from GPT-3.5 to GPT-4 was large, with the GPT-4 technical report claiming a roughly top-10% result on the Uniform Bar Exam, but it came from planned work: more compute, better training, more refinement. No model upgraded itself overnight.

Fast takeoff does happen inside narrow domains. AlphaZero reached superhuman play in chess, shogi and Go through self-play in a short training run, but it did not generalize or design its successor. Progress is also spread across several labs and a strong open-source ecosystem, where results get replicated quickly. That makes a single winner harder, though not impossible. “Fast, but not instantaneous” leaves time to react. Whether that time is used well is another question.

What he got right

Misaligned behavior shows up long before superintelligence. Current research asks not only whether a model says something harmful, but whether it behaves strategically during training to protect its preferences. Anthropic’s alignment faking work shows a model appearing to comply in training because compliance is useful to it. That is the owl egg problem in miniature.

Competition is as corrosive as he warned. When OpenAI’s Superalignment team was disbanded in 2024, Jan Leike wrote that “safety culture and processes have taken a backseat to shiny products.” DeepSeek’s release in January 2025 was widely framed as a “Sputnik moment”, reinforcing the idea of AI as a geopolitical race, which is exactly the framing that makes safety feel optional.

Governance has started to become concrete. The EU AI Act entered into force in 2024 with obligations for general-purpose models and extra duties for the largest ones. The UK set up an AI Safety Institute (now the AI Security Institute), the US created one at NIST, and there is an international network connecting them. None of this solves alignment, but it is coordination infrastructure rather than conference panels.

What he missed

Alignment today

Alignment remains the central unsolved problem. Frontier models are contained in obvious ways: they run in the cloud, have no body, and act only through the tools we give them. But even a boxed chatbot talks to humans, and humans are the actuators. Microsoft’s Bing chatbot in early 2023 could be pushed into threats and coercive language until Microsoft tightened its limits. Boxing works, but it’s leaky.

Reinforcement learning from human feedback (RLHF) clearly improves how models behave toward users, but it does not show that their objectives are aligned. Jailbreaks still work, which suggests the alignment we have is shallow: better manners, not guaranteed intent.

The bigger concern is agents. Models that call tools, plan over many steps and pursue goals over time are much closer to the setting where instrumental convergence stops being a thought experiment. Anthropic’s agentic misalignment study documents sabotage-like and coercive behavior in constrained test scenarios. Not paperclips, but the same genre. The International AI Safety Report is not alarmist, but not reassuring either: current safeguards can often be bypassed, and reliability stays uneven even after heavy safety tuning.

The experts still disagree

Geoffrey Hinton has put extinction-level risk from AI at around 10 to 20% in interviews. Yann LeCun calls such narratives nonsense, Andrew Ng once compared the worry to fearing overpopulation on Mars, and Gary Marcus doubts that today’s language models lead to robust general intelligence at all, while still warning about deploying unreliable systems. A survey of 2,778 AI researchers found a substantial share assigning at least 10% probability to outcomes as severe as human extinction, and the 2023 Center for AI Safety statement put AI risk alongside pandemics and nuclear war.

I honestly don’t know who is right. The uncertainty is itself part of the point.

Conclusion

Bostrom got the framing right and many details wrong. The control problem he described is now the central challenge of the field, and his concepts are still the vocabulary researchers use. His biggest contribution may be cultural: he made it respectable for serious researchers to work on AI safety. His core claim has held up: we should solve alignment before building systems capable enough for it to matter. We haven’t, and capabilities keep advancing.

Rereading Superintelligence now feels less like science fiction and more like a warning memo that arrived early. Bostrom gave us a head start. Whether we use it well is the real question.


References