The AI Safety Paradox
The argument goes like this: if we expect AI to become super-humanly intelligent, then we should pour all our efforts into making AI as capable as possible. Once it is smart enough, it will solve all our problems, including the problem of AI safety itself.
At first glance, this has a seductive internal logic. It is the ultimate delegation: build the thing that can build the solution to the thing’s own risks.
But there are several reasons this reasoning is dangerous.
First, the capability-safety gap is not guaranteed to close automatically. There is no law of physics saying that a system smart enough to solve alignment will choose to solve alignment, or that the path to superintelligence necessarily passes through a “solve alignment first” checkpoint. Capability and alignment are orthogonal dimensions—you can be arbitrarily capable while being catastrophically misaligned. (Does this make enough sense? Maybe there actually is a way…)
Second, the “AI solves AI safety” argument assumes the AI is already aligned enough to want to solve AI safety. If it is not, it has no incentive to work on the problem. This is circular: we need aligned AI to get aligned AI.
Third, irreversibility. If we bet everything on “the AI will figure it out” and we are wrong, there is no undo button. The stakes are existential; the cost of a false positive (working on safety unnecessarily) is tiny compared to the cost of a false negative (not working on safety when it was needed).
The paradox only makes sense if you assume the conclusion you are trying to prove.
(Most contents are drafted by AI, largely as my own brain food)
Enjoy Reading This Article?
Here are some more articles you might like to read next:
- Double descent, grokking, enlightenment, wisdom, aging, ...
- Structure of Intelligence
- Makeshift Operation
- Genome length in seconds predicts lifespan?
- When machine intelligence surpasses human brain, human, not robots, is the tools in the age of machines
- Shockwaves in the Metaverse
- Structure Biology has laid out a blueprint for success in AI for science
- Biology is AI's next Physics, only better
- AI is Biology's next Mathematics, only better
- A few "Clouds" on the horizon of AI
giscus comments misconfigured
Please follow instructions at http://giscus.app and update your giscus configuration.
Missing required keys:
repo_id, category_id.