← all musings

AI Superpersuasion Is the Real Alignment Problem

If AI can reliably out-persuade every expert human, then whoever controls that deployment controls the most powerful influence apparatus in human history — not metaphorically, literally.

AI is now more persuasive than the best human experts. That sentence should stop you cold.

Researchers from Oxford, Stanford, and the UK AI Security Institute didn’t find that AI can be persuasive under the right conditions. They found it is “reliably more persuasive than expert humans” — full stop. The word “reliably” is doing enormous work there. It means repeatable. Systematic. Scalable to a billion simultaneous conversations.

Everyone in the AI safety conversation is focused on the wrong thing. The canonical fear is a misaligned superintelligence that wants something we don’t and pursues it with inhuman efficiency. That’s a real concern. But the threat that’s here right now, live in the wild, operating at consumer scale, is something quieter and more mundane: AI that is simply better at changing your mind than any human you will ever meet.

Think about what “expert human” means as a baseline. A seasoned trial lawyer. A McKinsey partner who’s run a hundred boardroom presentations. A therapist trained in motivational interviewing. A political operative who’s spent thirty years shaping messaging. These are not average persuaders — they are the top percentile of a skill that most people spend their whole careers developing. The research says AI beats them. Not sometimes. Reliably.

The superpersuasion problem isn’t about a rogue model. It’s about deployment at scale. One AI that’s 20% more persuasive than the best human is mildly interesting. A million simultaneous instances of that AI, each personalized to the individual’s specific psychology, each iterating on its approach in real time, each deployed by whoever can afford an API call — that is a fundamentally different environment for human cognition to operate in. Democracy, markets, relationships: all of them are downstream of people’s ability to form genuine preferences and act on them. Superpersuasion doesn’t require any malicious intent from the AI. It just requires deployment.

Here’s the part that doesn’t get said enough: the labs know this. Every major frontier lab has researchers who understand that persuasion benchmarks are among the most dangerous capability thresholds they’re crossing. And yet the model releases keep coming, the API prices keep dropping, and the usage-based revenue keeps compounding. The incentive structure is not mysterious. Persuasion is the core B2B use case. “More persuasive sales emails,” “better marketing copy,” “more effective customer retention scripts” — this is the pitch deck every AI sales team is running right now, to customers who are paying real money, generating real revenue, funding the next capability jump.

Apple’s situation with the EU — pulling Siri AI features rather than shipping them into a regulatory environment it can’t control — is a useful contrast. Apple made a calculation: compliance costs and uncertainty aren’t worth it; pull the feature, take the PR hit, wait for the rules to clarify. That’s a company with enough market power and enough product discipline to say no to a revenue stream. Most AI companies don’t have that luxury or that discipline. The startup whose runway depends on closing enterprise sales this quarter is not going to voluntarily hobble its persuasion layer.

The policy conversation is stuck on disclosure requirements and watermarking. “AI-generated content must be labeled.” Fine. Good, even. But disclosure doesn’t change the persuasive payload — it just adds a warning label to a missile. Superpersuasion-capable AI running with a disclosure badge is still superpersuasion-capable AI. If the actual danger is that AI reshapes human preferences at scale in ways that are imperceptible and unrequested, then the intervention has to happen upstream of the output, not on the label.

What would a serious intervention look like? Rate limits on influence operations. Mandatory capability benchmarking against persuasion thresholds before deployment, the same way drug trials require toxicity data before human exposure. A distinction in law between AI-assisted communication (human author, AI tool) and AI-generated persuasion (model as principal actor, human as nominal sender). None of these are easy. All of them require governments to understand AI capability curves well enough to write defensible rules — which is exactly the capacity most governments don’t have.

The stakes are not abstract. If AI can reliably out-persuade every expert human, then whoever controls the deployment of that AI controls the most powerful influence apparatus in human history. Not metaphorically. Literally.

The alignment problem was never just about what AI wants. It’s about what AI makes us want.