35 Comments
User's avatar
Alexander McCoy's avatar

I appreciate you writing this, but as someone who is actually leading a team *fighting for* the regulation you say is needed, it is hard for me to get over the fact that the big Labs and their executives are certainly not ACTING like they want safety rules to solve the collective action problem.

Instead, they’re funding a political operation of unprecedented scale that is singlemindedly opposing ANY safety rules, seeking to overturn existing ones, and wielding hundreds of millions of dollars of super pac spending to intimidate all policymakers from going anywhere near the issue of AI safety. One in four federal lobbyists in Washington DC are currently employed by tech companies to lobby on opposing AI regulation. The forces that have been arrayed to explicitly prevent us from solving this problem are unlike anything I have ever seen, and I don’t think people in tech realize it.

If people in the Labs actually want safety rules that apply universally, then they need to do two things:

1) raise dissent within their companies over the actions of their government affairs teams and political action committees funded by their executives. Challenge the reassuring empty words being said at all-staffs and on Twitter about the leadership’s “good intentions” and contrast it with the concrete actions of the mercenary lobbyists and political operatives they have hired that say the exact opposite.

2) donate to pro-AI Safety organizations like Humans First, which are operating explicitly in the political/legislative advocacy domain. All the technical solutions and white papers don’t matter a jot if nobody will implement them. This is the high leverage, under-exploited area of AI safety.

Jay Dixit's avatar

Excellent point, and thanks for jumping in here. This is an important and thoughtful point.

To be honest, the lobbying side isn't my expertise, and you clearly know more about it than I do. You're no doubt right when you say "I don't think people in tech realize it."

Obviously lobbying against any specific regulation isn't itself proof of bad faith... and when it comes to calling out leadership, I can tell you that OpenAI's culture is already one of open, unrestrained dissent, both at all-hands meetings and on Slack.

But your larger point stands. I see what you mean about the disconnect between CEOs saying they want regulation and the lobbying operations they bankroll.

I think that tension is itself evidence of the coordination trap. The same competitive pressure that precludes labs from unilaterally slowing down ALSO pressures them to fight regulation that would constrain them before it would constrain their competitors.

To encapsulate it in a thesis: Lobbying is subject to the same incentive problem as everything else. The race dynamic doesn't apply only to testing timelines, it shapes politics too.

All of that still means, as you and I both agree, that the solution has to come not from industry but from public pressure, which is exactly your point. As I said in my article, that's the missing piece, so thank you for doing that work.

Tilia's avatar

> when it comes to calling out leadership, I can tell you that OpenAI's culture is already one of open, unrestrained dissent, both at all-hands meetings and on Slack.

Good.

> All of that still means, as you and I both agree, that the solution has to come not from industry but from public pressure, which is exactly your point. As I said in my article, that's the missing piece, so thank you for doing that work.

Why not both?

It seems to me that the leadership of OpenAI (and other labs too) are unusually wreck less, and unusually willing to gamble with risking the end of humanity, relative to you or me, or almost anyone.

The lobbying efforts doesn't look to me like just being against some specific regulation, that would put some specific AI company behind someone else. It looks like they are against anything that would make their wealth and power grow a bit slower in the short term, which includes regulations that would unilaterally slow down every AI lab.

I'm also not an expert on lobbying. But I'm a technical AI Safety researcher, so my information bubble is significantly different from yours. My general impression is that OpenAI (and others) are funding a lobby group, that just generally attacks, and tries to un-elect, anyone who proposes any AI legislations at all.

I.e. here's the latest news I saw:

https://thezvi.substack.com/p/ai-161-part-2-every-debate-on-ai?open=false#%C2%A7alex-bores-watch

(Under the heading "Alex Bores Watch")

I agree that the incentive situation is a large part of the problem. But on top of this, it does not look like OpenAI leadership is doing the best they can given their incentives. Since this is a hard coordination problem, we need everyone to actually try their best.

We need to both change the incentive (e.g. though public pressure), but also call out people doing unforced errors. You just told me that OpenAI leadership listens to insider dissent, which is great! Please use this power!

Please look into the lobbying that OpenAI supports.

Let me know if you want help in this. As I said, I'm also no a lobbying expert. But this is important. I can probably find someone for you to talk to who knows more.

Jeff Hatfield's avatar

Great piece. It made me think!

About incentives. It seems like one of the sub-incentives (is that even a word?) is benchmarks. Every new model that comes out dutifully highlights the benchmarks they excel at to prove their model is superior. I've never seen a benchmark on how moral or how ethical the model behaves. Maybe these benchmarks are out there, but I sure don't see them touted for new models. What if we made a benchmark like this that the AI labs could compete on?

Will Kiely's avatar

> But just because an argument is self-serving doesn’t mean it’s wrong. Critics refused to accept that two things could be true at once: (1) The CEOs have a financial incentive to make the claim; and (2) despite that incentive, it really is true that the frontier labs can’t afford to slow down because if they do, a worse actor will take the lead.

Good recognition of the fact that just because an argument is self-serving doesn't mean it's wrong!

I noticed that Karen Hao made this intellectual mistake *twice* in her recent Diary of a CEO podcast interview from last week, (1) when she dismissed the idea that AI poses an existential risk by saying that AI company CEOs profit by promoting this idea, and (2) when she dismissed the idea that scaling the current LLM-based AI paradigm will lead to AGI by saying that AI company CEOs profit by promoting this idea.

Will Kiely's avatar

The second example above: 1:14:56 Karen Hao: "And do you know what the common feature of all of them is? They profit enormously off of this myth." https://www.youtube.com/watch?v=Cn8HBj8QAbk&t=4496s

Jay Dixit's avatar

Exactly right. It assumes that self-interest is in itself proof of falsehood.

Valid skepticism: "The speaker benefits from saying this, so their claim merits closer scrutiny."

Fallacy: "The speaker benefits from saying this, so it can't possibly be true."

Will Kiely's avatar

Thanks for the write-up! It's useful for people to share the stories of when and why they changed their mind on important topics.

> The film makes a point I’d somehow never grasped. The problem isn’t that the people building AI are greedy, reckless, and unconcerned about the risks. The problem is that the system itself rewards speed over safety. Good intentions aren’t enough when the rules of the game punish restraint.

I notice you said you *somehow* never grasped it, presumably meaning that in retrospect it's surprising to you that you didn't grasp this point earlier. I too am quite surprised that you--a former OpenAI employee who has clearly paid a significant amount of attention to the discourse--did not grasp this point before seeing the film. What do you think prevented you from recognizing this fact? Any insight you can share would be helpful.

Jay Dixit's avatar

Good question! Answer: because all the criticisms I'd heard at the time were wrong on their face.

"The models are plateauing and will never get better" ➔ Wrong, I'd already seen the unreleased models.

"Everyone in the world should simply stop building AI" ➔ Great idea, let's also tell all the criminals to just stop doing crimes!

"AI companies don't care about safety" ➔ Wrong, I knew we cared deeply.

I dismissed them because they were wrong. The AI Doc was the first time I heard anyone make a case that didn't contradict the facts I knew to be true, i.e. "Yes AI companies care about safety, but we need better regulation anyway, and here's why."

Will Kiely's avatar

Wait, you're saying you literally never heard anyone point out that competition between AI companies means that if a company doesn't cut corners on safety then other companies that do will get ahead? This seems extremely implausible. I heard this literally hundreds if not thousands of times in the last decade and have articulated the point (many) dozens of times myself.

I had assumed you had heard the point made numerous times before but had just dismissed it for one reason or another.

Jay Dixit's avatar

The version I kept hearing was "Frontier labs are reckless and they're cutting corners because all they care about is winning the race." The film's more nuanced perspective — "some labs genuinely care about safety, but because of the race dynamic, that's not enough" — is a totally different argument. I dismissed the first argument because the premise was incorrect. The film's argument matches what I saw inside OpenAI. Plus, at the time, OpenAI was months ahead of any other lab, so the competition argument didn't feel that relevant.

That's what I mean when I say I didn't grasp it: not that no one ever mentioned competition, but that no one decoupled the race dynamic from the accusation of bad faith.

Will Kiely's avatar

Thanks, that helps somewhat. Though I'm still confused: If you dismissed the claim that OpenAI was cutting corners due to bad faith by telling yourself that you knew firsthand that OpenAI was operating in good faith, then what was your explanation for why OpenAI was cutting corners? If it wasn't due to bad faith and wasn't due to being forced into cutting corners due to the threat of competition catching up, then what was the reason for corner cuttinf, e.g. breaking the promise to devote 20% of compute to the superalignment team? Or did you just resolve this by telling yourself there wasn't any corner cutting going on at that time and that OpenAI's safety efforts were adequate? (E.g. Did you tell yourself that OpenAI promising to commit 20% of compute to superalignment in the first place was a mistake as there was no need for that much compute?)

Will Kiely's avatar

About a month after you joined OpenAI, OpenAI's Superalignment team was disbanded and the staff members who quit or were fired pointed out that OpenAI broke its public promise to commit 20% of its compute to alignment/safety work. When you heard of this, did you not recognize that the reason why OpenAI broke its promise was because of competitive pressure to use the scare computing resources on capabilities to keep up with competition? What did you tell yourself was the reason for this? https://finance.yahoo.com/news/exclusive-openai-promised-20-computing-105328622.html

Tony Northrup's avatar

Very well thought-out, Jay.

Jay Dixit's avatar

Thank you my friend! See what happens when YOU DON'T OUTSOURCE THE THINKING TO A MACHINE?

Colleen Avarene's avatar

Hey Jay — the two-part truth test is the most honest framing I've seen from someone who was actually inside the building. Both things being true simultaneously — "CEOs have financial incentives to say this" AND "the competitive dynamic is real" — is the part most people on either side refuse to hold at the same time. Critics want the claim to be pure rationalization. Insiders want it to be pure strategy. You're saying it's both, and that's where the actual problem lives.

I build custom AI agents and the race dynamic plays out at every scale, not just the frontier labs. Small builders cut corners on safety because the client wants it deployed yesterday and the competitor down the street doesn't bother with guardrails at all. The collective action problem isn't theoretical — it's Tuesday. The question is always: do you build the thing right and risk losing the contract, or build it fast and hope nothing breaks?

The nuclear arms treaty parallel is the right one but it took a near-miss at the Cuban Missile Crisis before anyone sat down at the table. I wonder what the AI equivalent of that near-miss looks like — and whether we'll recognize it when it happens or rationalize it away in a press release.

Becoming Human's avatar

‘“Why don’t we just ask the corporations to cease doing the thing they were founded to do?” is not a serious policy proposal.’

This may feel pragmatic, but it is not as justifiable a proposition as it sounds. Purdue had to be stopped doing what it was founded to do. Oil has to be stopped somehow. Clear-cutting corporate entities in South America have to stop.

When we take as axiomatic the existence of something that is existentially lethal, we are being complacent and intellectually suspect.

Being a corporation is not a license to existence.

Jim OB's avatar

It’s the byzantine generals problem. The AI industry should learn from the crypto industry.

Geoffrey Miller's avatar

Jay -- good post, and mostly reasonable.

But, you dismiss anti-AI activism too swiftly and mockingly, saying 'What’s needed, then, is regulation. Not some hippie-dippie appeal to a corporation to please just stop.'

Those of us involved in Pause AI, Stop AI, Control AI, and other anti-Ai activism groups are not just making 'hippie-dippie appeals' begging corporations to stop their reckless hubris.

Rather, we are revealing and amplifying the 'common knowledge' among citizens that we do not want Artificial Superintelligence, do not consent to it, and can fight against those who want to build it. We're not advocating for violence. But we are advocating for intense moral stigmatization against AI companies as institutions, and (at least some of us) against AI developers as individuals.

That kind of moral stigmatization of evil is an absolute prerequisite for AI regulation succeeding.

Historically, every successful fight against rich, powerful, institutionalized evils (e.g. the rush to build ASI before ASI alignment is solved) requires moral stigmatization of that evil. Not just national regulations and global treaties, but a public moral consensus that the evil is, in fact, evil. That's how slavery, sexism, and racism were (largely) overcome. Not just be writing laws against slavery, but my convincing people that slavery was morally wrong. Without that moral conviction, nobody's going to bother passing the anti-slavery laws, or enforcing them, or spreading them globally.

Lindsay Donaldson's avatar

I honestly believe slowing down to ensure we have more auditable AI that is actually aligned will make it better. Not just safer, but a genuinely higher quality product.

Currently, progress is mostly measured in benchmarks. But many benchmarks are brittle and fail to generalize. High scores don't necessarily mean the LLM performs reliably in the real-world.

Capability gains aren't useful if that capability can't be reliably used.

Therefore, I believe a slightly less powerful model with better reliability will win. Not just in safety, but competitively.

With such a model, its core capabilities would be consistently useful in real-world tasks. That makes it more competitively viable. What makes models "safer" can also make them perform better in most use cases.

If stricter regulations are the only force that will incentivize labs to design this way, so be it. But the competitive edge of making a better product design holds true either way.

---

EDIT: I realize “Therefore, I believe a slightly less powerful model with better reliability will win” might need to be clarified.

My point is, if we MUST choose between technically powerful models that crush benchmarks in the development race but generalize poorly, or models that are slightly smaller but generalize more reliably, then the smaller models are the better choice. They’d be more usable, and therefore, a better, safer, and more competitive product. The “G” in AGI stands for “general” for a reason.

Ron Bodkin's avatar

I agree with your analysis and recommendations that we need regulation to require safety for frontier AI and need international cooperation. It’s also unfortunate that the ill-informed arguments of the anti-AI left drowned out concerns about AI safety. EA has been talking about the incentive problem for a long time (e.g., see https://www.slatestarcodexabridged.com/Meditations-On-Moloch). I plan to write more about this kind of regulation vs pausing (spoiler alert - given that releases are already violating Preparedness Framework & Responsible Scaling Policies or at least tests are no longer able to rule this out, the labs would have work to do to continue the race before more releases).

I would also note that OpenAI co-founder Greg Brockman has donated to leading the future, a super PAC focused on blocking any AI regulation. And OpenAI has hired Chris Lehane to lead policy and aggressively attack those advocating for regulation: https://www.transformernews.ai/p/the-guerilla-warrior-who-taught-openai-chris-lehane?utm_campaign=post&utm_medium=web

Kevin's avatar

Thank you for writing this publicly!

I volunteer for a global community called Torchbearer, started by Connor Leahy, working on advocating for regulation and international agreements to prohibit the development of superintelligence. We work closely with ControlAI, started by Andrea Miotti. Would you consider working for a safety org to help stop the race to ASI?

https://www.torchbearer.community/

Also want to mention that OpenAI spends millions lobbying against regulation, which strengthens the argument that the least safe company will lead the race to superintelligence.

Jay Dixit's avatar

Thanks! I'll check out Torchbearer. Good point about the lobbying — another commenter made the same point. I think we can see it as further evidence of the coordination trap the film defines. The pressure to compete doesn't just affect how fast labs ship, it no doubt affects how they engage with policy too.

Sean Herrington's avatar

Thanks for the post, appreciate the insight into what people at the labs are thinking.

I feel like there's a point which has been missed by both you and commenters though which is that even with all your best people on the job doing their very best and caring as much as possible, if the problem is hard enough this doesn't matter.

Nobody wanted the challenger disaster to occur, and a lot of very smart very thoughtful people didn't want it to occur and spent their lives trying to prevent it from occurring. It still happened.

The Logosmitten's avatar

I hope you take this the right way, but the fact that you could not see this while being on the inside is more valuable of a statement in some ways than what you stated. Very meta.

Dimitry's avatar

Isn’t situation obvious? And how you’re include in ring and furthermore plan to control for example China?

Dominik Hörndlein's avatar

First of all, thanks for sharing your view from the inside, and reacting to some of the big critics!

Though I read some of your arguments with scepticism, I must say.

One of your main claims is that slowing down is dangerous, because if you do, somebody else (with worse intentions) will take the lead. But honestly, OpenAI started the race towards better and bigger models. This argument was there from the beginning, after the broad audience got aware of ChatGPT, and the "AI hype" took off. Other tech companies scrambled to catch up, and so did others beyond the industry ... Starting this race to get towards AGI asap, and then saying we cannot slow down because others could be faster is a circular argument.

With the huge (and well-earned) success of ChatGPT, OpenAI had the chance to lead with example and prioritize responsible AI before pure speed - but they chose a different path. I do believe that you have capable red teams and other units taking care of safety aspects - yet, you are fighting risks that OpenAI initially kicked off.

S. L. Sera's avatar

It's more than an AI arms race now, it has become a flywheel. Better AI means better production, which means better AI, which means better weapons. Weapons protect markets, which then funds more production, and so on...

Realpolitik projects this onto China and imagines a future where surveillance is administrative efficiency, censorship is social stability, and predictive policing is public safety.

The problem is that the machine is not culturally exclusive; the United States is already flirting with the same logic without needing to be conquered by it. Dissent has always been treated as a narrative problem. Now it is becoming a data-management problem. And power is hungry for data.

The Chinese model claims The State to be sovereign over narrative. It feels like the emerging U.S. model claims the Platform-State-Market complex is sovereign over risk.

The way out, if there is one, probably is not “stop the flywheel.” That seems unlikely; there is too much money, fear, ambition, and geopolitical pressure are already feeding it. The way out is to deny the flywheel moral sovereignty.

We can't let efficiency become the highest good. A society has to say, very clearly:

Some things must remain inefficient because humans are not inventory.

Some decisions require human judgment because prediction is not justice.

Some spaces must remain unmonitored because privacy is not suspicious.

Some speech must remain unruly because legitimacy requires dissent.

The practical counterweights are boring but crucial: Local resilience. Open-source tools. Data minimization. Strong warrant requirements. Public-interest technology. Antitrust enforcement. Analog fallbacks. Encryption. Unmonitored civic spaces. A political culture that treats dissent as a pressure valve rather than contamination.

A free society cannot define safety as the elimination of unpredictability.