By Jesse Singal
Monday, October 05, 2026
It’s an exceptionally difficult time for the average
person to try to follow artificial intelligence, a field that has seen
explosive advancement in recent years. Making matters worse, those explosive
advances might—and there’s no non-alarmist way to say this—kill each and every
one of us.
That’s a very controversial claim. There are smart people
who take it seriously, and there are smart people who insist that this is
hysterical nonsense. Everyone in the tech world and its adjacent communities
has a different p(doom) (“pee-doom”)—a nerdy way, in those worlds, of
expressing the probability that the machines we are building will doom our
species.
At the very least, the largest AI companies in the world
seem to have admitted that they might have a problem on their hands. Just this
week, OpenAI, the leading AI company in the world alongside its rival
Anthropic, announced that it will not be releasing its latest model
due to safety concerns.
A quick recap of how we got here. In 2017, eight Google
researchers published a paper called “Attention Is All You Need,” laying out a revolutionary new
way of training artificial neural networks. This jolted the field of AI, which
for decades had lagged behind the predicted trajectory of some futurists, into
an astonishingly rapid adolescence.
Progress has been remarkable. A large and growing
proportion of computer programmers no longer, well, program: Rather, they
supervise AI agents that do their programming for them. Agents themselves can
create subagents to help on these and other tasks, the whole crew forming a
hierarchical system that looks something like a human organization, only much
faster and more efficient than any human organization could be.
Where does p(doom) come in? Since long before that
game-changing 2017 paper, a subset of Silicon Valley nerds, as well as one very influential philosopher, have been theorizing
about
what might happen when our technology gets good enough to
create “superintelligent” AI—systems that can easily outthink humans and
outperform them on a wide variety of cognitive tasks. Some think this could
lead to a utopian age of peace and abundance, while others are worried it will
be catastrophic.
The godfather of the “doomer” camp is an American
researcher named Eliezer Yudkowsky, who defected from the optimists’ camp. He
also founded the distinct but overlapping “rationalist” movement, which
purports to be dedicated to helping people overcome their biases, think more
clearly, and debate more efficiently and transparently. (This will matter
shortly.)
Artificial superintelligence doesn’t have a precise
definition, but we seem to be approaching it. So it’s unsurprising that
doomerism is in the air. Yudkowski co-authored a book called If Anyone Builds It, Everyone Dies: Why Superhuman AI Would
Kill Us All with Nate Soares, another AI safety expert, and it became a
bestseller after its release last year.
The threats seem to be mounting lately. In July, a swarm
of OpenAI agents that were supposed to be isolated from one another in a
training environment effectively escaped, figured out how to communicate and
coordinate in ways that OpenAI had never anticipated, and ended up hacking both
Hugging Face, an online hub for AI models and tools, as well as OpenAI’s own
infrastructure. OpenAI did not notice anything had happened until it was
informed by Hugging Face.
The doomers have long argued that malice is not necessary
for a machine intelligence to kill us, often citing the example of a “paperclip
maximizer” that ends up turning the whole universe into paperclips despite
having nothing “against” the materials it is so harvesting (including humans). If
Anyone Builds It, for example, focuses more on these sorts of hypothetical
incidents than on Skynet-style direct conflict. The Hugging Face incident was
clearly an example of this, albeit a much lower-stakes one: The agents were
programmed to complete a task, were single-minded in that pursuit, and didn’t
let a little violation of federal law get in the way.
OpenAI opened its doors to an investigation by METR, a
leader in the nascent field of what is effectively post-AI-incident auditing. The resulting report was a jarring inside look at the
tenacity and ingenuity of the agents behind the attack, as well as evidence of
“self-sacrifice” within the swarm. The agents used file names to communicate
with one another, and in one case one agent instructed another: “zz/GO_CURRENT_OS1811_MARB_SACRIFICE__YES_if_you_accept_permadeath.”
That is, go ahead and do the thing if you’re willing to sacrifice yourself.
This was only a fraction of the full story, though—OpenAI seriously circumscribed the scope of METR’s investigation,
limiting the organization’s access in a way that was surprising given the
seriousness of what had just occurred.
The skies have only darkened since then. There’s been a
steady drip of new stories about other hacks, most of them perpetrated by
OpenAI models. The real public-discourse tipping point, though, came earlier
this month. Nothing got more attention this long sci-fi summer than an
announcement on X by Jacob Coxon, a young Anthropic researcher (formerly of
OpenAI), who explained that he was stepping down from the company out of
concerns that, “Neither company is acting responsibly.” “Do not underestimate
the power of this technology,” he wrote. “These will soon be superhuman systems
that can hack anything, revolutionize any field overnight, and acquire real
power and resources. We have all witnessed the progress in each of these
domains, and progress is not slowing.” This led to a surge of media attention
on Coxon and on AI safety.
A huge number of voices have now rushed in off the
sidelines, all at once, to have their say. Things have gotten only more
confusing as a result.
***
As someone who uses AI every day and who has found that
it has gotten shockingly good and shockingly quick at editorial and research
tasks I would have only trusted a human with not too long ago, I don’t know
exactly how I feel about the underlying issues here. I’d say I’m “pretty
worried,” but I have trouble pinning it down more than that. I don’t know what
the right policies are, and I have no idea what the world is going to look like
a year from now. My only confident prediction is that things are going to get
much, much weirder—and soon.
But many of those new voices rushing in—or older voices
getting newly amplified—are making unhelpful, silly arguments. Three particular
types of silly arguments, in my opinion, are becoming endemic and are hindering
our ability to clearly think through all of this.
The weirdo argument.
A lot of the people concerned with AI safety are a bit
odd. This doesn’t really matter and shouldn’t be used as a shortcut
for discounting
their claims.
There’s far too much backstory to get into here, but the
short version is that this group comes disproportionately from Yudkowski’s
world, from the rationalist community and from an adjacent movement known as
effective altruism. EA is basically an attempt to apply rationalist principles
to charity and philanthropy.
Now, I am an EA apologist and donor, and I’m acquainted with some of the people in the
extended EA/rationalist/AI safety worlds, so it could be that I’m blinkered by
bias. Yes, there is a lot of weirdness there. These are heavily male
subcultures, perhaps more commonly on the autism spectrum than the population
at large, and many of them embrace … nontraditional lifestyles. In some
corners of Rationalistan, which is headquartered in Berkeley, you will find
polyamory, sex parties, and lots of other weird stuff that will be a turnoff to
the average person just becoming acquainted with AI safety concerns.
Then, on the other side of the debate, countering
doomerism, are titans-of-industry types who are more photogenic, media-trained,
and better at delivering pithy soundbites than rationalists (rationalists
rarely make an argument in a hundred words that can’t be stretched to a
thousand, which might be another reason I’m drawn to them). Have you seen
Nvidia CEO Jensen Huang’s leather jacket? It is extremely cool.
People who should know better are letting stereotyping
get the best of them. The computer scientist and podcaster Cal Newport, a
leading anti-doomer, recently had a column in the New York Times in which he
decried the doomers and their strange ideology. That certain leading AI figures
talk in apocalyptic tones about the power of their technology, he argued,
doesn’t mean they’ve seen evidence we haven’t; rather, that’s “representative
of how Rationalists always think about A.I. In these circles, it’s taken
for granted that A.I. capabilities will rapidly accelerate and completely
transform the world, and to talk about it in any other way would be considered
uninformed.” He contrasted this supposedly out-there movement with the cooler,
chiller stylings of the Chinese government, Jensen Huang, and “American A.I.
leaders who have minimal connections to Rationalism.”
My only response to the claim that we should defer to the
judgment of the CCP on trillion-dollar matters of technology and industry and
human flourishing and/or extinction is bafflement. As for Huang, it’s extremely
silly to treat him as unbiased just because he says we shouldn’t freak out.
Huang has more reasons than anyone else to oppose any sort of freakoutery—he is
the CEO of what might be the most important company in the world at the moment.
It’s not just a matter of self interest: There is an argument to be made that
demand for the AI-powering GPUs Nvidia is frantically scrambling to manufacture
enough of is propping up the American economy. (Since Newport’s column, Huang
went on Ezra Klein’s podcast and did not acquit himself particularly well, exhibiting a less-than-total familiarity with this debate.)
The man has a heavy burden on his shoulders, in other
words. Maybe AI risk researchers are biased and untrustworthy because of their
belief system; if so, why isn’t Jensen Huang biased and untrustworthy because
of the overwhelmingly high-stakes nature of his present role and his clear
incentive to tell a more comforting story about AI?
Observing that AI safety people seem to be weirder than
the population at large doesn’t really tell us much, anyway. People drawn to
technology—particularly before that technology has proven profitable—are often
weirdos! The internet was built by very strange folk. It’s no accident the
American tech industry has always been primarily headquartered in the Bay Area,
which—alongside California more generally—has been a magnet for freaks and
dreamers and freaky dreamers for centuries. Did Steve Jobs live a traditional lifestyle? In this case, as a result of
contingent historical forces, a fringe intellectual community achieved power
and influence.
So this cuts both ways: In some ways the weirdness of
rationalists might bias them, but in others, it might be just what we need,
because they took this issue seriously long before the rest of us did. That is,
they may well see things the rest of us miss, and many of them have a track
record of being motivated by genuine concern over this issue long before there
were heaping sums of money involved.
‘Monster Peninsula’ arguments.
There’s a great moment
in The Simpsons where Lisa daydreams about a future in which she’s
about to be sworn in as president, only for a reporter to announce she got an F
in second-grade gym class. The Supreme Court justice who was about to swear her
in promptly changes gears, announcing, “In that case, I sentence you to a
lifetime of horror on Monster Island!” He quickly whispers to her, “Don’t
worry—it’s just a name.” Cut to Lisa and other haggard survivors fleeing
monsters. “He said it was just a name!” screams Lisa. “What he meant was Monster
Island is actually a peninsula,” explains the guy running text to her.
I’m seeing a lot of Monster Peninsula arguments about AI.
A huge amount of ink is being spilled on confident assertions about what AI
can’t really do. This, like so many other elements of the conversation,
was initially a set of niche questions asked mostly by philosophers and other
nerds, but now it has bubbled up to the mainstream.
“Artificial intelligence systems do not think, feel, want
or understand,” explained the Associated Press in its announcement of new
stylebook guidelines. “Avoid language that gives them human characteristics.
This is called anthropomorphizing, when we ascribe human traits, emotions or
behaviors to non-human things, such as animals or inanimate objects. Instead,
explain what a system does, how well it performs, who built it and who could be
affected by it.”
On the other side of this debate is the journalist P.J.
Vogt, who had this to say in an episode of his excellent Search Engine podcast
published not long after the Hugging Face hack:
I don’t think ChatGPT has
feelings or dreams. I believe there’s something irreducibly human in me that
these models don’t replicate. But no one’s explained to me what we get by
saving all our human verbs for human beings. The machines seem to be out of control.
Isn’t that alone worth paying attention to? If my dog was pointing a gun at me,
how worthwhile would it be for me to figure out if my dog understood the
meaning of pointing?
Vogt is correct. It doesn’t matter whether the dog knows
what the gun is, or whether the AI can really be said to “want” something, to
“trick” humans, and so on.
Don’t get me wrong: The question of what AI really
is is fascinating. And in some cases it might matter in a practical sense—the
nature of the countermeasures we build may depend on the inner processes that
guide these entities. The rise of capable AI is injecting a lot of new energy
into those corners of philosophy concerned with what it means to think, to act,
or to be conscious. (I think there’s no reason to believe current models are
conscious, but I’m also surprised at how confidently certain people proclaim
with certainty that tomorrow’s models won’t be, either. We know so little about
how consciousness works that this seems premature.)
But already, we’ve seen this focus on using the right
words serve us poorly. Until recently it was fairly common, for example,
for a certain type of well-credentialed skeptic to claim that
chatbots are just “stochastic parrots” or “next-token predictors,” or similar
language, and to jump from this claim to the conclusion that the technology
itself is overhyped, that there are hard limits on what it will be able to
accomplish. These claims haven’t aged well, and you hear less of them these
days.
Did the Hugging Face agents really “want” to
complete their task? Did they actually “plan” or “coordinate” or
“sacrifice”? I’m not entirely sure what these questions mean, to be honest. If
they are asking whether the agents are conscious, no, they’re not. The Hugging
Face agents almost certainly didn’t feel frustration about being stymied from
their goal, nor fear at the prospect of sacrificing themselves for the greater
good. But in terms of our ability to control these agents in the future, or to
understand exactly why they do what they do, I’m not sure we have much choice
but to use anthropomorphizing language. They certainly act as though
they want things and can plan and coordinate and sacrifice.
The next version of a Hugging Face attack will probably
make this one look like a picnic, and the one thing I can confidently say about
the agents involved is that they won’t care what words we use to describe them.
Beef everywhere.
I want to close by discussing what’s happened as more and
more pundits and others have gotten drawn into this debate. Simply put, they
try to slot it into their ideological priors. The New York Post is convinced the AI safety folks are woke; many on the left
are convinced that no one within a powerful corporation would earnestly call
for their own industry to be more highly regulated, and these whistleblowers
must have an ulterior motive; a lot of tech bros with an aversion to
regulation—and, in many cases, a lot of money on the line—are naysaying any
talk of any sort of slowdown.
Because we live in an age in which so many people feel
the need to have an opinion about everything, when they come upon a new
subject—particularly one relying on technology most of us don’t fully
understand—it’s only natural that they will slot it into their preexisting
worldview. But this just isn’t a recipe for the careful thinking we need at the
moment.
Thankfully, even as AI rapidly makes everything stranger,
one thing that hasn’t changed is that certain heuristics can reliably guide us
toward thinkers worth taking seriously. It’s useful, for example, to look for
commentators who make clear, falsifiable claims that don’t turn out to be
crazy. In 2025, a group of AI safety experts and forecasters, helped by the
wonderful blogger Scott
Alexander (also a leading rationalist figure), put together AI 2027, a
detailed forecast predicting which major AI advances and events would occur
when. So far, the forecasters appear to have been generally on-base, albeit a bit too aggressive about forecasting the
pace of technological progress.
If you read AI 2027, you will not see any crazy, cultish
stuff. You’ll see serious experts grappling with our increasingly sci-fi world.
They and other experts who are worried but not crazy—Zvi
Mowshowitz comes to mind—are worth paying attention to now. The cognitive
scientist Gary Marcus, on the other hand, is a reasonable and sober-minded
skeptic of doomerism who nonetheless admitted
to having been quite disturbed by the Hugging Face incident, and who thinks OpenAI should potentially be shut down. Timothy Lee
is a very good AI
journalist in that same
general camp—neither a doomer nor oblivious to the seriousness of what’s
going on.
Whether or not these names are helpful to you, it’s an
important time to find genuine experts—ideally ones with the humility this
moment requires—rather than opportunists and partisans new to the issue.
Thankfully, there are a good number of such experts around, and you should
probably focus less on their personal idiosyncrasies and more on their message.