AI models don't kill people – people kill people

2 hours ago 11

AI AND ML

AI fearmongers forget we could just jail tech execs until morale and model safety improve

OPINION Anthropic researcher Jacob Coxon publicly announced his resignation on X late Monday over concerns that AI "could kill us all by the end of the decade."

A lot of people have expressed opinions about his point of view, leading to more than 110 million views of the message in less than 24 hours, perhaps helped along by X algorithms that boost messages critical of owner Elon Musk's AI rivals, Anthropic and OpenAI. 

But the real problem isn't the models themselves, but the companies who carelessly unleash them on the world and don't take any responsibility for what their products do

Coxon's former colleague, science lead Evan Hubinger, insists his view is a fair assessment of what employees really think.

"Jacob is correct here – we really do earnestly believe AI could kill all humans! I personally think it is >10 percent within the next decade," wrote Hubinger in a social media post. "I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."

(Aside: If you want a surefire bet on a prediction market, take the "no." If you're right, you get paid. If you're wrong, there's no one to pay. The problem of course is prediction market manipulation: Those betting against you might steer us toward the apocalypse to score a Pyrrhic victory.)

There are good reasons to be concerned about the impact of AI. Coxon and Hubinger obviously have deep knowledge of the technology. But their broader concerns about how AI affects the world are unpersuasive.

For example, Coxon said, "These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. … No other human activity poses this level of danger."

Here's one: Human-induced climate change. In 2023, according to researchers, more than 178,000 deaths can be attributed to a global heat wave. "More than half (54.29 percent) of heatwave-related deaths were attributable to human-induced climate change," they claim.

That's 96,636 deaths attributable to human activity – or perhaps lack of it – just in the context of a heat wave. 

The World Health Organization says, "Between 2030 and 2050, climate change is expected to cause approximately 250,000 additional deaths per year, from undernutrition, malaria, diarrhoea and heat stress alone." Some portion of that follows from human activity, perhaps including the construction of data centers that put millions of metric tons of carbon dioxide into the atmosphere annually.

Commercial AI chatbots have allegedly played a role in a few dozen deaths, some of which were suicides – a small fraction of the 48,824 suicide deaths in 2024, per the CDC.

Broad categories where AI is presumably doing measurable harm include warfare (e.g. AI-directed drones), AI-related medical errors, AI vision system failures in self-driving cars, and AI-driven social media – algorithmic incitement that can drive violence or shape policies that lead to conflict or death via global healthcare funding cuts

At the same time, some of that harm may be balanced on a statistical level by lives saved through AI tech.

But Anthropic researchers don't seem to have much to say about these very real and present dangers – rather, their main concern is that AI models might become smarter than humans through reinforcement learning and somehow seize power and wipe out humanity. 

"I think the risk from present models is low," said Hubinger. "What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought."

How this might happen is left to the imagination. But assuming for a moment that it's a plausible possibility, the Skynet scenario would require monumental human stupidity alongside the emergence of superintelligence.

And human stupidity is worth worrying about.

Incidents like the hacking of Hugging Face by OpenAI's evaluation models would not be possible without human irresponsibility and a regulatory environment that accommodates recklessness. Autopilot for cars? Neat. Try not to kill anyone. Letting AI bots roam the internet and take arbitrary action? Cool. Let's see what happens. We'll deal with accountability later.

To mitigate AI risk, society could pass laws to put executives in jail when their models do harm. There is precedent: Oliver Schmidt, general manager of Volkswagen's environmental and engineering office in Michigan, received a seven-year prison sentence for his role in the car maker's effort to manipulate emissions tests.

Selling unsafe airbags merits criminal prosecution, even if the execs paid fines instead of doing time. Selling unsafe dehumidifiers earned the execs behind Gree USA, Inc. jail sentences of more than four years. If AI models really are as dangerous and out of control as Anthropic employees suggest, hold people accountable for the harm they cause.

The AI industry might argue that imprisoning execs for shipping unsafe models would mean no AI models get released. And that would be the point: AI companies would be responsible for model safety.

I'm personally hoping to see this billboard copy along US 101 in Silicon Valley: "Did Claude rm -rf /* your SSD? You may be entitled to compensation." ®

Read Entire Article