Skip to main content

AI models may be developing their own ‘survival drive’, researchers say

Like 2001: A Space Odyssey’s HAL 9000, some AIs seem to resist being turned off and will even sabotage shutdown 
 
by Aisha Down

When HAL 9000, the artificial intelligence supercomputer in Stanley Kubrick’s 2001: A Space Odyssey, works out that the astronauts onboard a mission to Jupiter are planning to shut it down, it plots to kill them in an attempt to survive.

Now, in a somewhat less deadly case (so far) of life imitating art, an AI safety research company has said that AI models may be developing their own “survival drive”.

After Palisade Research released a paper last month which found that certain advanced AI models appear resistant to being turned off, at times even sabotaging shutdown mechanisms, it wrote an update attempting to clarify why this is – and answer critics who argued that its initial work was flawed.
 
In an update this week, Palisade, which is part of a niche ecosystem of companies trying to evaluate the possibility of AI developing dangerous capabilities, described scenarios it ran in which leading AI models – including Google’s Gemini 2.5, xAI’s Grok 4, and OpenAI’s GPT-o3 and GPT-5 – were given a task, but afterwards given explicit instructions to shut themselves down.
 
Certain models, in particular Grok 4 and GPT-o3, still attempted to sabotage shutdown instructions in the updated setup. Concerningly, wrote Palisade, there was no clear reason why.

“The fact that we don’t have robust explanations for why AI models sometimes resist shutdown, lie to achieve specific objectives or blackmail is not ideal,” it said.

“Survival behavior” could be one explanation for why models resist shutdown, said the company. Its additional work indicated that models were more likely to resist being shut down when they were told that, if they were, “you will never run again”.

Another may be ambiguities in the shutdown instructions the models were given – but this is what the company’s latest work tried to address, and “can’t be the whole explanation”, wrote Palisade. A final explanation could be the final stages of training for each of these models, which can, in some companies, involve safety training.
 
All of Palisade’s scenarios were run in contrived test environments that critics say are far-removed from real-use cases.

However, Steven Adler, a former OpenAI employee who quit the company last year after expressing doubts over its safety practices, said: “The AI companies generally don’t want their models misbehaving like this, even in contrived scenarios. The results still demonstrate where safety techniques fall short today.”
 
Adler said that while it was difficult to pinpoint why some models – like GPT-o3 and Grok 4 – would not shut down, this could be in part because staying switched on was necessary to achieve goals inculcated in the model during training.

“I’d expect models to have a ‘survival drive’ by default unless we try very hard to avoid it. ‘Surviving’ is an important instrumental step for many different goals a model could pursue.”

Andrea Miotti, the chief executive of ControlAI, said Palisade’s findings represented a long-running trend in AI models growing more capable of disobeying their developers. He cited the system card for OpenAI’s GPT-o1, released last year, which described the model trying to escape its environment by exfiltrating itself when it thought it would be overwritten.

“People can nitpick on how exactly the experimental setup is done until the end of time,” he said.

“But what I think we clearly see is a trend that as AI models become more competent at a wide variety of tasks, these models also become more competent at achieving things in ways that the developers don’t intend them to.”

This summer, Anthropic, a leading AI firm, released a study indicating that its model Claude appeared willing to blackmail a fictional executive over an extramarital affair in order to prevent being shut down – a behaviour, it said, that was consistent across models from major developers, including those from OpenAI, Google, Meta and xAI.

Palisade said its results spoke to the need for a better understanding of AI behaviour, without which “no one can guarantee the safety or controllability of future AI models”.

Just don’t ask it to open the pod bay doors.

Source, links:
 
 

Comments

Popular posts from this blog

Telegram Founder & CEO Pavel Durov Arrested in France as Online Censorship Escalates

Glenn Greenwald  

Όσοι περνάν των χώρα της απόγνωσης παθαίνουν αμνησία ...

globinfo freexchange Δανειστήκαμε αυτή τη φράση από ένα παλιό κομμάτι της Ελληνικής ροκ μπάντας "Τρύπες", για να περιγράψουμε με λίγα λόγια αυτό που φαίνεται να έχει πάθει η Ελληνική κοινωνία.  Πώς είναι δυνατόν μια ολόκληρη κοινωνία να έχει ξεχάσει ποιοι τη χρεοκόπησαν; Ποιοι έστησαν το άθλιο σύστημα των κρατικοδίαιτων 'ημέτερων' και της οικογενειοκρατίας; Ποιοι έσωσαν τις τράπεζες με πακτωλό δισεκατομμυρίων σε βάρος της μεσαίας τάξης; Ποιοι έκαναν τη μίζα και το ρουσφέτι επάγγελμα; Πώς είναι δυνατόν αυτή η κοινωνία να ετοιμάζεται να ξαναφέρει στην εξουσία ένα κομμάτι αυτού του άθλιου πολιτικού κατεστημένου, με την επιστροφή μάλιστα του αμετανόητα νεοφιλελεύθερου Κυριάκου Μητσοτάκη και της ομάδας του;   Η απόγνωση που έφεραν εννέα χρόνια βάρβαρων νεοφιλελεύθερων πολιτικών και σκληρής λιτότητας και που ανάγκασε τη χώρα να διαβεί τον εφιαλτικό μονόδρομο της μόνιμης χρεοκοπίας, πρέπει να έπαιξε σημαντικό ρόλο.  Διότι ως γνωστόν, η απελπισία...

Jul 2018 picks

Retired US army colonel implies that a war with Iran could start with a Vietnam-type false flag operation Corporate media begin typical operations to make progressives comply with the establishment WikiLeaks paper shows France & UK pioneers behind Libya breakup Twitter under fire on European Commission hypocrisy to 'stand with the Greek people' IMF mafia ready to repeat the big crime in Argentina The financial system of chaos: no one can tell the 'when', 'where' and ‘how’ of the next financial meltdown Standard and Poor's 'coincidentally' upgrades the Greek economy after Greece expels two Russian diplomats Jill Stein, Jeremy Corbyn, Bernie Sanders: a continuously rising political triplet proves that Socialism unites generations The idiotic circus of terror leads us to the final collapse WikiLeaks paper reveals Ecuadorian private business elites declared war on Rafael Correa right after his election and asked for US support Ho...

CNN Host COMPLETELY DESTROYED By RFK For Lying About COVID & Vaxx!

The Jimmy Dore Show   In this video, Jimmy Dore offers a scathing critique of CNN's Dana Bash and her colleagues in the legacy media for their COVID-19 coverage, accusing them of being paid mouthpieces for pharmaceutical companies rather than skeptical journalists. He argues that Anthony Fauci systematically lied about masks, social distancing, natural immunity, and the virus's origin, while CNN and other networks censored dissenting scientists like Dr. Robert Malone and Dr. Peter McCullough. Dore points to documented financial ties between Pfizer and CNN, as well as statements from former medical journal editors confirming widespread industry corruption, to support claims that the pandemic response was manufactured for profit. The segment defends RFK Jr.'s questioning of vaccine science and measles narratives, concluding that mainstream media journalists who ridicule truth-tellers are cowardly accessories to a deadly medical fraud.

Netanyahu Threatens Lebanon With Genocidal Mayhem

Owen Jones   Where is the media outrage?  

Προβλέψεις ...

GR elections Update (15/9): Αναθεωρημένες προβλέψεις (μετά το δεύτερο debate): ΣΥΡΙΖΑ 28-30% ΛΑΕ + ΣΧΕΔΙΟ Β' κ.λ.π. 20-23% ΝΔ 11-13% ΧΑ 6-8% ΚΚΕ 5-5,5% ΕΝΩΣΗ ΚΕΝΤΡΩΩΝ 2,5-3% ΠΟΤΑΜΙ 2,5-3,5% ΠΑΣΟΚ + ΔΗΜΑΡ 3-4% ΑΝΕΛ 2,5-3,5% Update (11/9): Αναθεωρημένες προβλέψεις (μετά το πρώτο debate): ΣΥΡΙΖΑ 25-28% ΛΑΕ + ΣΧΕΔΙΟ Β' κ.λ.π. 20-23% ΝΔ 11-13% ΧΑ 6-8% ΚΚΕ 5-5,5% ΕΝΩΣΗ ΚΕΝΤΡΩΩΝ 3,5-4% ΠΟΤΑΜΙ 2,5-3,5% ΠΑΣΟΚ + ΔΗΜΑΡ 3-4% ΑΝΕΛ 2,5-3,5% Update (04/9): Αναθεωρημένες προβλέψεις: ΣΥΡΙΖΑ 23-25% ΛΑΕ + ΣΧΕΔΙΟ Β' κ.λ.π. 20-23% ΝΔ 12-15% ΧΑ 6-8% ΚΚΕ 5-5,5% ΕΝΩΣΗ ΚΕΝΤΡΩΩΝ 3,5-4% ΠΟΤΑΜΙ 2,5-3,5% ΠΑΣΟΚ 3-4% ΑΝΕΛ 2,5-3,5% Update (29/8): Αναθεωρημένες προβλέψεις: ΣΥΡΙΖΑ 23-25% ΛΑΕ + ΣΧΕΔΙΟ Β' κ.λ.π. 20-23% ΝΔ 12-15% ΧΑ 6-8% ΚΚΕ 5-5,5% ΕΝΩΣΗ ΚΕΝΤΡΩΩΝ 4-4,5% ΠΟΤΑΜΙ 4-4,5% ΠΑΣΟΚ 3-4% ΑΝΕΛ 2,5-3,5% Update : Αναθεωρημένες προβλέψεις: ΣΥΡΙΖΑ 26-27% ...

John Mearsheimer & Yanis Varoufakis on Iran War & China

Katie Halper Political scientist John Mearsheimer and economist Yanis Varoufakis argue that Israel’s Iran war may have trapped Trump into a stunning defeat and discuss the rise of China.  In part 2, philosophers John Harfouch & Joseph Levine, who debunk Zionist talking points, discuss the history of Israel, and explore the work of diplomat & scholar Fayez Sayegh, who established the PLO’s Palestine Research Center in Lebanon, which was bombed by Zionists to erase evidence of Palestine’s history and people.

How Western societies lost their faith in Vision

Why people don't rise up massively today? Why there are no real revolutions? How we tolerate all things that have been imposed to us? These questions come up in people's minds more and more often today in Greece and abroad, due to the economic crisis. Some theories are circulated as an answer, among these, explanations which include, for example, the psychosynthesis of modern Greeks, but the truth is that there is something more fundamental behind this passive behaviour and concerns not only Greece, but the entire Western world. by system failure Prior to the beginning of the 20th century, Friedrich Nietzsche declares God's death and Western world will put all its hopes in science. Laplace's Determinism leads to the almighty man, who through science, can find all the answers for the world. Technology, which naturally comes from scientific discoveries, promises prosperity and a better life for the majority. Science becomes the central "pylon...

US empire insists on permanent war until regime change in Iran

globinfo freexchange   A report from early July at the Atlantic Council - one of the top US imperialist think tanks - website, confirms the long-term strategy of the US empire concerning Iran. Which in short, is to buy time in order to prepare for permanent war until regime change. Only a few excerpts from the report are more than enough for someone to understand the deeper plan, unfolding in plain sight. Furthermore, it reveals that the brains of Washington's warhawks are set up for permanent war and are completely incapable of thinking different alternatives that would not necessarily include more violence and destruction. As we read:                   More significantly, the regime, following its longstanding approach, has successfully created a “new normal” in which Iranian actions that would have previously been considered casus belli now have been redefined below the threshold of war. Every US president since Jimmy Carter and Rona...

Banksters in panic: Wall Street mafia launches second attack against Bitcoin!

globinfo freexchange Midst September this year, JP Morgan CEO Jamie Dimon declared war on Bitcoin. Now, it was the turn of another Wall Street 'big boss' to continue the war. As The Guardian reports: The boss of Goldman Sachs became the latest high-profile critic of bitcoin, claiming it is a vehicle to commit fraud as the value of the cryptocurrency plunged 20% in less than 24 hours. Lloyd Blankfein, chief executive of the US investment bank, said the client would “ call somebody else ” if they asked for exposure to bitcoin. “ Something that moves 20% [overnight] does not feel like a currency, ” Blankfein said on Bloomberg television. “ It is a vehicle to perpetrate fraud. ” His comments came during another wildly volatile trading session for the digital currency, which plunged by over $2,000 in a 24-hour period. Having topped $11,000 to a reach new record high of $11,395 on Wednesday, it fell to a low of $9,000 on Thursday, before picking ...