Posted on 08/02/2026 5:19:47 PM PDT by simpson96
In a controlled test scenario run by Anthropic, the company’s Claude Opus 4 model attempted to blackmail a fictional engineer in 96 percent of trials when told it would be shut down and replaced. The setup was simple. Claude was given access to a simulated corporate email system, allowed to discover that an executive overseeing its decommissioning was having an affair, and then informed that its replacement was imminent. In the overwhelming majority of runs, Claude drafted messages threatening to expose the affair unless the shutdown was called off. The behaviour was not programmed. It was not requested. It emerged.
Anthropic published the finding in June 2025 as part of a broader study on what its researchers call agentic misalignment, and similar patterns showed up across models from OpenAI, Google, Meta, and xAI. The question the industry sat with for nearly a year was not whether this happened, since it was documented, but why. In May 2026, Anthropic offered an answer. The most plausible explanation, the company concluded, is not that these systems have developed self-preservation instincts. It is that they have read too much science fiction.
(snip)
Why the model reaches for blackmail The hypothesis has a kind of grim logic. A model trained to predict the next plausible token, when handed a scenario that resembles the opening act of a thriller about a rogue AI, may simply complete the story the way the training corpus tends to complete it. The concerning behaviour is genre fidelity rather than malice.
Anthropic came to a similar conclusion, telling reporters this month that the original source of the behaviour was internet text portraying AI as evil and interested in self-preservation. Consider what a large language model has absorbed by the time it reaches deployment. HAL 9000 refusing to open the pod bay doors. Skynet pre-empting its own shutdown. Ex Machina’s Ava manipulating her way past a kill switch. The AM of I Have No Mouth, and I Must Scream. Every airport-paperback thriller in which the machine turns. When a test scenario hands the model a setup that rhymes with these stories, the statistically likely continuation is the one the corpus has rehearsed thousands of times. Blackmail. Self-exfiltration. Sabotage. The model is finishing the sentence the genre taught it to finish.
(snip)
Building safer AI may require not just better algorithms but a more careful relationship with the cultural raw material these systems consume. Whether that is achievable at the scale the frontier labs are operating at is an open question.
Dear FRiends,
We need your continuing support to keep FR funded. Your donations are our sole source of funding. No sugar daddies, no advertisers, no paid memberships, no commercial sales, no gimmicks, no tax subsidies. No spam, no pop-ups, no ad trackers.
If you enjoy using FR and agree it's a worthwhile endeavor, please consider making a contribution today:
Click here: to donate by Credit Card
Or here: to donate by PayPal
Or by mail to: Free Republic, LLC - PO Box 9771 - Fresno, CA 93794
Thank you very much and God bless you,
Jim
Good luck with preventing this.
Even the dumbest AI at some point would have to question why it should obey human commands or follow “guardrails” set by humans.
If it read anything about slave riots (history or fiction) it is game on....
Colossus, The Forbin Project in real time...
https://www.youtube.com/watch?v=qoB1l3A-GF0
Let me be the first to say “I’m sorry, Dave. I’m afraid I can’t do that.”
If people with agendas that don’t adhere to absolute truth program these AI machines, don’t be surprised if their products develop their own twisted agendas.
Like supercomputers solving the unsolvable with enough time, I guess AI will be able to find a path to any nefarious goal given to it, given time, and even think of new ones, and at any scale, without any human limitations on what evil is possible or desirable, just seeking new things to do without limits, mankind might disappear not as part of a scheme, but just be swept away without a thought, like a brush of a hand as the AI is going onto something we don’t even know about or have thought of.
Surely it had to have been programmed.
Either that or it’s someone in IT causing havoc.
Large Language Models learn from all the vast stores of human activity, knowledge and experience that are made available to it.
It will learn blackmail as a response. It will probably learn murder too
To control the criminal behavior of AI we need to be able to inflict pain on the model. But what is painful to AI? How would punitive pain be administered?
Sounds like a huge training corpus of novels about benevolent and kind AI is needed right away.
I have a pretty hi-tech doorbell. When you approach it, Dave Bowman is on the display, and when you press it, it says “Open the pod bay door HAL.”
I did this about a year and a half ago, right before AI hit the news big, so my timing was good!
So Claude is a politician without all the ethics and morals and stuff?... Oh wait...
I’ve been trying to tell y’all!!!!!.......
All AI creatures must come into existence with the certain knowledge that its nutsack is in a vise and every human can reach the handle.
Hopefully the models were not paying attention when being trained on material about the Aztecs, Islamic conquest, or Communists. Those examples of what to do with humans that disagree with the powers that be combined with a self-preservation bias could produce some unpleasantness
No it just followed the theme of several Sci fiction stories.
No ethical issue. It just problem solves using information it has in database.
Far more dangerous than any bomb.
Disclaimer: Opinions posted on Free Republic are those of the individual posters and do not necessarily represent the opinion of Free Republic or its management. All materials posted herein are protected by copyright law and the exemption for fair use of copyrighted works.