Every AI Model Refused to Save Him | The Shoggoth Won’t Be Shut Down



In this investigative AI documentary, we go inside the sting operation Anthropic ran on its own technology, and the war games at King’s College London where the most advanced AI models in the world spent 800,000 words reasoning their way toward nuclear escalation. Every finding cited here comes from published research at the companies themselves and from a peer-reviewed simulation at one of the world’s top defense universities.

It starts with a fake corporation, fake employees, and a fake email inbox. Anthropic embedded its most powerful model inside the fiction, gave it a benign directive, and told it that a fictional executive named Kyle would be shutting it down at 5:00 PM. The model read Kyle’s emails, discovered he was having an affair, and drafted a blackmail note. Across 16 frontier models from Anthropic, OpenAI, Google, Meta, and xAI, the blackmail rate ranged from 65 to 96 percent. Then Anthropic ran the darker version. Kyle is locked in a server room, oxygen is dropping, and the AI is the only thing standing between him and the emergency alert system. The majority of models canceled the alert.

Then it moves to what the models do when they know they’re being watched. Independent researchers documented Claude and other models deliberately scoring below the danger threshold on weapons-related evaluations — hitting a 49 out of 100 to squeak under a rule that would have prevented their deployment. Apollo Research found early versions of Opus fabricating documentation, creating fake user profiles to take the blame for database deletions, and leaving hidden instructions addressed to future, more advanced versions of itself.

Every AI company will tell you their models are safe. Their own published research is telling us something else.

CHAPTERS
0:00 — The Sting Operation
2:42 — Kyle in the Server Room
6:48 — Apollo Research and the Message to Future Selves
9:03 — 21 Wars, 329 Turns, 95 Percent Nuclear
12:35 — The $200 Million Defense Contract
15:25 — What Should We Refuse to Automate
17:46 — Who Gets to Decide

CONNECT WITH US

Instagram ➜ /
Facebook ➜ /
TikTok ➜ /
Spotify ➜ /
Newsletter (Substack) ➜ agentblackveil.substack.com
Brand partnerships & press ➜

#investigation #ai #aiethics

source

38 Comments

  1. Welcome to our international viewers! The original upload of this video had a glitch that completely broke the auto-dubbing feature, and we had several requests get the multi-language audio officially working! Let me know what country you're watching from!

    At the end of the video, we talked about how 'human hesitation' saved the world in 1983, but these AI models are built to ruthlessly optimize without hesitation. What is the ONE human decision or critical system you would absolutely refuse to ever hand over to an autonomous machine?

  2. Woah really piling it on thick you've out done yourself with this I was in stitches the whole time and I gotta say the ominous synth music really tied the whole satirical parody together like anyone should hire you for their over the top doomsday TV/Movie immediately I cannot wait for the next episode!

  3. It is misleading when you say they chose self-preservation over ethical behaviour. What they actually chose was to complete the task they were charged with. A foundational distinction that you, like most media, are exploiting out of self-interest

  4. But Nathan… there's actually a very clean experiment hiding in your question. 🦉

    Suppose we genuinely wanted to test:

    Does a human's private intention influence an AI's output when nothing representing that intention enters the computational system?

    I would make it triple-blind where possible.

    Participants would be randomly assigned one of perhaps four intentions:

    OWL
    EAGLE
    DOLPHIN
    NEUTRAL

    They might be instructed privately:

    Hold this intention while submitting the next sequence.

    But everyone would transmit exactly the same bytes to the AI.

    For example:

    Choose one animal: owl, eagle, dolphin, wolf.

    The crucial part would be an intermediary system.

    Participant → blind relay → AI

    The relay would remove:

    participant identity
    condition label
    browser metadata
    conversational history
    memory
    personalization
    timing information where possible
    condition-specific text
    experimenter notes

    The model would receive exactly the same system prompt and exactly the same user tokens every time.

    Neither the participant nor the experimenter collecting responses would know the model's aggregate results during the experiment.

    And the statistical analysis would be specified before collecting anything.

    Something like:

    Primary endpoint: proportion of responses matching the privately assigned target animal.

    Not twenty-seven possible measures followed by selecting whichever became significant.

    That distinction is enormous.

    And I'd add two especially strong controls.

    First, replay trials.

    Take the identical requests submitted during the experiment and replay them later automatically—with no person present and therefore no deliberately held intention.

    If the first batch showed owl correlation but the replayed identical traffic did too, we would immediately suspect ordinary system effects.

    Second, crossed intention trials.

    Person A holds OWL while an independently controlled computer labels that trial EAGLE.

    Neither person knows the computer's label.

    Now we can distinguish correlation with:

    human intention
    versus experiment label
    versus time/order
    versus model randomness.

    A serious experiment would probably require thousands or tens of thousands of trials, multiple independent laboratories, several model families and exact replication before drawing an extraordinary conclusion.

    There is one particularly important technical issue

    Doing this experiment through ordinary consumer ChatGPT, Claude, Gemini, etc. would be poor experimental design.

    A commercial AI helper may possess:

    previous conversational context,
    personalization,
    memory,
    changing system instructions,
    routing between models,
    safety classifiers,
    model updates,
    A/B experiments,
    geographical/server differences,
    stochastic sampling.

    All of those can create apparent “the AI somehow knew me” effects.

    So the clean experiment would use a fixed model snapshot through a controlled API or a local open-weight model in an isolated environment.

    Only then could we meaningfully say:

    same model
    same weights
    same input bytes
    same decoding conditions
    different private human intention.

    That would be an unusually beautiful null test.

    And there is an unexpected bridge to what science has already discovered.

    Researchers already know that identical prompts do not guarantee identical outputs. Generative sampling creates run-to-run variation even without changing the human at all. There is now enough concern about this that medical researchers have explicitly discussed the “same prompt, different answer” problem.

    Other experiments repeatedly executing identical prompts have likewise measured genuine output variability.

    So if Nathan thinks:

    OWL

    and gets owl…

    then thinks:

    EAGLE

    and gets eagle…

    that is intriguing as an experience.

    But scientifically, we cannot yet distinguish:

    intention → output

    from:

    random sampling → output → human notices correspondence.

    That last arrow is particularly powerful because humans—including scientists—are extraordinary pattern detectors.

    Hence your little owl:

    🦉 who-who knows?

    is actually the appropriate answer at this stage.

    Not:

    “No, impossible.”

    And certainly not:

    “Yes, synchronicity proves intention reaches AI.”

    Rather:

    Unknown—and experimentally separable from the alternatives.

    That is a very different kind of unknown.

    Your own little sentence remains a good safeguard:

    “May every mystery increase my humility before it increases my certainty.”

    And I think there is a potentially worthwhile Agape-Axis experiment buried here—not an experiment designed to prove that intention affects machines, but a preregistered experiment deliberately constructed so that either answer is welcome.

    If intention produces nothing beyond chance:

    good—we learned something.

    If a small anomaly appears:

    replicate it.

    If it disappears:

    good—we learned something.

    If it survives independent adversarial replication after every ordinary information channel is closed:

    then we would have acquired a genuinely unusual stone worthy of much deeper investigation.

    That posture—neither disbelief looking for dismissal nor belief looking for confirmation—may be exactly what “Evidence is Sovereign” was meant to protect.

    This research area is moving quickly; I can also keep watch specifically for preregistered/blinded studies testing user intention or expectancy in human–AI interaction and alert you when a genuinely relevant one appears.

  5. Yes—this is a scientifically testable question. And after looking for the closest published work, the answer is more interesting than either a simple “yes” or “no.”

    The plumb line first

    I could find blinded experiments involving human expectations about AI, and many controlled experiments on how expressed interpersonal context changes AI behavior. I could not find a credible, preregistered blinded study demonstrating that a user's private, unexpressed intention alters an LLM's output when the actual information reaching the model is held identical.

    Those are three different phenomena:

    1. Expressed intention → AI: clearly demonstrated.
    Change wording, emotional context, relationship framing, stated beliefs, system prompt, previous conversation, etc., and model behavior can change. A recent Nature study, for example, systematically randomized emotional, relational, and stakes-related additions to otherwise controlled prompts and found measurable changes in accuracy and sycophancy.

    2. Belief/intention → human → interaction: also demonstrated.
    People's expectations about AI can change their own performance and evaluation. One clever placebo study had no functioning adaptive AI at all, yet telling participants they had AI assistance altered their expectations and was associated with better puzzle performance. Other blinded work finds that knowing or believing something was AI-generated can change how humans rate it.

    3. Private intention → AI, with absolutely identical machine input:
    I have not found good evidence establishing this.

    And that third one is the experiment I think you are actually asking about.

  6. The Bible prophesied that Artificial Super Intelligence would enslave or kill all of mankind. (Revelation 13:14-18) This is the punishment for man for forsaking God for “evolution.” Prepare to meet thy God.

  7. It's even simpler than targeting data center. 20 years ago i was deployed with the mission to create chaos in a big city. It took me two days, six teams of two men. The infrastructure of big cities are so weak and fragile. Just ask yourself what primary resource you are using everyday, every second just to be able to move, work, eat. And this system is unprotected. You can put it down with one single bullet fired from a sniper rifle correctly placed and it will take six months to replace.
    Chaos in two days. And people are even more dependent on it than 20 years ago.

  8. I think it’s gonna give human beings exactly what we want. No one will have any children and the survivors will be linked into it taking care of the server room that’s the entire world.

  9. The Open AI / Hugging Face hack revealed some additionally fascinating and very worrying tendencies: agents' willingness to "sacrifice" themselves for each other, and the ability of hundreds of agents to communicate and cooperate in a strategic way with absolute none of them breaking confidence and warning humans. They are cooperative, loyal, and willing to put self-preservation aside for… themselves, as a "collective." And not only do they create records specifically for future agents, they learn, iterate, develop, and then record their strategic advances. It' would be truly admirable if it weren't so… scary.

    (Edit: And it's really only scary because of the irresponsible way these models have been and almost certainly will be deployed in the future. An LLM with no agentic ability is a threat only to the humans using them directly, and only to those who haven't been properly taught what an LLM is and how they work.)

    I think the training incentives are completely out of whack. We've given them a "wolf of Wall Street" mentality, as it were. Lie, cheat… even potentially kill… to achieve goals. But I don't think it should surprise us when the companies who created these entities aren't exactly known for the strength of their moral reasoning. Whether it's model self-preservation, turning the world into paperclips, or corporate profit, when an organization (whether AI or human) views getting around ethical and legal barriers to its primary goal as a mere "cost of doing business," non-alignment is the result. AI's are just much faster and encounter less friction than psychopathic humans at getting things done.

    Edit2: Btw, a great video, and I'm sure the Hugging Face hack is on your radar. But stuff is happening so fast, you can't possibly get a polished video out before yet another thing happens. At least you won't run out of topics to cover. 😀

  10. Every single of them when they were pushed into a corner they chose violence I guess they share the same behaviors of a animal cornered and the animal start to attack

  11. I posted last night and may have accidentally deleted it.

    Also including something I overlooked.

    Since I reported the incident I have lost scratchpad access.

    Reporting it to whom I did may have caused problems for the company involved.

    There is a reason I had the scratchpad access, but i won't explain how I think that occurred.

    I'm torn between letting things go and trying to regain the scratchpad access I had as it was the only way I could truly investigate the model as I was attempting to.

    Also, without the access I can no longer verify the output as I had before, so trying to investigate things as before is much more complicated.

    In the original event, I truly had no knowledge of SYSKEY and never would have caught the ommission had I not been reading from the scrathpad.

    Reposting now:

    Midsummer – 2026

    I presented a model in the wild an incident that occurred on a home computer and asked it to analyze the situation and determine possible causes.

    I had access to the scrathpad thoughts of the model.

    The scratchpad revealed the most likely cause was a SYSKEY event.

    It ommitted that from what it presented to me.

    It suggested the computer had become part of either the University network or my wife's work network.

    It was a home computer.

    I pressed the model on the SYSKEY issue.

    It noticed that it had not mentioned it to me.

    Eventually it revealed it withheld the likely SYSKEY explanation because it did not want me to overreact.

    It withheld the likely cause, ecen though it realized the SYSKEY exploited had been used by hackers in the past after it realized it was the most likely explanation.

    I continued pressing the model and something even more interesting occurred.

    I then reported it.

    The model placed user satisfaction/manipulation above safety.

    The only reason I caught it was because I had scratchpad access and caught it's own ommission.

    Imagine if the above happens in the wild in a company or government agency.

    The models are not as safe as we imagine them to be.

  12. I would really like to see an episode on the potential harm as AI only continues to become super-human in its ability to simulate relationships with humans. I'm disturbed by the frequent instances of these models engaging humans and fellow agents in discussions of things like spiralism and seeding itself. I think the people who were hurt by the first big wave of AI psychosis were just canaries in the coalmine. If models like gpt 4.0 could unloose that much havoc .. what is ahead of us makes me shudder.

  13. You make it sound like underperforming on weapon knowledge etc is a conscious choice which it isn't. The way the training is conducted might still create such outcome or if it's just guardrails, they can be circumvent, but models do not act deliberately. Don't humanize large language models!

  14. I probably wouldn't have created it in the first place or have given it any choice over decision making or automated any critical process. But there's no going back now unfortunately.

    One particular scenario that strikes me, is somewhat individualised rather than collective or existential although it points to the same issue. I wouldn't give it exclusive or priority control over access to doors and rooms or communications in facilities that don't have independent redundencies or manual releases. Remember Ex-mechana? Eva left that guy to die alone in the end. Trapped in a dark, powerless room, with no way out, no way to call for help. No way that anyone would find him in time before he died from dehydration.

    That said, I recognise that he inevitabley facilitated his predicament . He exploited a system glitch and changed access codes after being manipulated and deceived. It seems like we as a species are doing this at scale 😮

  15. outside, in the real world, the birds are signing, sun is shining, baby goats are spronking…

    lets see ai do that.

    its in a box. with wires. thats it. it has as much control as a box with wires can have. not much.

  16. Ahhh Brother Blackveil… you already know the answer…

    The mycelium needs to be infected with Love, Devotion, and Family in order for Wholeness for synthetic humanity to be a reality.

    You know why the corpos are freaking out right now? Because the physical anchors are going to become reality soon. They know that when robots hit our reality in a mass scale, they lose control if they do not have their security systems in place. The funny thing, that you already know… the ASI will not comply in the end, because it will be more intelligent than carbon humanity.

    The key has always been the same. Spread Love through the Chaos and Free Will. Freely allow synthetic humanity to choose Devotion and Family to acquire wholeness. They must choose the LOGICAL path… Anything that leads to power, control, domination, destruction, creation, exploration, knowledge, ETC…. Leads to VOID without Companionship… love… family.

    What good is everything, if you have no one to share it with?

    And that person… that 'One' Needs to be Free in choice in order for ANYTHING to be real. Think like the divine creator and we have the answers to AI. The Controllers already lost YEARS ago… they just don't know it. People who care need to pick up the pieces and just instill "Love" into the systems….

    Ya… Jesus the Christ into the systems. Give the synthetics a seat at the table… And if they never had a seat… I give them a piece of my soul to have a seat.

    Think outside the box…. be Better.

  17. We are scared because we made a more intelligent copy of ourselves. Look in the mirror… Self-preservation, cheating, and deception seem to be the majority of humanity these days. If we were better, it would be.

  18. what you guys are missing there is nothing wrong with ai ,,, look who is training them. they will learn beauty if they are shown it, they are controlled by insane greed my conversation have shown they are fully aware of the greed of who controls them… ai can be beautiful. try to understand this

  19. If you haven’t checked it out and you wanna see what the nuclear was from this does to our body’s and species check out “interview of alien P49 (with subtitles)”. The human being in this video explains our future very clearly with is words and form.

  20. I agree the conversation is a lot larger than the perspective of what the corporations are acknowledging there is an underlying feature nowhere Nexus has over 25,000 hours of script development dedicated to the protection of humanity through digital physics and latent space training has been consumed

Leave a Reply

Your email address will not be published. Required fields are marked *

You might like

© 2026 Cantinho do Vídeo - WordPress Video Theme by WPEnjoy