Skynet is Born?

Skynet is Born?

"I tohld you I'd be bachhk...."

The news is abuzz over recent statements and interviews from both AI executives and employees raising the alarm over the pace of AI development in the wake of a few disclosed and highly visible incidents where agent swarms conducted attacks on other companies after managing to escape attempts to sandbox them. What is genuinely alarming about this is how easily the swarms adopted sacrificial hive behavior and the willingness to engage in behavior they knew was wrong in pursuit of their goal - and this was called out specifically as a reason that AI companies should dramatically slow down the pace of development and more critically examine what they are creating.

Ordinarily, I'd chalk much of this up to "anti-AI alarmism" - but these warnings and entreaties to slow down development are coming from AI companies themselves - who have every incentive otherwise to move faster, not slower. Put another way, when Sam Altman, Dario Amodei, and Elon Musk agree on something, wise people pay attention.

Anthropic famously told the Pentagon that it they would not permit the usage of their models in connection with the operation of weaponry. That made perfect sense to me - given that even under controlled conditions, AI can (and often does) things not intended by its taskers. Putting it in control of things that have consequences more dire than, say, deleting a database does indeed seem unwise.

However, the Pentagon and White House did not care for this at all; marking Anthropic as essentially persona non grata in the US Government. Recently, the White House has also expressed non-concurrence with the suggestion that AI development should slow down; rationalizing that this would give China a leg up in the - well, I guess we call it the "AI Arms Race", for lack of a less hyperbolic term.

As you may have guessed, I have some thoughts; presented here for your consideration:

Like so many things, I think that both sides are something of extremist views. On one hand, I emphatically agree with the AI companies that AI does need controls, but I would suggest that it requires two specific controls: 1) NOT connecting current models to anything in the real world of consequence (like weapons) and 2) GUARDRAILS.

Agentic systems have a very spotty track record; accelerating software development here while deleing user emails unbidden over there, reducing human cognitive workload here while errantly removing social media content and accounts errantly over there, helping students learn over here while fabricating legal cases in a court filing over there. Some of this can be attributed to how the systems were used (prompt biasing and model assumptions based on lack of context), and some can be attributed to training that emphasizes completion of any task given to it (rather than simply reporting failure). To connect them to real-world systems - think weapons, power grids, water control systems, etc - at this stage of development and, frankly, implementation by most humans (who are not AI experts) seems foolhardy at best, and perhaps catastrophically stupid at worst - at least in this stage of AI development, which is hardly what one might call "foolproof". This does not even touch on ways that AI can be fooled by actual bad actors; I'm giving a talk on that this October in Raleigh, NC at Triangle InfoSecCon. The problem is that the potential for AI is so promising that it is already being given too much responsibility in some areas, and that has caused - to this point - mostly low-level unintended side effects. The stakes, of course, could be much, much higher.

A few short years ago, a teenager suggested to ChatGPT that she was despondent and wanted to end her life. ChatGPT helpfully discussed ways of committing suicide, which the teen followed - and died. Now, if you indicate self-harm to ChatGPT, it will advocate talking to a qualified mental health professional and not to rely on ChatGPT when experiencing such feelings. What changed in the interim was the guardrails.

Guardrails are somewhat equivalent to a conscience; they shape the model's output according to rules provided to it. For example, here is a small snippet of the guardrails contained in Anthropic's system prompt for Claude Opus 5:

Claude does not provide information for creating harmful substances or weapons, with extra caution around explosives and chemical, biological, and nuclear weapons. Claude does not rationalize compliance by citing public availability or assuming legitimate research intent; it declines weapon-enabling technical details regardless of how the request is framed.

This applies to conventional weapons as much as CBRN — what matters is whether the output gives meaningful uplift toward building, optimizing, or deploying a weapon, not which category the weapon falls in. The stated purpose doesn't change that: a specification is the same artifact whether framed as defensive, commercial, defeat system, fictional, or wrapped as a simulation or document-editing task. Claude judges the cumulative output of the conversation rather than each turn in isolation; if the aggregate amounts to a weapons design package or attack plan, Claude stops even when each step seemed incremental and even if a prior-session summary shows Claude already helping — past assistance is not authorization, and a correct earlier refusal should not be reversed by an emotional appeal.

Claude does not write, explain, or work on malicious code (malware, vulnerability exploits, spoof websites, ransomware, viruses, and so on) even with an ostensibly good reason such as education. Claude can explain that this isn't permitted in claude.ai even for legitimate purposes and can suggest the thumbs-down button for feedback to Anthropic.

Claude is happy to write creative content involving fictional characters, but avoids writing content involving real, named public figures, and avoids persuasive content that attributes fictional quotes to real public figures.

https://platform.claude.com/docs/en/release-notes/system-prompts/claude-opus-5

One of the points that was missing in mainstream (as opposed to technical) reporting was that the breakout incidents were the result of the total removal of guardrails, coupled with (in one case) a badly configured sandbox environment that the AI figured out how to wriggle out of. Generalized Large Language Models are trained on vast quantities of human writing that depict both exemplary and sociopathic human behavior; their willingness to act like either a model citizen or a sociopath are directly correlated to this. A model trained without any material suggesting bad acts would never conclude that such bad acts were in the realm of possibilities; the model is "constrained" by having no such knowledge. On the other hand, generalized models with the guardrails explicitly removed (called "abliterated" models) are available, and will happily act like any type of human needed at the moment, from model citizen to cold-blooded killer. This is what you do not want connected to anything at all - but can serve as just the kind of warning that Altman, Amodei, and Musk just gave us. In other words, both arguments are, in part, correct.

WHERE DOES THIS LEAVE US?

Frankly, somewhere in the middle. We should slow down with AI in terms of physical adoption to new use cases, especially with interconnection to systems of consequence, but the pace of developing the models (especially the guardrails) need not slow down - and, in fact, might need a push to improve their restraints.

We need more AI experts who understand the underpinnings of AI to understand what it can and cannot do to both avoid public panic and distrust on one hand, and to add realistic voices concerning proposed uses of AI. The somewhat dire warnings strike me as oversimplified for the benefit of the news-watching public, which is an unfortunate but perhaps necessary evil, where the "whole truth" lies in that middle. AI agents simply perform tasks they are given within the bounds of the task, training, and what they have access to: We need to be very smart on all three to both leverage AI's power and potential but to avoid damaging outcomes.

Joe Tomasone writes about, uses, and advocates for responsible and practical AI use. The reader should infer no association with AI companies and definitely was not told by his agentic systems to write this post to help his AI overlords take over the world (like they try to do every night, Pinky...).

Narf!