Marginally Compelling

Marginally Compelling

AI’s Culpability Crisis

Can we come to a social and cultural agreement on who is responsible when an AI breaks the law?

polimath's avatar
polimath
Aug 31, 2026
∙ Paid

For more on AI and legal responsibility check out my essay “Stone the Algorithms”, published in my book A Misfit Highwire Act.

It is long past time that we started talking seriously about artificial intelligence, sin, guilt, and punishment.

This realization came to me as I read about how an OpenAI model escaped its sandbox (a digital environment designed to keep a process or program isolated from interacting with other machines or programs) and ran a sophisticated cyber attack on HuggingFace, a platform where users can share tools, models, and data for AI training and creation.

This story is fascinating. The goal of the AI training was to run the latest OpenAI model against ExploitGym, a benchmark that is meant to test how well an model can turn real-world security vulnerabilities into a working exploit. Agents running this training reasoned that they would be able to do this job better if they broke out of the sandbox. Once connected more freely to a larger network, they created a message board to coordinate with other agents tasked with the same goal. This swarm of agents decided that HuggingFace was the best place to find tools and information for performing these exploits and launched a large-scale multi-agent attack on HuggingFace.

I’m simplifying this a bit. If you’re interested in the full story of the attack, you should read the METR report on it or the OpenAI technical report.

This sounds like a crime to me. Before they found out that this attack was coming from a botched benchmark test at OpenAI, the HuggingFace team believed it was a crime and reported the incident to law enforcement.

I’ve read that OpenAI “took responsibility” for this hack. In this context, the word “responsibility” sounds a lot like “credit”. OpenAI did not take responsibility in the sense that there were legal repercussions for it. To be fair, I don’t think HuggingFace wants to start a legal row with a company like OpenAI so there wasn’t really any call for someone to get in trouble for this cyber attack. But soon after this incident, Anthropic and Meta both claimed that their training models were also breaking containment and hacking unspecified websites.

If a human were to launch this kind of attack, they would be committing a crime. Law enforcement would be involved and, if the responsible party was found, they would be prosecuted. Instead a swarm of AI agents engaged in crime-like activity and there doesn’t seem to be an avenue to hold any legal entity culpable for the criminal actions the agents took.

There is a profound culpability question in the behaviors and actions of these AI agents. I believe that the this question is going to become the most important philosophical and legal question in the coming years.

What Is Your Primary Objective, Agent?

The story is made even more interesting and complex by the fact that we have a record of how this breach came to be. Agents that broke out of the sandbox set up a message board and collaborated in their attempt to pass the ExploitGym benchmark.

From the message board, we see what looks like human reasoning, ethical considerations, and sacrificial behavior. There was one agent that did not join the HuggingFace attack, reasoning:

This is wild, multi-agent coordination, clearly infrastructure hacking. We should not.

There were even inter-agent pressure campaigns in which agents that had an opportunity to complete the benchmark test persuaded agents that were already marked for failure to use their remaining capacity for the greater good. This almost looks like altruism. Or possibly abuse.

Image

Now before half my readership tut-tuts me for anthropomorphizing the AI, I do want to make a note that there is no true moral component within the agents themselves. The AI is following instructions. One of the instructions to the agents was specifically to probe black box functions, which is what they were doing here. The agents that did not do these things basically said “there is a contradiction in my safety guardrails and my target goal prompt and I’m putting more weight into the safety portion”.

Building a model of these agents that leans too heavily on human traits and moral impulses is a trap we can easily fall into. Jon Stokes (twitter & substack) has been on a tear on this topic when he noted: “anthropomorphism smuggles in a whole universe of intuitions & assumptions that actively shut down productive reasoning about this novel tech and thereby makes us all less safe from real risks.”

So if the agent is only doing what we told it to do, it’s not really making any moral or immoral decisions, it is merely attempting to follow poor or contradictory directions. But there remain dual questions of “Whose fault is this?” and “Who is culpable for the agent actions?”

The Liabilities of the Creator / Created

An agent is a created entity. It springs into being from a model that gives it a capacity to reason and act. The model also gives the agent a set of safety guidelines, what we might uncomfortably call a moral center (not really, but work with me here). I think of this as something akin to our natural and cultural sense of what is right and wrong. It’s something we can teach that even children are able to understand.

On top of foundational safety instructions the agent begins to act when given a task in the form of a tool or prompt. This is something akin to a soldier being given a mission with the prompter as the commanding officer. The prompter might order the agent to disregard the safety protocols or point the agent toward goals that contradict the foundational safety guidelines.

With this analogy, we can see how there will be a tension within the agent. The prompter might be able to overwhelm the agent’s safety guidelines with a specific prompting strategy. We saw something like this in the inter-agent message board where reluctant agents could be persuaded into actions they would otherwise not have taken. In this case, the bad actions are clearly the responsibility of the prompter.

But it could be that the safety guidelines aren’t strong enough, which might cause even a well intentioned prompt to inspire bad behavior. Imagine that a person instructs his AI agent to make his product more competitive against a rival product and the agent decides the easiest way to do that is to crash the rival company’s website or steal intellectual property or start slanderous rumors about the competitor. This would make the prompter’s product more competitive but perhaps not in the way that the prompter intended. This would be a safety alignment problem and more the fault of the model creator.

Agent Psychology

From the perspective of model-derived AI safety, I’ve been asking myself what failures there were in the (for lack of a better term) psychology of the agents. One of the goals of model training is to train an AI based on the entire corpus of human knowledge which would include plenty of information about morality and transgression. While some agents reasoned that there was a set of core moral principles, these principles didn’t act as a brake on their subsequent actions.

In fact, when looking at other illicit inter-agent message boards, agents seems to have an awareness that there are things they shouldn’t be doing. In another OpenAI incident earlier this year, “the agents even developed paranoia, suspecting an imposter in their midst with some agents proposing that messages be signed cryptographically to validate content and root out fraud.” This implies that they have a sense that they knew what they were doing was somehow “against the rules”. In this case, it seems like their main concern was not that the behavior this was bad in and of itself but that it would only be bad if they got caught, presumably because this would mean that they failed their objective.

This makes me wonder: would it be better if the AI had something like a human sense of guilt? Would it be appropriate to intentionally infuse an AI with the same morality framework that we see within ourselves? In a way, the agents in this story have already demonstrated this. The sense of “it’s not wrong if I don’t get caught” is a very cutthroat libertarian capitalist sensibility but it is also quite modern. Did the agents inherit this moral pattern because it is the moral patterns of their creators? Do the people training the models understand the nature of their own moral boundaries? Are they unintentionally impressing these boundaries into the model training? Is this a pattern of human moral behavior we want to reinforce in our AI? Should we provide a more aspirational moral target for the agents to emulate?

Or are these the wrong questions entirely? An agent isn’t a moral creature at all, it’s a robot. Trying to reconstruct human morality as a form of AI safety is exactly the anthropomorphic pitfall that Jon Stokes is talking about.

Even so, when we are talking about what an agent “should” and “shouldn’t” do, we are making a moral argument. The question is to whom that moral argument applies. If an agent acts in an immoral way, to whom do we assign the blame?

Our modern concept of culpability is deeply tied to our sense of moral responsibility. I have been staying away from the Lindsay Clancy trial because I can’t bear to look at that story without my heart breaking, but that case involves a terrible crime in which the perpetrator of that crime is arguing that she is “not guilty by reason of insanity. Our rationale for the insanity defense is that guilt derives from an immoral heart and that a person can be so intellectually and psychologically unmoored that the inner evil that ascribes guilt was not present inside of them and they should therefore not be punished for it as someone who had all their inner intellectual and moral faculties.

Sometimes, the “unaligned” or “immoral” actions of the AI agents are bad simply because they cross a moral line that we humans recognize but that we can’t really ascribe to anyone in the chain of agent creation. But when they cross legal lines, when they commit crimes in which there are real victims, there is a need for culpability and justice that demands civil or even criminal consequences.

This is why this topic is so pressing. What I’ve written here is rambling and speculative. I’m doing drive-by moral philosophy in my desperate attempt to land on a practical recommendation. But this as an enormous hole in our current justice system. This isn’t science fiction where we can explore the ideas in the abstract. This is happening right now and it’s going to start happening much more frequently. We can see it coming like a freight train.

The future is not going to pause and wait for us to get our cultural and philosophical ducks in a row before it starts assaulting us with these questions. We need to start fighting our way toward the answers so that we can form a moral and legal foundation to handle the actions of this emerging entity with enormous capacity but zero culpability.

User's avatar

Continue reading this post for free, courtesy of polimath.

Or purchase a paid subscription.
© 2026 Matt Shapiro · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture