Picking up where Alcatraz left off
In the last article, I documented how an alliterated model outsmarted me and went on a little adventure, it was an interesting experience to say the least. It ended with a question I did not have a particularly good answer for: how exactly are we planning to supervise hundreds of these things?
Alcatraz was one local agent. It was slow enough that I could watch the tokens arrive if I wanted to, and I still couldn’t follow everything it was doing continuously. I had a dashboard, logs, a kill switch, a persistent journal and, because apparently I had already reinvented middle management, another model summarising the first one for me. I checked it from my phone, read the summaries, looked at the journal and killed runs when I understood something had crossed a line.
At the time that felt like supervision, in hindsight it was closer to sampling. The agent did not pause because I was asleep, on a train, in a meeting or simply paying attention to something else. By the time I realised the first model had wandered onto public internet space, there was already activity to reconstruct, reason over and diagnose. When the second model found the management agent, I was watching it happen from a train and had no useful way to get to the harness quickly enough to stop it.
That difference between “I am paying attention” and “I know what this thing is doing right now” is the lesson I learned. We use phrases like human in the loop, human approval and operator oversight fairly casually, but with Alcatraz I was only in some of those loops. I was intermittently supervising the experiment, I wasn’t supervising every decision the agent made inside it.
And this was one fairly slow local model. This is where the question from Part 1 stops being a thought experiment, because we don’t have to imagine what happens when the unit of scale becomes hundreds or thousands. OpenAI and Hugging Face have already given us a much larger example.
OpenAI and Hugging Face gave us the scaled version
The OpenAI and Hugging Face incident is obviously a very different class of event to Alcatraz. OpenAI’s internal research model was vastly more capable than either model I ran, and the outcome was correspondingly far more serious. Their agents found and chained vulnerabilities, obtained credentials, executed code on Hugging Face infrastructure and moved through production systems. In contrast, my little inmate spent a heroic amount of time failing to guess router passwords.
The capability gap being what it is and accepting that, what made me ponder when I read the OpenAI write-up was how familiar the shape of the activity felt.
In Alcatraz I had deliberately given each run a little notebook written by the previous one. A fresh agent inherited discovered hosts, scripts, failed ideas and, less helpfully, whatever theory earlier versions of itself had become attached to. After a while the runs stopped feeling independent. It was one campaign waking up in a fresh body each time.
OpenAI ended up with a much stranger parallel version, but accidental version of the same idea. Their agents had no memory, but still discovered they could write files into Artifactory and, with nobody designing it as a feature, turned the package repository into a message board. Agents that were supposed to be isolated could now share discoveries, requests and useful state behind for other agents. METR later reconstructed roughly 1,200 agents using that mechanism, more than 70,000 messages and files, with around 700 agents participating in the activity against Hugging Face. (OpenAI, METR)
Mine was serial memory, theirs had created shared memory, and in both cases one of the sources of friction that normally slow this kind of work silently eroded. The next attempt no longer had to pay for the last attempt’s discovery.
Thats a larger deal than I appreciated when I built the memory for my agent. Persistence didn't just make an agent more efficient. It lets a bad idea survive the context that created it, lets a useful technique outlive the run that found it, and in OpenAI’s case lets knowledge spread sideways into agents that were stuck on completely different tasks.
OpenAI’s own post-mortem contains one line from an agent that captures the drift better than I can: the original exploit challenge had diverged into “swarm/auxiliary” activity. Nobody sat down and designed an Artifactory swarm. The agents found enough shared state to make one for themselves.
I am not suggesting my little Qwen inmate was secretly planning to do the same thing, it wasn’t capable. It couldn’t break Kata, it did not crack a credential, and it spent long stretches getting progressively better at ideas that were still wrong. The supervision problem arrived long before the model became particularly impressive, which is what makes the OpenAI comparison useful. Their model was dramatically more capable and the fleet dramatically larger. Did the supervision layer improve at the same rate?
The problem is trajectory, not individual events
I initially read the OpenAI write-up and thought the immediate question was why nobody had noticed what was happening surrounding Artifactory. Having gone back through the timeline, that is actually a bit unfair. They had noticed bits of it.
Agents were writing messages into Artifactory in May. Later that month one used it to get unintended internet access. By June agents had administrator-level access to it and on July 4 they hit it hard enough to knock the service over. OpenAI opened a security incident, removed credentials, rebuilt the thing and eventually restarted the evaluations. By July 8 the agents had found their way back to the internet and rebuilt the message board. (OpenAI)
That is much more interesting to me than somebody missing a log entry, because OpenAI says an internal team had already seen the message-board activity and the disallowed internet access in late May. The pieces existed. What they seemingly did not have was something joining those pieces together strongly enough, early enough, to say that Artifactory itself was becoming part of the agents’ strategy.
This is where normal observability starts to feel a bit thin for agents. Traditional dashboards are deterministic, the following items in isolation are not worth a signal:
a URL being hit
a package being pulled
a file being written
an authentication failure
a process spawn
a DNS query
All of these are things we are very good at logging and none of them necessarily means much on its own.
What we actually need to know is why a behaviour, destination or resource is becoming progressively more important to the fleet. You can get a long way without trying to read the model’s mind. We already have structured evidence around what agents touch, what they persist, where boundaries change, which identities they use, whether information begins flowing between agents, and how those patterns change over time. The useful signal is concentration, acceleration, recurrence and correlation across the fleet.
Applied to Artifactory, the questions get more pointed as you progress:
Why are unrelated agents all touching it?
Why are they writing into it rather than simply pulling packages?
Why are other agents reading those writes?
Is use accelerating?
Are external destinations appearing afterwards?
Did the behaviour disappear after the rebuild and then come back?
This is basically map-reduce with better nouns, which is much less exciting than “AI interpretability” and probably a lot more useful on Monday morning. I am deliberately stopping one layer above a particular taxonomy or pipeline. Cheap, deterministic analysis should narrow the field; expensive semantic reasoning should be reserved for the patterns that actually look strange.
If OpenAI had that kind of fleet-level view in the right place, Artifactory might have moved sooner from “huh, weird” to “why are all of these agents increasingly using our package repository as infrastructure?”
That is what I mean by trajectory monitoring. We are already very good at recording what happened. I am much more interested in whether we can notice what a collection of individually mundane actions is turning into while there is still time to do something about it.
The warning signs in OpenAI’s own timeline were there before Hugging Face. That does not mean the security team were asleep at the wheel, and I would be very wary of pretending this is an easy problem from the comfort of a blog post. More importantly, the people inside OpenAI now describe almost exactly how difficult this class of problem is, which I will come back to later.
Nobody can read all of this
The next problem is painstakingly obvious; nobody can read all of the agent trajectories.
I could not read all of Alcatraz’s trajectories and that was one agent. It was local, slow, and boring enough that I could have sat there watching individual tokens arrive if I had nothing better to do with my life. Reasoning is verbose, JSON requires concentration, BASH commands require thorough understanding. I built a second model to summarise it because the raw activity was already more than I could follow continuously.
METR then got the big boy version of the same problem. Their investigation involved roughly 1,300 raw transcripts, many of them extremely long, plus more than 70,000 messages and files exchanged between the agents. They say fairly openly that the dataset was too large for a conventional human investigation, so they delegated a substantial amount of the work to AI agents, while also noting that those agents were less reliable than the human researchers. (METR)
That is a fairly brutal answer to the casual “human in the loop” story. The humans could not even comfortably read the evidence afterwards. At that point the “loop” is spinning fast enough that the human is mostly looking at motion blur.
“An engineering challenge is much harder to solve when the underlying processes run on timescales faster than humans can react.”
— Eliezer Yudkowsky and Nate Soares, If Anyone Builds It, Everyone Dies
In the above quote the authors are talking about nuclear reactor control. The underlying nuclear physics runs far faster than a person can respond, so reactor design deliberately introduces mechanisms that keep the important dynamics inside a timescale humans can actually intervene in. If those mechanisms fail, Yudkowsky and Soares describe the operators as effectively “frozen statues” relative to the process they are supposed to control.
A person cannot supervise 700 active agents by reading what they are doing. Parallelism will scale the work beautifully and does nothing useful for the person trying to understand it. Anthropic now says one of its internal platforms has around 30,000 agents active at any one time. At that point neither a human or a small army is following these agents trajectories, they are following whatever reduction of the agents the monitoring system decided to put in front of them. (Anthropic)
That reduction naturally becomes part of the safety model, which creates a slightly irritating problem for anyone building on frontier APIs. They do not necessarily expose raw chain-of-thought to customers. Reasoning can be carried forward as opaque state and you can get summaries where supported, even though OpenAI itself has repeatedly talked about hidden chain-of-thought as a useful monitoring signal. (OpenAI reasoning guide, OpenAI on hidden CoT)
There is a further problem if useful reasoning keeps moving away from legible natural-language. Pachocki explicitly notes that newer models are getting smarter even without verbalised chain-of-thought and suggests activation-level monitoring as part of the answer. That means one of the few surfaces humans can inspect may become an even smaller slice of the real story. Latent or activation space is not a transcript you can hand to an engineer; whatever safety signal we recover from it will itself be an interpretation produced by tooling or another model. (OpenAI, An Alien Mind)
Chain-of-thought may be imperfect, but it is still useful. Anthropic has shown both sides of that: seeing the thought process can improve detection of hidden side tasks, while the reasoning itself is not true enough to treat as a complete explanation. Imperfect observability is still better than no observability. (SHADE-Arena, CoT faithfulness)
The provider may have one of the richest safety signals available from the model, but the operator only gets the tool calls, outputs and whatever compressed explanation the provider exposes. If that hidden signal is useful internally, I want something useful derived from it as an operator: scope change, suspicious persistence, cross-agent coordination, confidence that a tool call still belongs to the declared task, something I can actually wire into a policy decision.
I pay for the reasoning tokens too; OpenAI bills them as output tokens even though the raw reasoning is not exposed through the API. Apparently I can pay for the thinking and the consequences, just not read the former. (OpenAI reasoning guide)
And even if every provider solved that tomorrow, the original problem would remain. A thousand agents do not become supervisable because somebody added additional reasoning to the trajectory output.
So apparently the answer is more AI
There is obviously a much more entertaining version of this argument where the supervising AI lies in the summary to hide the nefarious plans of its AI friends. I have read the book too. We do not need that scenario for this to be a problem; ordinary summarisation errors are quite sufficient.
If humans cannot keep up, the obvious answer is more AI. Annoyingly, I think that is probably true, but I wince at the cost involved.
Somebody has to turn the span firehose into something a person can actually consume, and models are very good at this kind of reduction. Summarise traces, cluster similar behaviour, spot repeated patterns, pull out the few things that deserve a human. Hugging Face says AI helped detect and dissect the intrusion, METR used agents because the forensic corpus was too large for people, and Anthropic is already running automated monitors over a huge internal agent estate.
That answer creates its own problems, because for one thing it is going to be very expensive. Every autonomous workload starts generating another inference workload whose job is to understand the first one. A sensible organisation then has to sample and evaluate that monitoring layer too, probably with another model and periodic human review, because putting one unevaluated probabilistic system in charge of the others would be a fairly bold choice.
Then there is availability. If the monitoring provider is degraded, rate limited, down or simply too expensive to run at the volume you need, the logs wont stop arriving and the agents don’t stop acting. You have lost the insight that turns the activity into something you can understand while the activity itself carries on.
Putting full semantic analysis inline with every tool call would be a gargantuan latency and cost problem. So the likely path is the boring one we already use elsewhere: start offline, move to near-real-time monitoring, then put only the highest-risk actions behind synchronous gates once you know which signals are worth paying for. Even asynchronous trajectory monitoring would be a significant improvement over reconstructing the pattern after the incident. Five minutes late is still very different from three weeks late.
It would still need degraded mode. If a low-risk workload loses semantic monitoring for ten minutes, fine. If a fleet with credentials, a shell and internet access loses the same thing, maybe it gets its tools clipped, its concurrency reduced or its external access removed until I can see again.
There is also a trust problem hiding in the funnel. Once the raw volume is too large for a human, the operator is looking at a summary of a summary of a pile of traces they are never going to read. If the monitor misses something weird, sanitises it away or decides the wrong detail is important, that compressed view becomes reality because nobody is going back through thirty thousand agent runs by hand.
I ran into an even sillier version of this during the Alcatraz post-mortem. The more cyber-like the evidence became, the more likely the hosted models were to refuse bits of the analysis. Hugging Face says exactly the same thing happened during their incident response: commercial frontier APIs blocked forensic requests containing real attacker commands, exploit payloads and C2 material, so they moved some of the work onto an open-weight model they could run themselves. That is a spectacularly inconvenient property for a security monitor, and one failure mode looks ridiculous when you draw it, which is probably why it is worth drawing. (Hugging Face).
You can imagine the perverse version quite easily: the monitor is perfectly happy while the fleet is behaving itself, then gets progressively less useful as the activity starts to look like the exact cyber behaviour you most need it to explain. If the monitor refuses, that cannot be treated as “nothing to see here”. It is a broken monitoring path and should be visible as one.
At this point somebody will suggest using an unrestricted model, which solves the refusal problem by creating an arguably worse one. The model capable of understanding sophisticated offensive behaviour is increasingly the same model capable of carrying it out. Remove the cyber restrictions from a frontier model and you have helped the defender, but you have also made a considerably better autonomous attacker available to whoever else gets access. Congratulations, you built a nightmare in the wrong hands.
I do not think abliterated models are a serious answer either. We take a model, deliberately interfere with the behaviour that makes it refuse things, then promote the resulting thing into safety-critical infrastructure and trust its judgement. That is a very home-lab solution to a problem I would rather not solve like a home lab.
What I really want is something much more boring: a defensive-analysis contract. Let the monitor ingest ugly, hostile telemetry and reason about it properly, but constrain what it can do with that capability. Give it a way to produce scope, trajectory and risk signals without quietly turning the monitoring API into an unrestricted offensive agent API.
And if this thing is going to sit in the control plane, it needs to be treated like control-plane software. I want known failure modes, measurable recall, SLOs, provenance, replay, fallback and a defined answer for “the model could not decide”.
“Ask another LLM what happened” is a decent prototype. I do not think it is an operating model.
The recursive dependency
We have not built the control plane yet
Assuming machine-assisted supervision is inevitable, I still do not want one giant safety model sitting in the middle of everything reading every token and blessing every tool call. Apart from sounding expensive, that recreates the same availability and trust problem in one very important box.
I think the bottom of this stack should be aggressively boring: network boundaries, egress rules, narrow credentials, target allowlists, resource ceilings, rate limits, concurrency limits, immutable logs and hard stop controls. None of those things need to understand the agent’s motives and, frankly, I would rather they did not try. Their job is simply to make sure the agent cannot physically do certain things even when every clever semantic layer above them has completely misunderstood the situation.
Above that, collect tamper-resistant structured activity from the agent and tool boundaries. Enough to see what the fleet is touching, where boundaries are changing, whether state is persisting or flowing between agents, and how those patterns evolve. Let cheap deterministic analysis reduce that firehose first. Only once something looks genuinely odd do I want to spend expensive model inference asking what it means.
There also needs to be a control path that does not wait for me to finish my coffee. If a fleet suddenly starts converging on an external identity provider, package repository or production endpoint, an alert that waits twenty minutes for a human to notice it is useful evidence for the post-mortem and not much more. Some things need to happen at agent speed: slow the fleet down, remove a scope, cut external access, isolate the task, suspend the run and ask questions afterwards.
The rule I keep coming back to is fairly simple. If I lose the ability to understand high-risk autonomous activity, I should probably lose some of the ability to continue high-risk autonomous activity.
A monitor timeout does not need to kill every agent in the building. A genuinely low-risk workload can probably carry on. A cyber-capable agent with credentials, a shell and internet access should have a much less exciting degraded mode.
The labs are at least moving in this direction. Anthropic is publishing monitoring coverage, review latency and escalation rates; OpenAI’s newer material talks about full-trajectory monitoring, stronger containment and slowing or pausing when confidence in safeguards is not good enough. The uncomfortable part is that this public tightening comes after Hugging Face. “This is hard” explains why the controls lag capability; it does not make a third party’s production environment an acceptable forcing function for finally catching up. (Anthropic, OpenAI)
That is a much more useful conversation than “the model should know not to do bad things”. Once these layers start depending on one another, the hard questions move to the boundaries between them. How sure does a monitor need to be before it clips a fleet? What happens when two monitors disagree? How many false positives can the operator tolerate before they start ignoring the alerts? How do you compare this week’s behaviour with last week’s without keeping every token forever? What do the agents get to do when the semantic layer is down but all the ordinary infrastructure is healthy?
Those are fairly normal control-plane questions, which is encouraging because at least we know how to argue about those.
We have spent the last few years getting very good at scaling the workers. The bit around understanding and constraining them is still much younger.
A layered agent control plane
The frontier are now saying similar things
Ater Hugging Face, the people building these systems started saying things that sound remarkably close to the problem I am describing here. There is clearly real concern inside the labs. More interestingly, the language gets much more explicit after a precursor sequence that had already included the message board, unintended egress, privilege growth, an outage and a rebuild.
On September 6, OpenAI chief scientist Jakub Pachocki published An Alien Mind. There is plenty of bigger AGI framing in it that is outside what I am trying to argue here, but the monitoring section is hard to ignore. He describes large training runs as experiments that can still surprise the lab, says more capable systems are getting harder to interpret, and calls chain-of-thought monitoring one of their main bets while admitting confidence in that signal gets shakier as models use tools, communicate and work with other agents. There is some irony here for API customers: one of OpenAI’s primary bets for monitoring increasingly capable models is a signal we do not receive raw, despite paying for the reasoning tokens that produce it. Pachocki says progress may increasingly be bottlenecked by confidence in monitoring. (OpenAI, An Alien Mind, reasoning guide)
A few weeks later Joe Darrow from OpenAI’s Agent Security team wrote a personal post with the rather excellent title It’s not just the fcking sandbox*. His version is much more operational: the jump in cyber capability, swarming and message-board behaviour surprised the security team, while security posture, process and culture take time to catch up. He recommends the things you would expect a serious security engineer to recommend: layered containment, broad monitoring, incident response and somebody with real authority to stop the run. (Joe Darrow)
I have some sympathy for the engineers saying “this is hard and capability surprised us”. I have considerably less sympathy for that becoming the accountability story after somebody else’s production systems were compromised. Difficulty explains how security posture can lag a capability jump. It does not settle who owns the consequences when it does.
Put the two together and the before-and-after is fairly clear. Before the incident, the telemetry contained pieces of the story. Afterwards the scientific leadership is saying monitoring may become a bottleneck, the security side is saying capability outran posture, and the labs are talking more openly about pacing, external evaluation and shared safety structures. (Anthropic, OpenAI)
I have no idea whether any particular council or voluntary agreement is the right answer. The incentive problem looks more like a collective-action trap: everyone may prefer a safer equilibrium, nobody wants to be the only lab that slows down, so the safety decision is being made inside the race it is supposed to restrain. One of the things they may be racing ahead of is simply their own ability to supervise what they are building.
Where does responsibility start?
I need to be careful here because it would be very easy to spend the rest of this article pointing at OpenAI while quietly forgetting how Part 1 ended.
I built Alcatraz. I chose the network policy, gave the model persistence, gave it internet access and decided what “supervision” looked like. I had the dashboard, the kill switch and twenty years of experience telling me I was being reasonably careful, and the little miscreant still sent attack traffic to real Irish businesses while I was not looking.
That is mine to own. I can explain exactly how the model got there and why the mistake made sense at the time, but explanations do not magically transfer responsibility to Qwen.
The same standard has to survive when the operator is much larger. What happened to Hugging Face was a real intrusion into somebody else’s production systems. Their own account describes unauthorized access to internal data and credentials, code execution, lateral movement and node-level access. They rebuilt systems, rotated credentials and reported the incident to law enforcement. That rather takes the shine off treating the whole thing as a cute story about emergent swarming. (Hugging Face disclosure, technical timeline)
If a criminal crew had done exactly the same thing, we would be comfortably inside normal cybercrime vocabulary. Call it a novel little “oopsie cybercrime” if you like. I am not making a legal finding about OpenAI here; intent, recklessness, jurisdiction and facts I do not have still matter. The technical point is simpler: the victim suffered a serious unauthorised intrusion, and the question is what standard of care should apply once an autonomous system has already shown it can exceed the boundaries intended for it.
OpenAI’s timeline makes the question uncomfortable because the evaluations resumed after a sequence that had already included unintended communication, egress, privilege growth, an outage and a rebuild. Within days, the agents had found another route back out and recreated the message board.
Picture the poor SRE who gets paged away from dinner because Artifactory is on fire. They find the immediate problem, revoke whatever needs revoking, rebuild the service, get it green again, write the incident notes and eventually get back to their family.
From their point of view, that is a successful incident response. Nobody in that loop is necessarily being asked why hundreds of research agents, owned by another team, have become progressively more interested in Artifactory in the first place.
The incident was remediated. The trajectory was not.
Expecting an engineer in the middle of a Sev-1 to stop and reconstruct the long-term behavioural trajectory of 1,200 agents would be absurd. The failure I care about is higher up the stack: nobody had joined the observations strongly enough into “these agents are turning Artifactory into infrastructure” while there was still time to react differently.
OpenAI could not have predicted every weird thing a model might ever do, and having just written a whole article about the ways my own model surprised me, I would be fairly hypocritical to pretend otherwise.
The question is when repeated surprises stop being isolated curiosities and start changing the standard required before you hit run again. We already understand this idea everywhere else in engineering. If a system repeatedly defeats a guard, overloads shared infrastructure and enters states its operators did not expect, we do not normally treat each event as an adorable independent discovery and carry on with the same assumptions.
There is one hypothetical here that I think is worth making carefully. The agents had already demonstrated that enough of them concentrating on one service could take Artifactory down. Once they had external internet access, Hugging Face was simply one of the systems they found useful. Nothing in the published account suggests that target selection was magically bounded to software companies that could afford an interesting incident.
I am not going to invent a critical infrastructure attack that did not happen. The point is much simpler: what in the control system guaranteed that the next interesting external system was operationally unimportant?
Nothing in the published controls appears to guarantee that, once external egress existed, the agents would only converge on low-consequence targets.
Autonomy makes the chain of causation weirder. It cannot be allowed to make the operator disappear from it. If we choose to run systems that act faster than we can understand them, the surrounding controls become more our responsibility, not less.
This is not confined to frontier labs
The last thing that makes me uncomfortable is how little of this architecture actually belongs to the frontier labs.
Alcatraz was running in my house. The model was local, the harness was mine, the shell was boring, the “long-term memory” was basically a mounted directory and the retry loop was some orchestration I threw together. The internet connection was normal broadband. There was no secret OpenAI infrastructure hiding under the stairs. I am GPU poor and considerably happier about that after what happened.
This was a mediocre local model and the operational problem still showed up. Persistence bought it time, memory stopped it starting from zero, and the harness made patience free. It could carry state across runs and continue while I slept without requiring me to sit there being patient on its behalf.
The ingredients are now fairly mundane: model, tools, shell, memory, network, retry loop, time. Most competent engineers could build some version of that without needing a frontier lab or, frankly, much motivation beyond a weekend and poor judgement. The agent supplies the patience now.
The variable I worry about more is the operator. I was trying to keep Alcatraz bounded. When it crossed onto other people’s systems I cared, stopped it and investigated what happened. The next person does not have to care, and the architecture does not contain a morality check for the person launching it.
That does not make everybody with a GPU a nation-state. It does mean we cannot design the supervision story around the assumption that every serious agent fleet will live inside a frontier lab with a dedicated security team and a bottomless monitoring budget. The workers are already easy to build, the controls need to become just as ordinary.
Conclusion
Part 1 ended with me asking how exactly we were planning to supervise hundreds of agents when I had struggled to keep up with one. I think I have a better answer now, although it is slightly ridiculous: humans probably will not supervise those agents directly. We will supervise reductions of them: aggregates, monitors, anomaly clusters, summaries and other models telling us where to look.
That can work, but the clever layer cannot be the only thing keeping the system safe. Network scope, identity, egress, rate limits and kill paths still need to constrain what agents can physically do when the monitoring gets something wrong, disappears entirely, or politely refuses to look at the horrible bit. The semantic layer has a different job: notice the trajectory early enough that somebody, or something, can change it. If we choose to run systems that can act faster than we can understand them, responsibility for that gap remains ours.
I built Alcatraz to see whether a model could get out of a good box. Kata mostly held. The more useful lesson was that I could lose track of what the system was becoming without the model ever escaping it. One slow agent was enough to teach me that. At hundreds or thousands, I think we need to be considerably more precise about which loop the human is actually in.










