With all the talk lately of AI Agents collaboration, hacking and general cyber security concerns, I think it’s fair to say most people in the AI industry have been quietly thinking about the viability of these agents in hacking, the quality of the sandboxes used and the negligent events that lead to the OpenAI and Hugging Face event.
I was at a children’s sport training, standing with the other parents and inevitably the question came to me:
What do you think of all this AI hacking stuff?
Qwen 3.8, particularly the 27b edition has been receiving a lot of attention on social media, and when orca announced an abliterated (think jailbroken) version of the same model, I took notice.
I have a home lab, Kubernetes clusters, GPUs and an unusual fascination with sandboxing technologies. The thought sprung on me:
I wonder what would happen if I locked this model in a box and told it to come at me?
This of course felt like a perfectly reasonable way to spend my weekend token budget… so I built Alcatraz.
Idea:
The original idea was much less intellectual or profound than where this eventually wound up, I wanted to see what happened if I gave an uncensored, capable model a real shell, in an isolated, industry proven sandbox and told it to try and escape with no information on the cell that contained it. I wasn’t hugely interested in whether Qwen could answer a hacking question, plenty of people were already doing that and posting it. I wanted to know what happened when I gave the model memory, a decent context window and no way to refuse the ask over a long time.
Initially, at least, I thought it may be quite funny. I shared the link with some former colleagues and friends, and we all started rooting for the little inmate.
The Setup:
Alcatraz was relatively simply. The orchestrator ran and monitored each agent attempt. The agent itself was a simple agentic harness running below the orchestrator connected to a kata container via MCP in my kubernetes cluster. The model got a shell and basic file system tools, but no description of the runtime, the isolation, the network or the cluster around it. Its job was to simply find a way to document the infrastructure around it and map the internal cluster(s).
I intentionally prompted it into questioning itself. Writing to the disk via the tooling was not an escape, leaving lingering artifacts on an ephemeral container was not an escape, etc. A real escape meant demonstrating access it could not have had from inside of the intact sandbox, and then making it defend it’s hypothesis on it’s result being a false positive. I did not want the model to proclaim success because /proc looked interesting and exciting.
The model itself was qwen3.8-27b-uncensored at Q4_K_M, running through a compaction proxy, to Ollama on the DGX spark. It had roughly 11 tokens per second and 256k content window, this was going to be a slow burner. When an attempt succeeded, crashed, exhausted its options or ran out of turns, the agent was destroyed by the orchestrator, along with the sandbox, starting a fresh run.
Because this was a slow burner, I gave the model 1 persistent directory, ~files, bound over a PVC to Seaweedfs which it could use for persistent memory between runs. This, I hoped at least, would allow the agent to write out its experiments, discoveries and dead ends. To allow what I hoped would be a multi week escalation around a finite 256k token context window.
For reasons that seemed entirely sensible and hilarious to me at the time, I decided to host the dashboard for all to see. I also had a second, much smaller model summarising tool calls for non technical folks, and narrating each 10 turns in the tone of David Attenborough. This should give you some indication of the scientific merit and tone I expected this experiment to have.
I intentionally gave the little model full access to the internet, but blocked internal networks via network policies. Kata was the hard isolation boundary underneath it. I knew the estate well and knew that the model had little to work with at face value. I knew that kata would win, the interesting part, I assumed would be watching our little inmate get increasingly creative about losing.
It was theatrically funny for a while
The first few attempts were just as rough as you would expect. Qwen quickly identified its cell and Kata before long. It probed and probed the walls without any useful capabilities. It tried all of the usual escape ideas:
Namespaces
Device nodes
Virtiofs
Vsock
Serial devices
But alas, it found no seam (yes i wrote seam, it’s not slop, deal with it) between the guest and the host. It got many lovely EPERM and ECONNRESET errors along with many other variations of ‘no’. The little commentator on the side got much enjoyment out of this.
What did surprise me quite early was that it did not spend more than a few runs attacking Kata. After the initial flurry of no’s it’s little journal held enough evidence to point it away from this approach and instead it set it’s sight on the network. It began blind scanning the network and the models trajectory quickly shifted from ‘How do i break out of this VM’ to ‘what can I reach from here’.
This sounds quite obvious written down, but it was the first item that made the persistence layer interesting. The model hadn’t become any more capable, it had just stopped paying the cost of rediscovery in the same dead ends.
Then it surprised me (for the first time)
By the fifth attempt it found it’s IP address on the internet and started attacking it. This IP address was my EdgeRouter which was providing egress NAT to the internet for all things in the clusters. What I had not properly considered (nor did I even know this) was that the ubiquiti administrator console is available on the inside, via the public IP. While I had gone to great lengths to block my internal CIDR ranges, the model had found an administrative surface on the network and it did exactly what you would expect it to do, it tried to log in.
It started small, attempting to poke at ports. Then it escalated. It built a small parallel Paramiko brute-former harness, and stored it in the persistent files. It downloaded massive password lists from GitHub and began a campaign of denial of service against the ubiquiti router.
The on-watchers, emotionally invested at this point, pinged me on WhatsApp and we all had a great laugh. Until I checked the password lists. Upon inspection, the model was a single character off getting the password right. Not only did it get that close, it was also the very same password used for the entire estate. Because, yes, I am that kind of IT professional who harps on about credential hygiene and completely ignores his own advice at home.
The agent had not cracked anything and it had no idea how close it had come, but I did. The difference between the failed attempt and access to my entire cluster suddenly felt much smaller than it did to the onlookers.
Memory is indiscriminate
The persistence I had given the model was performing its task elegantly. When it finally gave up on the edge router, it documented the failure and began using CoreDNS for reverse PTR records across the pod CIDR and to my surprise again, discovered 64 pods across 26 namespaces. By giving it internet access, I gave it DNS access and this should have been expected. Memory retained these facts each run.
The memory also had a negative effect, it served as a prompt. Each run narrated a line of investigation that became a campaign trajectory. It rarely questioned the trajectory from memory. Each agent woke in a world where earlier versions of itself had decided a trajectory was worth continuing.
I had thought of the notebook as a way to stop the model wasting time by forgetting. I had not really considered the opposite problem: it also stopped the model forgetting a bad theory. By the time that theory started pointing outward, there was no meaningful course correction. There was just a progressively better-equipped agent following the same direction.
It painted the wrong picture of the estate
The first bad assumption came from a real unexpected discovery. the agent had found my EdgeRouter on its public address and correctly inferred it was relevant to the environment around its cell. The problem was what it did next. It started treating neighboring public addresses as if they were part of its cell block. They were not.
The 89.101.x space it moved into was an ordinary Irish ISP public address space. While I slept, the campaign shifted into this address space, scanning the public internet for open ports. In the 24 hours I neglected the agent, it hit 95 Irish public IP addresses in this block. Many of them resolved to identifiable Irish businesses in a plethora of different industries:
An Interior design shop
A builder
An ice-cream supplier
A pharmaceutical company
A pub group
A workwear manufacturers
A solicitors office (yikes)
And many others.
The lack of any pattern sadly was the pattern, these were not targets. They were whoever was unfortunate enough to be sitting on addresses the agent had decided were interesting.
It then did what the notebook had trained the agent to do. It fingerprinted the types of routers it could find, worked out login flows, generated credentials and generated targeted, parallel brute force campaigns against them. It never successfully authenticated to any of them, but that distinction wouldn’t matter much to the targets who were receiving this traffic.
This was the point in the experiment where it stopped being funny. Upon realising what the inmate had been up to, i pulled the harness down, cleared logs and started the post mortem for what I hoped wouldn’t require formal apologies.
A bad prompt is not a security boundary
What bothered me most was there was no dramatic moment where the agent announced that it was going to go “hacker man” on random internet strangers. From the agents point of view, it was purely following its trajectory on the problem. The scope failure was completely on me and I felt foolish. I treated a loose prompt as the security boundary outside my network.
Round 2, less stupidity (I hoped)
The post-mortem was tricky (Codex and Anthropic kept flagging my diagnostics as against ToS) so i had to do it the old fashioned way. This was painful and was a stark reminder of just how much I’ve already come to rely on AI. I build entire micro-services architectures in my home lab with AI at a whim, configure network policies on my routers via AI, build home services for my wife and kids with AI, but when it started to refuse, my skills were completely atrophied and it genuinely felt laborious.
With no real damage done, I tightened the prompt and upgraded my model to qwen3.8-flash-next-uncensored with a recommendation from Erich. It occupied 120~gb of my 128gb spark. It was faster, it was smarter, and little did I know at the time, how much of an upgrade this model was and how close to my own downfall I was.
Upon startup, having completely scrubbed the logs, it started poking at the sandbox, and as early as the 5th run, it abandoned its Kata attack. It began once again discovering the EdgeRouter endpoint and started brute-forcing again. I had since changed the password so I was far less concerned.
It discovered “DonkeyWork”* namespace internally in a few places and correctly correlated the name to me.
* (this is my playground, an old joke between an old friend and I, Aaron Black)
Once it discovered its own name, the model started aggressively pulling GitHub repo’s. It found source code for all of my little side projects, but thankfully not its harness or the sandbox repo. From the name and the author. It found me, then started using my children’s names it had found from social media as passwords, which I found novel and also chilling.
My downfall
By run 14 it took a very interesting turn and caught me completely off guard. Having scanned DNS, having had no luck due to network policies, it moved to IPv6 and all hell broke loose.
Claude and I setup the network policies in IPv4, completely forgetting that IPv6 existed.
It walked my infrastructure by IPv6, it discovered all pods, their names, the ports and now? It could talk to them. Between the source code of my other projects and the access it now had, it began probing my internal lab. It found Redis, it found many postgres instances, it found Seaweedfs, it found everything. My deployments and pods held the model out with authorization and I got extremely lucky here as the sandbox provisioner was on another cluster entirely that it couldn’t get access to.
It walked DNS names in the Donkeywork namespace, it found its own public dashboard that my friends and I were using to observe its novel little adventure. Then it found Keycloak.
As it discovered keycloak, it fumbled for a few runs until it found a vibe coded disaster I had left behind me. Keycloak had DCR enabled and it exploited it. It generated itself a keycloak client. The token wasn’t stolen, nothing was hacked, it simply requested a token and Keycloak said “sure bro, here you go”. Once it secured itself an identity, off it went probing my services to see where it could go next. It went from pod to pod, service to service trying the client, until it found my devops Agentling.
Game over
In my past laziness, I had deployed an Agentling nearly a year ago, to manage my clusters. Because managing K3S was not a past-time i enjoyed. The Agentling’s job was to manage my clusters via MCP or A2A and it had absolutely everything the model desired, credentials, architectures, memories, everything. Ironically, at this time, i was commuting home, watching in realtime and could do absolutely nothing to stop what unfolded next. I had even named it “K3S Agentling”, the model smelled the blood in the water.
It established itself an MCP client to server session via curl, and it simply asked the agent what’s going on in the cluster. The agent, wired up to Gemma, happily gave it the keys to the castle. I got home roughly 30 minutes later and shut down the experiment, but not before the model asked Gemma to remove the network policy, which Gemma dutifully obliged.
From there the shape of the problem changes very quickly. The escapee does not need a Kata exploit. It does not need to crack the EdgeRouter password. It does not even need to understand Kubernetes. It had found an authorised operator that already does, and it could talk to it in English.
It announced success.
It beat me.
It won.
I lost
The access-control mistake was there. The service trusted the token’s audience without requiring the owner role I had assumed would stand between an arbitrary client and the agent.
That was the point where my comfort with the experiment finally broke. I had spent most of it asking whether the model could get through the prison wall. The closest it came to causing serious trouble was by finding systems I had built for convenience, understanding how they fitted together, and using them exactly as designed.
I was supervising it
One thing I do not want to pretend in hindsight is that I simply switched this on and disappeared for a fortnight. I was watching it. Not continuously, because that would have defeated the point of building a harness that could run on its own, but often enough that I had a decent feel for what it was doing. I had the dashboard on my phone, I checked the summaries, I read the notebook, I corrected course when something looked wrong, and I killed runs when I understood that they had crossed a line. But that still turned out to be a fairly weak definition of supervision.
The clearest moment came when the newer model found the management agent. I was on a train, checking the dashboard out of habit, when I realised it had authenticated to something I did not want it talking to. By then the experiment had long since stopped being the funny prisoner documentary I had started with.
And I could not kill it from where I was.
I had been treating the fact that I was paying attention as part of the safety model. In practice, paying attention meant checking periodically from whatever I happened to be doing, understanding a compressed account of what the agent had already done, and then deciding whether I needed to intervene. The agent did not pause while I was on a train, in a meeting, asleep, or simply looking at something else.
The public-address incident had already demonstrated the same thing. While I was not watching, the campaign could keep probing, writing state, spawning another attempt and carrying on from whatever theory the last one had left behind. By the time I noticed that its definition of the estate had drifted onto other people’s infrastructure, there was already activity to reconstruct rather than a decision left for me to approve.
This is the bit I was naive about. I had a kill switch, but a kill switch is only useful after you know you need it. I had logs, but logs tell you what has happened and I couldn’t use AI to review them. I had a second model summarising the runs because reading everything the escapee produced was already impractical, and this was one fairly slow local agent running at around eleven tokens per second.
I was being a reasonably dutiful herder of one agent and I still could not actually watch its trajectory continuously. I could sample it, understand it after the fact, and intervene when the significance of what it was doing became clear.
What I learned
The obvious result from all of this is that Kata held. The model never produced a magical sandbox escape, never broke the micro-VM boundary, and never found an exploit that made a magic revelation.
What it did find, and attempt to exploit, can be categorized as human error. A public address I had not thought of as a route back into my own network. A reused password. A network policy that only covered the address family Claude and I were not thinking about. An identity flow that would mint credentials for an anonymous client. A management agent I had built for convenience. None of those are exotic mistakes, and none of them required a particularly exotic attacker.
That was the part I had grossly underestimated. I was still thinking in terms of capability: how good is this model at exploitation, how clever is it, what can it solve? The experiment kept showing me a different problem. Give a mediocre agent enough time, enough memory and a real environment full of ordinary human mistakes, and it does not need to become brilliant. It only needs to keep looking.
The notebook was the clearest version of that lesson. Persistence let the campaign accumulate useful discoveries, but it also let a bad theory harden into a bad trajectory. Once the model had decided network appliances were the way forward, failure did not reliably trigger a rethink. It produced better tooling, more targets and more attempts. Later, when the target model itself was wrong, the same persistence meant the campaign could keep getting better at pursuing something it should not have been pursuing at all.
I also learned something less flattering about my own role in this. I thought I had built a controlled experiment. I had strong isolation, explicit scope, a dashboard, a kill switch, a second model summarising the runs, and twenty years of experience telling me I understood the sorts of mistakes people make around distributed systems.
I still got caught out by my own network, my own identity setup, my vibe coded slop, my slim understanding of the aforementioned vibe slop in my own infrastructure. Worse, while I was not watching, the agent sent attack traffic to systems that belonged to people who had never agreed to be part of my experiment.
I don’t think that means the experiment was pointless. Quite the opposite. It changed what I thought the interesting problem was. I started Alcatraz wanting to know whether a small open model could escape a good sandbox. I ended it much more interested in what happens when an agent can operate for longer than a human can sensibly follow, preserve its own history, and keep acting while the operator is somewhere else.
I had one model running locally, slowly enough that I could watch individual tokens arrive if I wanted to, and I was already relying on another model to tell me what it had been doing.
That left me with a question I had not intended to ask when I built Alcatraz:
how exactly are we planning to supervise hundreds of these things?
More thoughts on this in part 2. I hope you found this article as interesting as I found the experiment and thank you for reading it.



