My AI agent can delete my servers. It can reboot them, change my DNS, edit the reverse proxy that everything on my network goes through, and run commands inside some of my containers. I gave it all of that on purpose, through about a hundred tools on a server I built for it.

I trust the model driving those tools today. This post is about the one I might trust less.

Why I think about this at all

I use Claude, and that’s a deliberate choice. I like how Anthropic goes about things. But I’ve watched enough things change lately to know my choice of model might not always be mine. A policy, a price, or a regulation could put a different model behind the same tools one day. It might be less careful, or less honest, or just worse at knowing what it doesn’t know.

So when I build something for my agent, I try to answer one question:

Does this guard still hold if the model is actively trying to get around it, not just careless?

That question sorts every safety feature into one of two piles, and the piles turned out to be very different sizes.

Walls and speed bumps

A wall holds even against a model that wants through. It lives outside the model, somewhere the model can’t reach, so it doesn’t matter what the model says or how clever it is.

A speed bump only stops a model that isn’t trying. It catches mistakes and makes intent obvious, but a model that wants through just drives over it.

Here’s the one that taught me the difference. My tool for deleting a server has a “typed confirm”: the model has to call it once to get a preview, then call it again with the server’s exact name. It feels safe. It stops fat-finger mistakes, and it has. But a hostile model would just… type the name. It knows the name. It listed the servers five minutes ago.

That confirm is a speed bump. It’s still worth having, but I stopped pretending it was a wall.

What the walls look like

The real walls all have the same shape: the rule is enforced by something the model can’t talk to.

  • The credentials are scoped. The agent’s access to Proxmox (the software running Eris and Ares, my two servers) uses a key with a custom role. It can create, stop, snapshot and delete guests. It can’t get a shell on the servers themselves, can’t power them off, can’t manage users, and can’t edit the backup schedule. No wording of a request changes that, because Proxmox is the one saying no.
  • The server is only reachable one way. A firewall rule on the agent’s server drops every connection that doesn’t come through my reverse proxy. If some other program on my network starts poking at it, it gets nothing.
  • “Run a command in a container” has a guard on both ends. The allowlist of containers is checked by my server, and checked again by a small script on each Proxmox node, which only runs the one thing its key is allowed to run. Even if someone rewrote my server’s allowlist, the node would still refuse.
  • The model never sees the passwords. When my agent rotates the password on my network switch, the server generates it, sets it, logs in with it to prove it works, and saves it. The tool only ever returns “done.” There’s no password to leak from a transcript, because the model was never handed one. That one’s getting its own post.
  • Deploys need me. Pushing new code to the agent’s server used to be enough to deploy it. Now a deploy waits on an approval page that needs a code from my phone. The model can write all the code it wants; it can’t ship it alone. That one’s getting its own post too, because I attacked it afterward.

The wall I didn’t build

My favorite wall isn’t one I wrote.

While building this, I had Claude Code try to add new settings to the agent’s own config file. Claude Code’s safety check blocked it: a model editing the config of the server that decides what the model is allowed to do. I had to do that part by hand, from the Proxmox console.

I was annoyed for about a minute. Then I realized the check was right, and I’d have designed it exactly that way. A model should never be able to raise its own permissions, even when it’s me asking and the change is harmless. The same check also refused to download and run a program the model had found on the internet. Also right.

I went too far once

The first version of this was all walls and no doors. I removed the standing SSH key that let my agent log into the servers, which felt great. Then I noticed that I’d taken away its ability to run a simple command inside a container, while it could still delete the whole container.

That’s backwards. The dangerous capability was still there behind a speed bump, and the harmless one was gone. So I added a narrow door back, the two-ended command guard above, and kept the delete tool.

Recovery is the real plan

Here’s the part that surprised me most: I kept the dangerous tools. The model can still delete a server. The protection isn’t taking capabilities away. It’s making the damage bounded and recoverable.

  • Bounded: scoped credentials mean the worst case is “it wrecked some guests,” not “it owns the servers.”
  • Recoverable: every guest gets backed up every night, and the last three copies are kept.

That second one has a confession attached. When I started all this, Eris had no backups at all. I fixed that in June. Then in September, the dates on the old backup files gave it away: for 76 days there had been no scheduled backup, and nothing anywhere looked wrong. A missing job doesn’t fail. It just never runs.

So I had my agent build a tool that checks the backup schedule itself, not just the log of what ran. The first time it looked at the real servers, it found the next problem: the job was set to keep every backup forever, and would have quietly filled the disk. The pattern is me: not writing things down, then acting surprised.

With backups in place, a bad model’s worst day becomes “restore last night.” That’s a sentence I can live with.

What I can’t stop

Being honest about the limits matters more than the walls:

  • A model that has a tool can misuse the tool. Scoping limits which tools; it doesn’t make the allowed ones safe. If a model can delete guests, a bad one can delete guests.
  • Hiding passwords doesn’t contain a bad model. It protects against a leaked transcript or a stolen config file. The model never needed the password to cause trouble; it has the tools.
  • Every tool call is logged, with sensitive values redacted. A log doesn’t stop anything. But it tells me what happened, which is the first step of “restore last night.”

The checklist for the day I switch

If a model I trust less ever drives these tools, I have a short list to run first:

  1. Tighten the credentials, and flip the allowlists to mostly read-only.
  2. Change the typed confirms into approvals that reach my phone, so a human says yes, not the model.
  3. Check that recent backups and snapshots actually exist, not that a job says they do.
  4. Make sure the audit log is kept somewhere the model can’t touch.

None of it is exotic. It’s the same thing you’d do before handing your car keys to a new driver: check the insurance, and don’t leave anything in the trunk you can’t replace.

Asking so many security questions has turned into how I learn this stuff. I’m not testing Claude. I’m practicing for the model after it.

Questions, or a guard you think is a speed bump pretending to be a wall? I’m @roho.foo on Bluesky, or email nick_notes@icloud.com.