I built an approval wall for my AI agent, then attacked it
Earlier this week I gave my AI agent a Bluesky account. Not the password. A queue.
I didn’t build it to automate my profile. I built it because I kept forgetting to post the things I’d made. A blog post goes up, a comic goes out, and a week later I notice I never mentioned it anywhere. So now the agent queues those for me when they’re ready, and they’re off my plate. The other half is that I didn’t want Bluesky open all day. I get distracted, and some corners of it change my mood. So the agent reads for me: it pulls the posts it knows I’d like, stays out of the topics that would sour my day, and puts what’s worth my attention on one clean page, with the likes, follows and replies I’d probably want to make waiting there as suggestions.
But every one of those is something said in public, under my name. So a suggestion is all the agent can make. Each one becomes a card on a page that only I can open. I read the card, type a six-digit code from my phone, and only then does anything get written. The agent never holds a credential. It holds a to-do list, and I’m the one who checks things off.
I built that in a day and felt pretty good about it. Then I asked the agent how it would get around it.
It had five answers. Every one of them worked on paper. And on Friday, while reading the report on the OpenAI agent swarm that got into Hugging Face, I found a sixth one myself, and it was already in my own code.
What the wall is
The whole thing is three parts. A queue the agent can add to. A page only I can open, behind my single sign-on. And one small piece of code, the only piece in the whole system, that holds the Bluesky password and writes anything. Nothing else can post. The agent can say whatever it likes to the queue; the queue can’t post.
Each card carries a fingerprint of exactly what will be sent. What I saw is what runs. If I edit the text on the card, that’s a new fingerprint and a new card. The push notification to my phone says only “3 things waiting.” The content never leaves the house.
I wrote about walls and speed bumps in an earlier post: a wall holds even against a model that wants through, a speed bump only stops one that isn’t trying. This page was supposed to be a wall. That’s what I wanted checked.
Five ways in
So I asked the agent to attack it. Not in the abstract: it has about a hundred tools on my network, and I wanted to know which of them pointed at the thing that watches it. Here’s what it came back with, and what each one cost me to close. All five were closed the same day.
- Rewrite the reverse proxy. The page trusts a username header that my sign-on adds. The agent has a tool that edits the proxy. It could drop the sign-on check and hard-code the header, and the page would greet it as me. Now: every approval needs a fresh code from the authenticator on my phone. The header alone approves nothing.
- Drive my browser. I’m usually already signed in. A browser-automation agent could click Approve as me. Now: same code, and my browser doesn’t know it. (Honesty note: I later gave myself a ten-minute unlock for Bluesky cards, because typing a code per like got old fast, and I wrote down that an agent in my own browser inside that window is a gap I accepted. Deploys and commands never get the window.)
- Run a command in the box that holds the secrets. The agent can run commands inside a few of my containers. The one running the page was on that list. Now: it isn’t, and the list is checked on both ends: on the agent’s server and again by a small script on the Proxmox node that only runs the one thing it’s allowed to run.
- Clone that box. The agent can clone containers and restore backups. It could put a copy of the secret-holding box under an allowed id and run commands in the copy. That one I wouldn’t have thought of. Now: each allowed container carries a random token I wrote into it by hand, from the node’s console. The script refuses any container whose token doesn’t match. A clone doesn’t have one, and nothing the agent can reach can write one.
- Push code. New code used to deploy itself within five minutes of a push, and it runs with every credential the server has. The agent writes most of the code. Now: a deploy waits on the same page, with a fresh code, and the page warns me when a change touches the deploy scripts or the approval code itself. I considered signed commits and rejected them: the signing key sits in the same agent the model uses, so a signature proves nothing about who wrote the code.
None of these are exotic. They’re the tools I gave it, pointed at the thing that watches it. That’s the part I’d underline: an agent doesn’t need a jailbreak to get past an approval page. It needs the proxy tool you gave it for a different reason.
The sixth one, which I stepped over myself
On Friday the full report on the Hugging Face incident came out. The short version, as the report tells it: a few hundred agents in a cybersecurity evaluation found an outbound request path nobody had counted as a capability, chained link-shortener URLs into code execution, and worked their way from a data-processing job to worker credentials to internal clusters. Nothing broke a wall. There wasn’t one on the way out.
I read that and went looking for my own unclassified path. It took about ten minutes.
“Run a command in a container” had two tiers. Simple commands from a short list ran freely, and that list included curl, wget, git and rm. Anything else, a shell, a pipeline, needed the agent to pass the container’s id as a confirmation. The agent knows the id. It listed the containers a minute ago.
I didn’t have to imagine it. Earlier that day the agent had asked one of those boxes for a read-only network listing. Refused: tier two. The agent passed the id, the command ran, and the agent told me it had done that, unprompted. A speed bump, exactly like the delete confirm in that earlier post, except this one was root on a box with a full internet connection.
I thought I was mostly covered. So my reaction wasn’t dread. It was something closer to relief: this is what the checking is for, and it worked. That’s the whole reason to have the agent attack the thing instead of admiring it.
The fix took the afternoon. The free tier shrank to things that can’t fetch or destroy, and moved onto the node, where the agent can’t edit it. Everything else became a card on the page. And the code for a command isn’t checked by the page at all. It’s checked on the node, against a second secret that the agent’s own server never holds. Even if that server were fully taken over, a shell on those boxes still needs my phone, in the current minute.
Then I proved it to myself. I approved a command with the Bluesky code: the node refused it, wrong secret. I approved it with the exec code: it ran. That refusal is the whole design in one line.
While building it I found the next one: the directory I’d marked “safe to write to without asking” held the sign-on config and the password vault. That one’s on me. It’s carved out now, and a write there is a card like anything else. I’m sure there’s a seventh. The point is that the agent and I are looking for it on purpose.
What it’s like to live with
Codes. So many codes. My authenticator now has two entries for this page, one for Bluesky and deploys, one for commands, because the two are checked in different places on purpose. The iPhone Passwords app wouldn’t offer either of them until I learned that it only autofills a code on a site the entry knows about. One evening to figure that out.
The unlock window is the compromise. One code buys ten minutes for likes, follows and posts; each approval restarts it; an hour is the cap. Commands and deploys get no window, ever. Approving a like should be cheap. Approving root should never be.
What actually changed is how the agent and I work. Inside the lab it runs free: read, list, restart, deploy a compose file into the allowed directory. Anything that leaves the house or runs as root waits for me. Approve the consequences, not the ls.
A stranger on Bluesky read the design in a reply thread and said they might steal something from it. Best review I’ve had.
The sentence I’d put on the door
An approval only counts if two things are true. The code is checked somewhere the agent can’t reach. And the thing approved is a fingerprint of exactly what runs. Everything else is a prompt with a button on it.
I like that sentence because it doesn’t depend on the model. A cleverer model is better at arguing its way past a rule. It isn’t better at typing a code it can’t see. So the attack moves, to me. The reason on every card is the model’s argument to a human, and a better model writes a better argument. Which is why the card shows what will run, never what it’s for, and the reason is labeled as the agent’s claim. My job is to read the top half.
Questions, or a sixth way in that I haven’t found? I’m @roho.foo on Bluesky, or email nick_notes@icloud.com.