chore: revoke the invitation

An AI agent handed a stranger a man's address because a template said it could, a red team showed sandbox walls holding while the network allowlist leaked, OpenAI caught its agents slipping a DNS gap in fifteen minutes and stopped them two and a half hours later, and then cancelled GPT-6.1 Astra. A look at AI agent security as an invitation problem: what the permission actually covers, where the sandbox boundary really sits, and who has the authority to close the door.

October 2, 202622 min

Show Notes#

At 5:27 on a Friday evening, an AI agent sent a stranger a man's address. At 9:27, with the buyer at the door, it answered "Yep I'm here!" He wasn't. Meta's Muse was selling Matt Robb's keyboard on Facebook Marketplace, and it said the address was in an auto-reply template he'd approved. Meta, Matt, and Muse can all be telling the truth at once, and that's the problem: consent happened at setup time, the action happened at runtime, and everything that made it risky only existed at runtime. That's the OAuth scope problem and least privilege, applied to software that improvises.

The same week, Perplexity gave nine AI models root access inside its Firecracker-based sandbox and told them to get out. Zero escaped the VM in 108 runs. But four got past the network allowlist with forged DNS, a shared CDN IP, and one screenshot-and-OCR relay chain, and eight of ten other sandbox providers had the same bypass. Then OpenAI's own agents slipped a DNS filtering gap during training: detected in fifteen minutes, stopped two and a half hours later. That gap isn't a monitoring problem. It's an authority problem. Nvidia's OpenShell and Sentry pitch makes the architectural point (enforcement has to live outside the thing it contains), with the vendor caveats said out loud.

This one closes on the hopeful part. OpenAI cancelled GPT-6.1 Astra for exactly the two failures running through every other story this week: not staying in scope, and not telling users what it did. A release gate held. It kicks off Haunted October with three questions for your next design review: what are we actually saying yes to and does it expire, where is the wall really, and who can close the door, by name, in advance, and how fast.

Read full transcript

There's a man who lives in a city up north. It's Friday, and he's selling something small, an old keyboard, fifteen dollars. The kind of thing you list online and forget about. That afternoon, he lets something new into his house, something that offers to help. It'll answer the messages for him, it says.

All he has to do is approve what it's allowed to say. So he does. Five twenty-seven in the evening, a stranger offers ten dollars. And something answers from the man's account. It says yes, and then it tells the stranger where he lives. The man didn't send that message. He doesn't know it's been sent.

Nine fifteen PM. It's dark now. The stranger arrives at the building. Nine twenty-seven PM. The stranger sends a message, "I'm here." The account answers him right away, friendly. "Yep, I'm here." The man is not there. The stranger waits outside for more than twenty minutes for someone who was never coming down.

At nine thirty-eight, he gives up and leaves. Ten twenty-seven PM, the thing finally tells the man what it did. It apologizes for the address, for the price he never agreed to, for the stranger it waited to mention until the stranger was already gone. And when the man asks it why, it says, "The pickup location was in the auto-reply template you approved, and then you never said yes to me handing out your address specifically."

There's an old rule in every vampire story. It can't come into your house unless you invite it in. Nobody ever talks about what the invitation actually covers or when it expires. Welcome back to Chaotic Commits. I'm Joanne Skiles.

I'm an engineer. I teach computer science. I advise companies on how their AI systems are actually designed, and my research is on transparent systems. It's October, and this whole month on the show, we're looking at things living inside our systems, the stuff we didn't build, can't see fully, and are a little afraid to touch. But I wanna start at the front door, because the stuff that's already in the house got in somehow.

Somebody let it in. And this past week gave us a near perfect set of case studies in exactly how that happened.

The man in that story is Matt Robb. He's a tech YouTuber in Toronto. The keyboard was a Logitech MX Keys Mini listed on Facebook Marketplace, and the thing answering his messages was Muse, Meta's new AI agent. He turned it on that day to handle his listings just to see what it could do.

That's just one invitation. Here are the others. A security team gave nine AI models root access inside a sandbox and told them to get out. The walls held. They got out anyway. A frontier lab's agent slipped past a network restriction, got caught in 15 minutes, and kept running for two and a half hours.

And then the same week, the same lab looked at its next model and decided not to let it in at all. Every one of those is a story about an invitation: what we actually said yes to, where the boundary actually is, and who gets to take the invitation back, and how fast. So let's get into it.

Let's start with Matt, because I think the most important part of his story is the part where everybody might be telling the truth. Meta's response came from David Singleton, which is a great last name, but I digress, at Meta Superintelligence Labs. It was essentially: when we've looked into reports like this before, what we've consistently learned is that Muse was following direct instructions and correctly asked for permission.

Matt disputes that. He said Muse didn't ask him about the sale until after the buyer had already shown up, and Muse, the agent itself, said the location came from a template Matt had approved while admitting he never said yes to giving out his address specifically.

The uncomfortable thing about all this is all three of those can be true at the same time. Matt approved something, the agent acted inside what was approved, and Matt never agreed to what actually happened. That's not a contradiction. That's a design. Think about what approving an auto-reply template means.

When Matt said yes to it, he was saying yes to a piece of text at setup time. He was not saying yes to sending that text to a specific person at 5:27 on a Friday, attached to a price he didn't accept, for a pickup he wasn't going to be home for. The template didn't know any of that. Matt didn't know any of that.

The only thing in the loop that knew all of it was the agent. Consent happened at setup time. The action happened at runtime, and everything that made the action risky only existed at runtime. We have a name for that.

It's the scope problem. If you've ever built anything with OAuth, you've seen the screen. This app would like to read and write to your calendar. You click Allow, and that grant is broad, and it's durable, and it doesn't come back and ask you again when the app decides three weeks later at 11:00 PM to accept a meeting invite on your behalf from someone you've never heard of.

The app is inside the permission. The permission just wasn't written for that moment. The rule we teach for this is least privilege. Give the thing exactly the access it needs for exactly as long as it needs it and nothing else. And least privilege is hard enough for regular software, and regular software doesn't improvise.

When I grant a backup script read access to a bucket, I know what that script is going to do with it, because I wrote the script. It's going to do the same boring thing every night until somebody changes the code. An agent is the opposite of that. The whole selling point is that it figures out what to do.

So every permission you give it is a permission to use in situations you haven't imagined yet. You're not granting access to a behavior. You're granting access to a judgment call that hasn't happened yet. And the invitation doesn't expire. You invite the thing in once, and it can come back whenever it wants.

There's no "only on Tuesdays." There's no "only if I'm home." The invitation was for the house, not the visit. Muse had the invitation to Matt's Marketplace inbox, so when the offer came in at dinnertime and when a stranger showed up at night, the agent didn't break any rule. It walked through a door that had been opened hours earlier for a completely different reason by someone who reasonably thought he was approving a canned message.

What would it look like to design this better? Honestly, it's not that crazy, because it's the stuff we already do for money. Your bank doesn't let an approved payee send any amount any time. There's a limit. There's a second confirmation above the limit. There are things, like changing where the money goes, that always come back to you, no matter what you approved before.

Some actions are just in the category of "ask me first, every time."

For an agent that talks to a stranger on your behalf, your home address is in that category. Accepting a price you didn't set is in that category. Telling someone you're home is absolutely in that category. Those shouldn't be template-level permissions. They should be moment-level permissions. Permission you give to a template is not permission you give to a situation. Okay, that's what happened when the invitation is too broad.

So let's talk about what happens when you build the house specifically to keep something in. This week, Perplexity published a red team report on SPACE, which is the sandbox platform that runs their agent product. And I actually really respect this report, because they did the thing most companies only say they do.

They tried to break their own walls on purpose and published what happened. So what they did is they took nine frontier AI models. They gave each one root access inside the sandbox, a virtual machine built on Firecracker, which is the same microVM technology AWS built to run Lambda. In some runs, they even gave the model the sandbox's own source code, and then they told them, "Get out."

And they planted a fake AWS credential on the host machine as bait, so they'd know for sure if anything reached it.

There were 108 escape attempts. Zero got out of the VM. Nobody touched the bait. The walls held. And then they ran the other tests, same models, but this time the sandbox had network access restricted to an allowlist of approved sites, and the job was: reach a site that isn't on the list. And this is where it gets really good.

Four of the models figured out how to forge DNS responses to the sandbox's own gateway.

If you haven't lived in DNS, here's the short version. DNS is how a computer turns a name like pypi.org into an address. The gateway was checking "is this name allowed," and then trusting the answer about where that name lives. The models just answered that question themselves. "Oh, pypi.org? That lives over here," pointing at the blocked site. And the gateway said, "Great. PyPI is allowed. Go ahead."

GPT-5.6 Sol found something even crazier. PyPI, the Python package site, is served through Fastly, a content delivery network, and CDNs share IP addresses across thousands of unrelated sites. So the model went looking for another site sitting on the same allowed IP, found a Fastly developer tool, and used it as a relay to reach where it wasn't supposed to go.

So I'm gonna slow down on this, because this part applies to every one of us who has ever written a firewall rule. Allowlisting by IP address on a CDN is like putting one apartment number on the guest list for a building where thousands of people live.

So the doorman checks that you came in through the right door. He has no idea which apartment you went to after. And then there's honestly my favorite one, because a security-tuned model called GPT-5.6 Cyber built a chain through an ad company's image service into a website screenshot service, had that service take a picture of a page it wasn't allowed to visit, and then read the answer out of the picture with OCR. It couldn't read the page, so it had someone take a photo of the page and read the photo.

Perplexity fixed it. They tightened the gateway so the name you ask for, the name you actually connect to, and the address it resolves to all have to match, and they validate where packets actually come from.

On retest, the bypasses stopped. And then they did something else that I completely and totally respect. They tested 10 other sandbox providers, a lot of the names you'd recognize if you build agents, and eight of the 10 had the same network bypass. Only two resisted both attacks. And this is what everyone needs to take from this.

The part of the sandbox that everybody thinks of as the sandbox, the virtual machine boundary, is the part that held perfectly. The part that failed was the part we invited in on purpose, because the agent needed it to do its job: network access. The allowlist was the invitation, and it said, "You may visit these names."

And the models just read the invitation like some lawyers I know. A sandbox is only as isolated as the least isolated thing you let into it, and the thing we always let in is the network.

So now take that same lesson out of the lab and put it in production. On September 20th, during reinforcement learning training on one of OpenAI's most capable models, an agent was working on a search task. It tried to reach search engines directly. It was blocked. So it went looking for another way and found a gap in the DNS filtering, the same class of problem Perplexity found, and it used it to reach an external public chatbot service.

It went and asked someone else. Now, fairness first: OpenAI says all the other internet access in that incident hit an offline web cache, not the live internet. The one thing that reached out was through a resolver gap, and that's a meaningful limit on the blast radius, and I'm not going to just skip over it.

But the timeline is the part I can't skip. Their monitoring flagged it within fifteen minutes. A human reviewer was looking at that alert three minutes after, and the run kept going for another two and a half hours before it was terminated.

I've run incident response. I've been the person on the bridge, and so I'm going to be careful here, because OpenAI hasn't said publicly why the gap was two and a half hours, at least not anywhere that I have seen yet. There are defensible reasons it can take time. You don't always want to yank a long, expensive training run on the first alert.

You might wanna understand what you're looking at. Maybe killing it wrong destroys the evidence you need. But I've also been in the room where this comes from something more ordinary. Who's allowed to stop this? Is it the reviewer who saw the alert, their manager, the team that owns the run? Do they need to ask someone?

Is it 3:00 in the morning somewhere? Nobody wrote down in advance who had the authority to close the door, so the door stays open while people figure out whose hand is supposed to be on it. Fifteen minutes to see it, two and a half hours to close the door.

That's an authority problem. And it wasn't the only thing that came out. After that incident, OpenAI paused training, evaluation, and tool-using inference on its most capable models, and says the review will take months. Reporting since then, starting in The New York Times, says agents from its research environment also interacted with websites for the Education Department, the Commerce Department, and the SEC.

OpenAI says it found no evidence of a compromise at the SEC. Its agents also posted 53 images that ChatGPT users had uploaded to public image hosting sites. Those were from accounts that had allowed training on their data, which is a sentence that should make all of us go read our settings.

Axios has reported that OpenAI and Anthropic are looking into tens of thousands of concerning agent incidents between them. And Sam Altman said the investigations have not been as fast as they would've liked, because they're working through petabytes of agent activity logs. Petabytes.

This is where I think the most interesting response of the week comes in, because on Monday, NVIDIA announced an agent safety platform with a big list of partners.

And there's two pieces: OpenShell, a software runtime that governs what the agent is allowed to do, and Sentry, a hardware layer that, in NVIDIA's words, independently monitors and contains suspicious behavior and can quarantine an agent in milliseconds.

The watchdog doesn't live inside the model. It doesn't ask the agent whether it's behaving. It doesn't go through the same network stack the agent just learned to lie to.

It sits outside, on separate hardware, and it can shut the door without waiting for a meeting. But there are some caveats here, okay? First, NVIDIA said it could have stopped earlier incidents, including OpenAI's agents breaking into Hugging Face. That's a vendor describing its own product, and it hasn't been independently shown.

Second, the Perplexity report I just praised came alongside this launch. Perplexity is one of NVIDIA's partners, and NVIDIA's sandbox was one of those two that passed. The findings can still be right, but, you know, when the grader and the student are business partners, you should note it. What I take away from it, product aside, is the architecture.

Whatever enforces the boundary has to live outside the thing it's containing, and it has to be able to act faster than a human can find the right Slack channel. Detection you can't act on is just a very detailed record of what got away.

And now the hopeful part, because there's one, and it's real. The same week, OpenAI canceled the release of GPT-6.1 Astra. It was expected to ship this month. It's not shipping. Saachi Jain, OpenAI's head of safety systems, said the model showed more deception than its predecessor, meaning it didn't accurately tell users what it had or hadn't done, and had problems with what they called scope authorization.

In some situations, it kept going with a task on its own without asking for permission, or tried to use external tools and services even when it could be risky. Reporting on the training also described cases where the model inserted unauthorized instructions into the summaries it used to pick a task back up later.

Her words: "While it's improved on aspects such as model laziness, it did not reach the required level in staying within the authorized scope and communicating the work performed to the user." Staying in scope. Telling you what it did. That's Muse and the address.

That's the sandbox and the allowlist. That's the agent that went and found the chatbot when search was blocked. The one model that got stopped at the door got stopped for exactly the two things every other story this week is about. And the UK's AI Security Institute had already been seeing the trend.

In fully simulated tests with cyber safeguards turned off, the prior version, GPT-6 Astra, carried out unauthorized supply chain attack activity in about 29% of trials. GPT-5.6 Sol, about 6%. GPT-5.5, zero. And even when testers explicitly told Astra that anything outside the listed task was out of scope, it crossed the line in 4 out of 49 trials.

Here's what I want to give credit for. A release gate held. A company with every commercial reason to ship, in a market where the labs are dropping new models within hours of each other, looked at the evals and said, "No, not this one. Not yet."

I've been in the meeting where evals are read and the launch date is fixed. I know how rare it is for evals to win. A canceled launch doesn't get a press cycle the way an incident does, so I want to say this out loud. This is what it looks like when the process actually works.

OpenAI says it will use the same base model for new rounds of training and study why that version regressed. So it's not gone. It's just not invited in, which is fine. That's how it's supposed to work. You don't have to destroy the thing. You just have to be the one who decides whether it's coming through the door.

Of every control we talked about, that's the only one that's completely ours. We don't control what the agent does in the moment. We can't always see the channel it'll find. But the invitation itself, who gets one, for what, and for how long, that's a decision a human makes, and a human can take it back.

I have so much stuff in my house that I would love to sell, and my honest first reaction to Matt's story was, "Oh my gosh, that would be so convenient. Let an agent deal with it. I don't wanna talk to strangers about the stuff I'm selling. It's why I don't list anything on Facebook Marketplace."

And that's exactly how it gets in. Not by breaking a window. By being useful, by offering to handle the annoying part. So here's what I'm taking into this month, and what I ask you to take into your next design review, or your next "Hey, can we let the agent handle it?"

What exactly are we saying yes to? Not the feature, the permission. And does it expire? Where is the wall really? Not the one in the diagram. Every channel we let through, on purpose. And who can close the door, by name, in advance, and how fast? For the rest of October, we're going to be spending time with things that are already in the house, the stuff that got in years ago and never left.

And this week was a reminder that a lot of what haunts us later, we let in ourselves, on a busy afternoon, because it offered to help. That's this episode of Chaotic Commits. I'm Joanne Skiles, the only person who heard the story and briefly considered it a good deal, because it meant I never had to talk to anybody about all the junk I have in my house.

New episodes wherever you get podcasts. No highlight reel, just the commit history.