stale read: down forty minutes after it wasn't
Show Notes
It's the morning of July 16, 2026, and I'm refreshing a status page like reloading it faster will change the answer. AWS's CloudFront service has been down since before I woke up, half the internet's education platforms have gone with it, and the page keeps promising an ending. At 11:27 UTC: initial signs of recovery. At 11:57 UTC: significant recovery, full recovery in forty-five minutes. What I didn't know at the time: the outage had already ended, completely, at 11:18 UTC — nine minutes before the first of those updates, and thirty-nine before the second one handed me a countdown to something that was already over.
This isn't "AWS had an outage, let's dunk on AWS." Every major cloud provider has had one in the last eighteen months. It's about the forty-minute gap between a system recovering and a system telling you it recovered, and about a trade-off baked into the feature that caused the outage in the first place: CloudFront's VPC Origins closes off public access to your origin server for real security reasons, but that also means there's no way around it when it breaks. AWS's own workaround during the live incident was to tell customers to switch the feature off.
I'm upfront that this wasn't the biggest outage of the year — not close — and that's kind of the point. A clean, small incident with one root cause and one specific forty-minute gap is easier to actually see through than a headline-sized one.
If you've ever refreshed a status page like reloading it faster would change the answer, or signed off on an architecture decision without asking what you were quietly trading away, this one's for you.
Read full transcript
It's the morning of July 16th, 2026, and I'm refreshing a status page, which is already off to a great start. And I'm refreshing it not once, but repeatedly. The way that you refresh a post that has a thread where someone said something dramatic and you know, the internet is just going to explode. AWS's CloudFront service has been throwing errors since before I woke up. Half the internet's education platforms are down, and I'm doing the thing every engineer does during an outage, which is stare at a page as if staring harder will make the words change.
At 11:27 UTC, the status page tells me there are initial signs of recovery. At 11:57 UTC, it tells me the recovery is significant now and to expect full recovery within 45 minutes.
Here's what I didn't know at the time and what AWS's own retrospective says happened. The outage had already ended completely at 11:18 UTC, 9 minutes [clears throat] before the first of those two updates and 39 minutes before the second one gave me a countdown to an ending that had already happened.
I was reading a forecast built on a version of the system that no longer existed.
That's the whole episode actually, not the outage, the nine minutes and then the next 30. The gap between something being true and something being told to you.
Welcome back to Chaotic Commits. I'm Joanne. Today's episode is called Stale Read Down 40 minutes after it wasn't.
And I want to say upfront what this episode is not. It's not AWS had an outage. Let's dunk on AWS. Every major cloud provider had an outage in the last 18 months. Google Cloud in June of last year, Azure in October, Cloudflare as we all remember wonderfully in November, and of course AWS itself catastrophically in October of 2025, one of my favorite days. Distributed systems at that scale fail. That's not news and it's not the interesting part.
The interesting part is narrower and I think more useful. It's about what happens in the space between a system recovering and a system telling you it recovered. It's about a trade-off baked into the architecture that caused this outage in the first place. A trade-off that as far as I can tell, nobody wrote down anywhere a customer evaluating that architecture would have seen it.
We're going to walk through what actually happened hour by hour, then the mechanism, then the tradeoff nobody documented, and then why the size of this outage almost doesn't matter and why I think that's the whole point.
So, let's get into it.
Here's the sequence, and I'm giving you real time stamps because the timestamps are the story.
7:45 UTC, that's 12:45 a.m. Pacific, customers using a CloudFront feature called VPC Origins start seeing elevated 500 errors. VPC Origins is a routing option that lets you point CloudFront at an origin sitting inside a private VPC instead of a publicly reachable one. And we'll come back to exactly why that matters.
8:44 UTC AWS posts its first public acknowledgement investigating increased errors related to VPC origins.
9:21 UTC AWS confirms the 7:45 start time, clarifies that other origin types are unaffected and publishes a workaround. Hold on to that workaround. We're coming back to it because it's doing more work in this story than a one-line migration usually does.
9:57 UTC root cause identified internally though not yet shared publicly.
10:18 UTC a public update references a packet processing subsystem. The workaround gets repeated.
10:52 UTC multiple mitigation actions get kicked off.
11:16 UTC AWS scopes the problem specifically to routing table capacity and plans a phased roll out of the fix.
11:18 UTC. According to AWS's own retrospective summary published later, this is the moment full recovery actually happened.
11:27 UTC. A live update reports initial signs of recovery.
11:57 UTC. Another live update repeats significant recovery with full recovery expected within 45 minutes. Both of those land after the moment AWS's own retrospective marks as full recovery.
12:21 UTC. A resolution summary finally post and customers are told to revert whatever workound they put in place.
Total impact window from 7:45 to 11:18 UTC. That's 3 hours and 33 minutes.
Meanwhile, this wasn't contained to AWS's own dashboard. Canvas and Blackboard, two of the platforms a huge chunk of higher education runs on, went down. Hugging face, a home base for a big piece of the AI developer ecosystem went down. So did the UK national lottery tail scale, an identity platform called Frontegg, a telehealth company called Doxy, a collaboration tool called Coda, and a Japanese blogging platform called Hatena. A single capacity limit in one availability zone in Frankfurt turned into a multicontinent, multi-industry outage touching education, healthcare, AI tooling, and apparently whether the UK could run its lottery.
That's worth sitting with for one second before we move on. A regional problem produced a global outage. That's not a fluke. That's architecture. And that's part two.
Here's the mechanism as plainly as I can put it. The fleet that manages connections to private VPC origins hit an internal constraint. Specifically, AWS traced it to the availability zone euc1-az2 in the Frankfurt region. that fleet's routing table capacity inside its packet processing subsystem got exhausted. When that happens, the system responsible for distributing updated networking configurations out to CloudFront's edge network processors could not load that configuration correctly.
And I want to pause on the word load because it's doing something specific.
If you followed cloud outages over the last year and a half, you've seen a pattern. Google Cloud. In June of 2025, a bad automated policy hit a null pointer bug in a core service. It propagated globally Azure. In October of 2025, an inadvertent front door configuration change slipped past the validation that was supposed to catch it and it propagated globally. Cloudflare. In November of 2025, a malformed feature file got pushed to their bot management system globally, and it blew past a hard-coded size limit. The pattern in every one of those. Something wrong got pushed out everywhere, all at once.
This one is a variant, and the variant matters. It's not that a bad configuration went out. It's that the correct configuration failed to load in the first place. Different failure, but same blast radius. One is we shipped the wrong thing everywhere. The other is we couldn't get the right thing to load anywhere. AWS hasn't published the specifics of what the internal capacity constraint actually was, only that it existed and that it was scoped to the routing table capacity. That's a real gap in the account, and it's fair to name it as one.
But I want to spend the rest of this section on the other half of the story, the part that isn't about AWS's infrastructure at all. It's about the status page.
There's a term in distributed systems, and I promise this is the last piece of jargon. I'll ask you to hold on to a stale read. It means you ask a system for its current state and the answer you get back was true a moment ago, but it isn't anymore. The write already happened somewhere. Your read just hasn't caught up to it yet.
That is exactly what was happening to every engineer refreshing that status page between 11:18 and roughly 11:57 UTC. The system had a new state fixed, recovered, done. The read hadn't caught up.
And notice what the stale read actually produced at 11:57. Not just a vague not yet, a specific number. 45 more minutes. That estimate was built entirely on information from before 11:18. It wasn't a bad guess. It was a correct calculation performed on an input that had already stopped being true.
So, here's a dumb analogy, but it's one that I keep coming back to. It's the unread message badge on your phone. It doesn't clear the second you actually read the message. It clears when the app decides to tell you it's been read, which is sometimes now, sometimes 40 seconds from now, and you have no way of knowing which. For about 40 minutes on July 16th, AWS's status infrastructure was the badge that hadn't cleared. The underlying thing was already fine. The indicator hadn't caught up to reality and there was no way from the outside to tell the difference between still broken and fixed telling you slowly.
That's worth naming clearly because I don't think people build this mental model by default. a status page, a monitoring dashboard, an incident feed. That's a system, too. It has its own pipeline. It has its own latency. Usually, it has a human somewhere deciding when it's safe to say the words resolved instead of monitoring because saying resolved and being wrong is its own kind of failure. That pipeline can lag the thing it's reporting on. Not because anyone's lying because confirmation takes time and the page doesn't know what it doesn't know yet either. You were reading a value that already changed. Nobody told you it changed. Those are two different facts and only one of them is technically a bug.
Okay, so here's where I actually want to land this episode because outage happened, status page lagged is a fine technical story, but it's not the one that's been sitting with me. Let's talk about why VPC origins exist at all.
Before this feature, if you wanted CloudFront in front of your application, your origin, the actual server or load balancer doing the work, generally needed some kind of public reachability. That's a real attack surface. People scan for it. People try to find your origins IP and go around your CDN and your WAF entirely. VPC Origins closes that door. Your origin lives inside a private VPC with no public exposure and CloudFront reaches it through a private path. That is a legitimate meaningful security improvement and I don't want to undersell that anywhere in this episode.
But that improvement quietly requires something. If CloudFront's private connectivity path is the only way in, then CloudFront's private connectivity path becomes the only way in. When that path has a bad day, such as July 16th, you don't have the fallback you have had with more conventional, more exposed setup. There's no just hit the origin directly for a minute option because the entire architectural decision was to make sure that option didn't exist anymore.
Here's the detail that turned this from a hypothetical into something I watched happen in real time. AWS's own publish workaround. The thing they told customers to do at 921 UTC while the outage was still active was to change the origin type. move your distribution off VPC origins and onto a public origin and you were back in business.
Sit with that for a second. The fix during a live incident was to temporarily undo a security feature. To route around the outage, you had to give up the thing VPC Origins was built to give you. That's not me speculating about a theoretical trade-off 3 years from now in a blog post. That's customers in the middle of the real outage being told, "Turn off your sole ingress protection if you want your site back up."
Genuinely great if what you're worried about is someone getting in who shouldn't. Also locks out the fire marshall who just needs to check the panel during a drill. And the fact that it does both of those things is not a flaw in the door. It's the door working exactly as designed. Nobody hands you that trade-off sheet when they sell you the lock. You find out what the lock costs the first time you need to get past it in a hurry.
I went looking for where AWS documents this specific inseparability. Security through sole ingress. Fragility from sole ingress. a thing you should weigh before adopting VPC Origins. I didn't find it. Maybe it exists somewhere I didn't look, but near as I can tell, this is a trade-off that customers discover by living through it on July 16th. Not one they were handed before they made the architectural decision.
And I want to be careful here because there's a version of this argument that's just software has trade-offs. More news at 11. That's not what I'm saying. Making the tradeoff is fine. Security for flexibility is a completely legitimate design decision, and plenty of teams would make it again, knowing everything we know now. What isn't fine is a customer finding out what they gave up only during the incident that finally prices it. An architecture decision that trades flexibility for security is a decision. An undocumented one isn't a decision anymore. It's a discovery. And discoveries made at 8 a.m. during an active outage are the expensive kind.
I want to be honest about scale here, too, because it would be easy to make this sound bigger than it was. It was not the largest outage of the year. Not even close. 3 and 1/2 hours isn't a record. The October 2025 AWS DNS incident, the one that took down a huge share of the internet through a Dynamo DB endpoint failure in US East1, generated more than 6 and a half million down detector reports across over a thousand companies and ran for something like 15 hours. And that was an amazing day that I will never forget. This July incident logged around 350 down detector reports. That's not the same order of magnitude. It's not the same order of magnitude twice over.
So why am I building an episode around the smaller one? Because the size of the outage tells you how many people it touched. It doesn't tell you anything about whether you could trust what you were being told while it was happening. Those are different measures and I think we conflate them by default because the big headline outages are the ones that also happen to have the messy communication story. This one is useful precisely because it's small enough to see clearly. 350 reports, one clean root cause, one specific 40minute gap between actually fixed and still telling customers it's broken. No noise to hide behind.
The size of the outage tells you how many people it touched. The size of the gap tells you how much you can trust thing that's supposed to tell you when it's over. Those numbers don't have to move together.
So, let's go back to where we started. Me refreshing a status page at 11:57 UTC. reading a countdown that told me I had 45 more minutes of this ahead of me. While according to AWS's own account, the outage had already been over for 39 minutes.
I don't have a tidy resolution for you here. I don't think I should manufacture one. AWS hasn't published exactly what the internal capacity constraint was. I don't know if they will. that's allowed to just sit there unresolved because it is in fact unresolved.
What I do want to leave you with is this. What you're monitoring tells you and what's actually true are not automatically the same thing. Most of the time the gap between them is small enough not to matter. Sometimes it's 40 minutes wide and during those 40 minutes an entire status page worth of decisions. Do we escalate? Do we page someone? Do we tell our own customers we're still down? All that gets made on stale information. That gap is exactly where trust gets spent or saved. And it's worth building systems and organizations that know the difference between fixed and confirmed fixed. And say so honestly while they're waiting to find out which one is true.
That's this episode of Chaotic Commits. I'm Joanne Skiles. And yes, I did in fact refresh that status page probably a dozen times in a row. Like reloading it faster would change the answer. So if you ever done the exact same thing during an outage that wasn't even yours to fix, this one's for you. New episodes wherever you get podcasts. No highlight reel, just to commit history. See you in the next