debug: how would you know
A non-technical friend forwards a wire story about an AI researcher resigning from Anthropic and asks 'should i be worried' — and the honest answer takes twenty minutes to write, because neither that article nor the bestselling AI doom book making the same rounds ever answers the load-bearing question: how would you know, specifically, if the risk were starting. This episode takes AI safety fear seriously without feeding it: why stopping any single lab just hands pace-setting to whoever is left racing, why real documented failures like Air Canada's hallucinated refund policy, COMPAS recidivism bias, and the Dutch childcare-benefits algorithm all trace to fixable engineering gaps rather than malice, and why mechanistic interpretability and deception-probe research — early and unfinished — are the field's honest attempt at the missing instrument panel.
Show Notes
A friend who has never worked in tech sent me a wire story about an AI researcher quitting Anthropic — with a quote about the companies playing with our lives — and one line underneath it: "Should I be worried?" It took me twenty minutes to answer a text.
Not because the article was wrong. Because it was the third thing that week shaped exactly the same way, including a bestselling AI doom book I'd just finished, and none of them ever answer the one question that actually matters: how would you know, specifically, if the risk were starting.
This episode takes AI safety fear seriously without feeding it. I walk through why stopping any single AI lab doesn't remove the risk from the world, why real documented AI failures — Air Canada's hallucinated refund policy, biased recidivism scoring, a Dutch government algorithm that wrongly flagged tens of thousands of families for fraud — all trace back to fixable engineering gaps rather than malice, and why mechanistic interpretability and deception-probe research, early and unfinished as they are, are the field's actual, honest attempt at building the instrument panel that's missing from the discourse. If you've ever gotten a scary AI headline forwarded to you by someone who doesn't work in tech and didn't know what to say back, this one's for you.
Read full transcript
Wednesday night. And I'm living the dream. The kids are asleep and I'm in my bed playing the Sims. More precisely, I'm looking at my phone, running through the different combos to splice plants, and I get a text.
It's a friend. Not in tech. She's never worked in software a day in her life. Her interests do not overlap with mine at all in the tech world. We have Broadway to bond on instead. So she does not read the newsletters I read. She has no idea who Evan Hubinger is and really no reason to. She sends me a link. Wire story, Tribune paper, AFP byline, a headline built around a quote: an AI researcher quit Anthropic and said, on the record, that the companies building this stuff are playing with our lives.
Under the link, one line. "Should I be worried?"
I sat there, cozy in bed, staring at that message for longer than I want to admit.
It wasn't that the article was new information to me. It was that this was the third time that week I'd run into something with the exact same shape. Different source. Different format. Same feeling underneath it. Like a hum of a song you know on the tip of your tongue but can't place until somebody names it for you. I remember thinking, lying there in bed, there's probably a German word for this. Turns out there is. Zukunftsangst. I can't really say it that well, and I won't say it again, but it's the fear of the future. Not fear of a specific thing. Fear of the shape of something coming before you even know what it is or where it lands.
I didn't answer her right away. I paused my Sim (she's unreliable on her own) and just sat there rereading the article on my phone a second time, then a third, looking for the part I apparently skimmed past that would tell me what to actually say back.
And I'm going to tell you why I took so long, and it's not the reason you'd guess.
Welcome to Chaotic Commits. I'm Joanne Skiles. Sixteen years writing software and a PhD. My dissertation was vehicular ad hoc networks, cars talking to each other and predicting routes before there was infrastructure to support any of it. Half my actual dissertation wasn't working the routing protocol. It was building the logging and metrics just to prove the routing protocol was doing what I claimed it was doing. Long before this showed up in my texts as a scary headline, my day job was the exact same question: how do you actually know what a system is doing versus what it's telling you?
So I'm going to do something a little different this episode. I'm not going to try to talk you out of being worried, because some of what's in that article is real and I'm not in the business of telling people their gut reaction is stupid. What I want to do instead is take the fear apart with the same tools I'd use on any other unverifiable system, and hand you back the one piece of it that's actually important. Because right now, it's buried under a pile of other pieces that frankly aren't. And that pile is exhausting.
And I will tell you what I sent back to my musical loving friend. It took me twenty minutes to write two sentences, which for anyone who knows me is a personal record in the wrong direction.
Let's get into it.
So let's not undersell what's actually sitting in that article, because it isn't nothing.
Jacob Coxon spent three years doing pretraining research, first at OpenAI, then at Anthropic. He quit. Publicly, on X, no hedging. He said neither company is acting responsibly, that both are racing towards self-improving superintelligence and gambling with all of our lives to get there first. He said it wasn't a marketing stunt, which tells you something about the world we're in right now: that a person has to specify that out loud before anyone will take the resignation seriously.
And then Evan Hubinger backed him up publicly. Hubinger isn't some guy with a newsletter. He's an actual safety researcher inside Anthropic, the name behind the sleeper agents work if you follow interpretability research at all. And he put a number on it. He said he personally puts the odds of AI killing everyone at more than ten percent within the next decade.
That's not some dude with a blog. That's not a guy selling a book on a press tour. That's someone inside the building saying it on record with his name attached.
If you felt something drop in your stomach, you're not being irrational. So sit with that for a second. Because the rest of this episode is not about talking you out of it. It's about telling you what question you should be asking instead of the one that message asked me.
So here's what actually got me. That article was not the first thing that week shaped like this. I just finished a book, If Anyone Builds It, Everyone Dies, by Yudkowsky and Soares. Bestseller, blurbed everywhere, a title doing exactly what a title is built to do the moment you see it on the shelf. And the thing that actually bothered me about that book wasn't its conclusion. It's that neither author works on a frontier model. Their case is built almost entirely on decision theory and thought experiment, not on anything resembling engagement with how today's actual systems behave, fail, or get caught failing in the wild.
Then the Tribune piece lands in my texts. Real people, real quotes, a real number, but strung together with almost no reporting in between them. Coxon's tweet, Hubinger's tweet, Hubinger's follow-up clarifying that he means future systems, not the ones running today, a line pulled from OpenAI's chief scientist, a mention of a Senate bill somewhere in committee. Five sources, zero connective tissue, and the piece reads less like an investigation than like someone screenshotted a group chat. And nobody in that piece, not the reporter, not an editor, thought to ask Hubinger the one obvious next question, which happens to be the same question the book never answers either.
How would you know.
Not "will it happen." Not "is it possible." How would you know, specifically, in advance, if it were starting. What would you actually see on a screen somewhere. What's the mechanism, the readout, the alarm that goes off. Neither the book nor the article gets within shouting distance of that question. And once you notice it's missing, you start seeing the hole everywhere you look. That's the Zukunftsangst engine right there. It's dread, with no instrument panel. You can't do anything with a feeling that has no gauge attached to it except carry it around.
So now here's something a little uncomfortable.
Say Anthropic stops tomorrow, shuts down, doors locked. Hubinger's ten percent never gets tested because the company that produced the number simply exits the field entirely.
Does the risk go away?
No. Someone else builds it. Maybe slower. Maybe with a fraction of the safety research budget Anthropic currently spends. Maybe in a lab with zero interest in ever publishing a sleeper agents paper for anyone outside the building to read. One company stepping back doesn't remove the capability from the world. It just removes the most cautious player from the room and hands the pace-setting to whoever's left standing in it.
I keep coming back to nuclear weapons for this one, and I know how that sounds, but stay with me, because I'm not claiming AI development and nuclear proliferation are the same thing. I'm making one specific, narrower comparison: when a dangerous capability already exists and more than one actor is capable of building it, unilateral restraint by the most careful actor doesn't remove that capability from the world. It just changes who develops it, who controls it, who gets to set the norms around it.
Here's where it holds. The US has more nuclear warheads than any other country on earth. I don't think many people think that's a good thing. There's no defense budget line item anywhere that reads "we should acquire more of these because they're nice to have." And yet nobody is unilaterally disarming, because getting rid of your own stockpile doesn't make nuclear weapons disappear from the planet. It just changes who holds the most of them and who gets to set the norms everyone else has to live inside of.
Same narrow logic here. It's a little arrogant to think one company stepping back solves anything. The bad actors don't stop being bad actors because the most conscientious lab in the industry left the table. Arms control treaties never worked by asking the most careful country to unilaterally hand over its weapons and hope everyone else followed along out of good manners. They worked, when they worked at all, through verification. Inspectors. Instruments. A way to check what the other side was actually doing instead of taking their word for it. Nobody ever got safer by convincing the most cautious party in the room to leave.
So you can shut down Anthropic. You can shut down every American frontier lab. You haven't shut down the idea, you haven't shut down the capability, and you haven't shut down every country on earth.
That's why the answer "stop AI" doesn't work. No matter how many tweets, blogs, and bestsellers get written demanding it, the only fight that was ever actually available is making sure whoever's racing ahead, all of them, every lab in the race, gets forced to answer the how-would-you-know question before they ship. That's not the comforting version of this argument. It's the only version that survives contact with the fact that nobody, anywhere, is slowing down.
So let's talk about what "how would you know" looks like when it actually works. Because it already has, more than once, and none of these are hypothetical.
Air Canada's chatbot invented a bereavement fare refund policy that never existed, told a real grieving customer he qualified for it, and a tribunal held the airline to its own bot's word. COMPAS, the recidivism scoring tool half the courts in the country leaned on for years to help decide bail and sentencing, turned out to be carrying a measurable racial bias baked into its scoring, and it took investigative reporters rebuilding the model's math themselves to prove it, because the vendor treated the internals as a trade secret nobody got to inspect. And in the Netherlands, a tax authority's fraud-detection algorithm used dual nationality as a risk factor for flagging childcare benefit fraud, wrongly branded tens of thousands of families as fraudsters, mostly immigrant families, forced them into ruinous repayments, and in the worst cases fed directly into custody proceedings that took children out of their homes. It brought down the sitting government.
Every single one of those is a genuinely bad outcome. I'm not standing here telling you these systems don't screw up. They clearly do. Sometimes in ways that wreck a real person's actual life. Sometimes in ways that topple a government. But look at what all three of these have in common. Someone, eventually, could start pointing to the mechanism. Not "the algorithm wanted to hurt someone." A specific, traceable, fixable input, a specific line of logic, a specific weight on a specific variable that nobody had stress-tested against the people it would land on. In every case, the failure got a name and an address when someone was allowed to open the hood. That's the engineering problem I keep coming back to. And it is the opposite lesson from the one the fear-mongering version of this conversation wants you to walk away with.
The problem was never that these systems are malicious. The problem is we keep building things that can act faster than we build the instruments capable of reading them. That's an engineering problem. It's not a horror movie premise. And treating it like one is how you end up unable to tell the difference between a system that needs a better dashboard and a system that's secretly plotting against you.
So let's go to that instrument panel. This is the part where the field has quietly moved a lot faster than the discourse around it has.
A smoke detector doesn't promise your house will never catch fire. It promises you'll know before the smoke fills the hallway. Nobody stands in their kitchen complaining that the smoke detector isn't also a fire extinguisher. That was never its job. Its job is that fifteen seconds of warning that let you act instead of finding out after the fact. That's the entire category of thing I'm about to describe. And it's why I don't think the honest answer to "should I be worried" is either "trust us" or "shut it all down."
Interpretability research broadly has leaned for years on explaining an output after it already happened. Feature importance scores, saliency maps, the SHAP bar chart somebody drops into a slide deck to make a black box feel briefly less black. Useful, but still working backward from what the model already said. Mechanistic interpretability is a narrower, more specific effort sitting inside that broader world. And one direction of it, the one Anthropic's own published circuit-tracing work is a visible example of, is trying to reverse engineer the actual internal computation, the causal circuit that produces a given behavior in the first place, instead of a plausible-sounding story bolted on afterward.
There's active research into deception probes: work investigating whether a model's internal representations can be read for warning signs of deception or strategically concerning behavior, not just taken at their word for what it says out loud. It is not a finished tool sitting on a shelf. It's early, it's hard, and a lot of it doesn't work yet, but it's the closest thing anyone has to a real answer to the question nobody in the Tribune article thought to ask Hubinger. Not "trust us." Not "shut it down." Try to build something that would actually tell you before the smoke fills the hallway instead of after.
Now, I'm going to be careful here, because there's a version of this argument that oversells the fix just as badly as the article undersold the question. Detecting dangerous behavior is one problem. Detecting a dangerous internal state is a harder one. Telling actual deceptive alignment apart from a model just doing a weird model thing is harder still. Catching any of that early enough to matter is its own separate problem, and none of it holds up if the measurement stops being trustworthy the moment the thing you're measuring gets smarter than the tool doing the measuring. Those are five hard, unsolved problems stacked on top of each other, not one solved problem wearing a research-agenda name tag.
An instrument panel doesn't make the airplane safe. It makes it possible to know when something is going wrong. That's the whole claim. It's a smaller claim than it sounds like, and it's also the only one that's actually true. And even a working gauge doesn't do anything by itself. Somebody still has to be watching it, believe what it says, and be willing to act before the smoke fills the hallway. Whether institutions actually do that when the number comes back bad is a different question than whether we can build the gauge. It's probably a different episode.
That's not a hypothetical someday sitting on a whiteboard somewhere. That's a research agenda with actual people and actual funding behind it, happening in parallel with every scary headline you're going to keep seeing this year. It is not, on its own, the whole answer. It's the first honest step toward one.
So, back to the text message.
I sat on it for about twenty minutes before I answered her. Like I said, anyone who knows me knows that's roughly nineteen minutes longer than I usually take to respond to anything, including actual emergencies. But here's what I finally sent. We don't yet have a reliable instrument panel, the kind you can trust the way you can trust a smoke detector, that could tell us with real confidence a catastrophic capability or deceptive behavior was actually starting to show up. We have pieces. Evaluations, interpretability work, behavioral tests, the early probes I just walked you through. None of it adds up yet to something you'd stake the outcome on. That gap, not the ten percent itself, is the actual thing worth losing a little sleep over. And closing that gap is exactly what a real chunk of this field is working on, right now, today, to do.
Be worried about the right thing. Not the plot of a horror movie we've already written in our heads. The instrument panel we haven't finished building yet. Worry gets you somewhere when it's pointed at something you can actually check on. Pointed at everything, it just sits in your stomach and does nothing.
That's this episode of Chaotic Commits. I'm Joanne Skiles, who apparently needs twenty full minutes to answer a text but will happily talk at you for half an hour, unprompted, about mechanistic interpretability. New episodes wherever you get podcasts. No highlight reel. Just the commit history.