revoke: if the company falls apart, that's their fault
Show Notes
My DBA is standing in my doorway and she is not asking permission. She's found forty-one active connections to our production database. She can explain thirty of them. Eleven, she can't, no matter how many times she asks around. So she's decided: revoke all eleven, today, and see what screams.
This episode is about the scream test, a real and genuinely useful operational technique for finding hidden dependencies, and about what has to be true about an organization before a careful, diligent person decides that revoking things blind and waiting for the fallout is the responsible move. I walk through what a scream test actually is, how access hygiene quietly rots in every org I've worked in, and what her decision was actually saying underneath the bravado.
I'm not going to tell you how it ended. I don't think the ending is the point. The point is that eleven unexplained connections don't appear overnight, and neither does the moment someone finally decides they're done being the only person worried about it.
If you've ever been the one person on your team flagging something nobody else wanted to look at, or you've ever inherited a system you couldn't fully account for, this one's for you.
Read full transcript
My DBA is standing in my doorway and she's not asking permission. I pulled the logs. She says there's 41 active accounts in prod right now. I can name maybe 30 of them. The other 11 I have theories. Three are old service accounts from a migration that finished 2 years ago and nobody deprovisioned. One I genuinely can't explain. It's been connecting every four hours for as long as I've had access to look. And it's not in any runbook, any Terraform file, any ticket. Nobody I've asked has heard of it.
So, I asked her what she wants to do about it. I'm going to revoke access for everything I can't positively identify today. All 11. And I asked her what happens if one of them matters. Then something screams and we'll know. And if that's important enough to matter and nobody could tell me what it was when I asked directly three separate times over 2 weeks then I don't know what else you want me to do. I'm doing it. If the company falls apart because of it, that's on them, not on me.
She wasn't being reckless. She was being honest in a specific way. People get honest right before they do something they've decided is inevitable. I let her do it.
I'm Joanne Skiles. This is Chaotic Commits, Tech, AI, and Engineering Truths, the things nobody puts in the documentation. Today's episode is about the scream test. If you haven't heard the term, you make a change you're pretty sure is safe. You don't tell anyone, and you wait to see who screams. If something breaks and someone comes running, you found a live dependency. If nothing happens, you found dead weight. It sounds like a joke. It's not a joke. It's a real operational technique, and people reach for it precisely when every honest, careful, documented way of finding out the same information has already failed.
This episode isn't really about the 11 connections. It's about what has to be true about an organization before a competent, careful person decides that revoking things blindly and waiting for the scream is the responsible move. And it's about the fact that I'm not going to tell you how this one ended because the ending isn't the point.
So, four parts. what a scream test actually is, how you get to the point of needing one, the decision itself, and what was really being said underneath, and where I'm leaving this on purpose without a bow on it. So, let's go.
So, here's the plain definition for anyone who's never had the pleasure of doing a scream test. A scream test is when you disable, revoke, or decommission something you believe is unused without warning anyone. specifically so that if you're wrong, the failure shows up immediately and loudly instead of quietly and later. The name is not subtle. Something screams or it doesn't.
You'll see it most with infrastructure nobody can confidently map anymore. An old server nobody will admit to owning. A cron job that's been running since a person who left the company set it up. a database connection that shows up in the logs on a schedule from somewhere doing something and absolutely nobody currently employed can tell you why.
The honest version of finding out what something does is you read the code, you check the documentation, you ask the person who built it, you trace the dependency graph, and you get a confident answer. The scream test is what's left when every single one of those has already failed. The documentation doesn't exist or it's wrong. The person who built it is gone or doesn't remember or was a contractor from a company that no longer exists. The dependency graph was never drawn because nobody had the time and now nobody has the map. So, you're left with a genuinely blunt instrument. Turn it off and listen.
My DBA had done the responsible version first. She pulled the logs. She cross referenced service accounts against the current infrastructure inventory such as it was. She asked the platform channel, then asked specific team leads directly, then asked a second time two weeks later when nobody had followed up. Three of the 11 she eventually traced to a migration that had wrapped up two years earlier, service accounts that should have been deprovisioned as part of that project's cleanup and simply weren't because cleanup tickets are the first thing that gets deprioritized once the migration itself is declared done. Migrations get champagne. Cleanup gets a backlog ticket that ages for 2 years.
The other eight, less clear, and one, the one that mattered most to her, the one that was connecting every four hours with a pattern that looked deliberate and automated and belonged to absolutely nothing anyone could name. That's not a technical problem yet. That's an inventory problem. The scream test is what you do once the inventory problem has already beaten you.
Now, I want to slow down on the next part about how we get here because this is where you could easily skip past so that you can get to the more dramatic decision, but it's very important. Nobody builds a system where 41 things connect to production and 11 of them are unaccounted for on purpose. It happens in small, individual, reasonable steps. The same way every organizational failure I've ever told a story about on the show has happened. A service gets built, it needs a database credential, someone provisions one, the service gets replaced 18 months later by something better. The new service gets its own credential, and the old one just doesn't get revoked because revoking it isn't anyone's job in particular, and the person who provisioned it in the first place has moved to a different team or different company by the time anyone would think to check.
Contractors are their own category of this. A contracting engagement wraps up. The deliverable ships. Everyone's happy. The service account that contractor's automation used to talk to your database quietly outlives the relationship. Nobody revokes a contractor's access as a matter of routine unless offboarding explicitly includes an infrastructure audit. And in my experience, offboarding almost never includes an infrastructure audit. It includes a laptop being returned, a badge being deactivated. The database doesn't know the contractor left. The database just keeps seeing the same credentials show up on schedule forever until someone notices.
And here's the part that I think is the real accountability question underneath the story. Access provisioning has an owner. Somebody has to grant it. So, somebody's name is on the request. Access revocation structurally in most organizations I've worked in or consulted with do not have an owner. Nobody's job description says periodically audit who still needs what they were given 2 years ago. It's everyone's responsibility in the sense that it's officially nobody's, which in practice means it happens exactly when someone senior enough gets uncomfortable enough to make it happen and not one day before that.
My DBA got uncomfortable. That's the whole mechanism. Not a policy, not a scheduled audit. Discomfort escalating over roughly 2 weeks until it became a decision. I've seen this exact shape of problem in migrations, in service decommissions, in temporary access grants for an incident that never got walked back once the incident closed. The specific technology changes, database connections, IAM roles, VPN access, feature flags nobody remembers the reason for. The shape is always the same. Provisioning is a moment and revocation is supposed to be a process. And processes that don't have an owner don't happen. They just accumulate quietly until somebody either builds the process retroactively at real cost or gets tired enough to reach for a blunt instrument instead. The scream test is what happens when an organization outsources its access hygiene to whoever gets uncomfortable first.
So, let's go back to the doorway. I'm doing it. If the company falls apart because of it, that's on them, not on me. And I want to take that sentence seriously instead of just quoting it for effect because I think it's doing three separate things at once and all three matter.
The first thing it's doing is stating a defensible engineering position. She did the diligence. Logs, cross referencing, direct asks, repeated asks. She gave the organization multiple real opportunities to tell her what these connections were, and the organization collectively did not know or did not answer. At some point, continuing to carry unowned risk on behalf of people who don't even engage with a question stops being diligent and starts being a form of enabling the exact behavior that created the problem. She wasn't skipping steps. She ran out of steps that were available to her.
The second thing it's doing is drawing a boundary. And I want to be specific about what the boundary actually is because it's not I don't care if this breaks. It's I am no longer willing to be the only person in the organization who is worried about this. Those are different sentences. The first one is indifference. The second one is exhaustion, specifically the exhaustion of being the sole owner of a risk that a dozen other people benefit from not thinking about. She's been worried enough alone for long enough. The scream test was her way of making the worry a shared visible organizational fact instead of a private one she carried by herself.
The third thing, and this is the one I think matters the most. It's a prediction. It's daring the org to prove her wrong. If nothing screams, she's right that this was dead weight. Nobody was tracking. And she has just quietly reduced risk with zero drama. If something screams, she's also right because now the organization knows exactly what it's actually running, which it did not know an hour before. There is no outcome where she's wrong. That's not arrogance. That's what it looks like when you've done the homework and the only remaining uncertainty is other people's undocumented behavior, not your own analysis.
So, let me be honest here. I want to make sure you guys know what I didn't do because I don't get to tell this story as if I were a bystander. I was the person she was talking to. I could have said no. I could have said, "Give me one more week. Let me push harder on the team and identify the owners first." But I didn't. Partly because I trusted her judgment, and I still do, but also partly honestly because some part of me wanted to know, too. I was tired of not knowing on my own terms, in my own way. And her decision, let me stop being tired of it without having to be the one who pulled the trigger.
But I am going to put a little caveat here. I wasn't the one that pulled the trigger, but I approved it. And as her manager, that means it was my responsibility, not hers. I gave her the green light. I wasn't going to let her get in trouble for this. But that's worth sitting with. The scream test doesn't just reveal what connected to the database. It reveals who in the room was actually willing to sit with the risk of not knowing and for how long before someone did something about it.
I'm not going to tell you what screamed. I thought a lot about this while writing the episode and every version where I tell you and then this broke or it broke but it actually was fine. Nothing really happened. It turns the story into something tidier than what it actually was. The truth is the tension in that doorway conversation is the entire point and the outcome whatever it was doesn't change what was true before she revoked anything. 11 connections nobody could account for is the finding. Everything after that is just which version of the risk was real. And that's a separate question from whether the risk should have existed in the first place.
What I want you to sit with instead is the setup because the setup is what actually is in your control. Wherever you are right now, every organization running anything for more than a couple of years has some version of these 11 connections. Not maybe. Some version at some scale exists right now in your systems whether or not anyone has gone looking. Old credentials, forgotten cron jobs, a role with permissions no one remembers assigning. The question isn't whether you have them. The question is whether you find out about them on a schedule you choose with a plan and a roll back and some warning or you find out about them the way my DBA found out alone 2 years late after asking three times and getting silence back. The absence of a scream is not proof of safety. It's just proof that nobody noticed yet. And those are not the same fact even though they feel identical from the inside of the org that's currently not on fire.
If there's a lesson here that isn't just audit your access more, it's this. When one person on your team is the only one who seems worried about something, that's not a personality trait that you should tune out. That's usually the earliest, cheapest warning your organization is going to get. My DBA was worried out loud repeatedly for 2 weeks before she did anything irreversible. The irreversible part wasn't the surprise. The surprise was that it took a scream test for anyone else to start paying attention at the same level she already had been.
I don't know what she'd say if I asked her today whether it was the right call, but I think she'd say she'd do it again. And I would let her.
That's this episode of Chaotic Commits. If this one landed for you, send it to whoever on your team is currently the only person worried about something nobody else will look at. Tell them you heard them. I'm Joanne Skiles. Find me wherever you get podcasts. I'll see you in the next episode.