fix: employee scores don't reset for maternity leave

July 24, 202615 min

Show Notes

Six weeks into parental leave, your Slack status turns green by accident, and a coworker texts asking if you're back. You're not back. You're up at 2am with a newborn. But somewhere, a dashboard is still counting your tokens, your commits, your keystrokes, and for those six weeks, the number is zero.

Twenty-six Meta employees just filed over exactly that math. They allege the scoring system used to decide who got laid off pulled from keystroke and activity monitoring data, AI token-usage dashboards, and algorithmically assisted performance rankings, and that the resulting score structurally could not be earned by anyone on protected medical, parental, or family leave. Not unlikely. Structurally impossible.

I'm not a lawyer and this isn't a legal deep dive. I want to talk about what a token-usage dashboard actually measures, what it silently assumes about the person generating those tokens, and how objective quietly becomes punishes anyone without a keyboard open eight hours a day. Nobody sat in a room and decided to target people on leave. Nobody had to.

If you've ever built, or been ranked by, a system that called itself objective, this one's for you.

Read full transcript

You're 6 weeks into parental leave when your Slack status turns green by accident. Some app you forgot to close. And within an hour, a coworker texts you and says, "Hey, are you back? You're not back. You're on the couch at 2 a.m. with a kid who won't sleep." And the last thing on your mind is your commit graph, but somewhere a dashboard is still running. It doesn't know you're on leave. Or maybe it knows technically. There's a field somewhere that says leave start date, end date, but that field isn't the one anyone's actually looking at. The one they're looking at is the one that counts. Lines committed, tickets closed, tokens spent talking to whatever AI tool your company plugged into everyone's IDE that quarter. And for 6 weeks, that number is zero. Not low, zero.

You come back in March, nobody says anything. You catch up, you ramp back up, everything feels normal enough, and then in May, you get a meeting on your calendar with a title that doesn't tell you anything and a name attached that you've never had a one-on-one with. You already know. What you don't know yet, what you won't know for months is that the number that put you on that list wasn't really about March or April or how you performed once you were back. It was still carrying those six weeks. The system did the math it was built to do. It just didn't know there was a version of that story where the zero wasn't a performance problem. It was a person having a baby.

Nobody sat in a room and decided to target people on leave. That's actually the whole point of this episode. Nobody had to.

I'm Joanne Skiles. This is Chaotic Commits. Tech, AI, and Engineering Truths. The things nobody puts in the documentation.

That scene isn't hypothetical. It's the shape of a claim 26 people are making right now in an actual filing against an actual company. And I want to be upfront that this isn't a legal episode wearing a tech costume. I am not a lawyer. I'm not going to explain arbitration to you, and I'm not going to pretend I understand the legal standards well enough to weigh in on them. What I can talk about is the system. What it measured, what it assumed about the people it was measuring and what happens to somebody standing in the exact spot the system was never built to see.

So four parts. What actually got filed against whom? What one specific number in that system quietly assumes about a person's life who ends up holding the bill for what the system got wrong. And what it would actually take to build one of these that doesn't do this. So, let's go.

Just recently, 26 Meta employees filed to challenge how they were selected for a layoff. Separations start at July 22nd, which if you're doing the math against when I'm recording this, it's basically right now. The allegation isn't the usual layoff complaint. It's not my manager didn't like me. It's more specific and honestly more interesting from where I sit. They are alleging that the process ranking them pulled from keystroke and activity monitoring data from dashboards tracking usage of company's internal AI tools, specifically how many tokens people were spending talking to the models, and from performance rankings that were algorithmically assisted rather than written by a human who actually watched them work all year. And their core claim is that however the scoring worked, it structurally could not be earned by someone who was out on protected medical, parental, or family leave during the window it measured. Not that it was unlikely, not that it was harder, that it could not happen. The math has no slot for this person was having a baby or this person was in the hospital. It only has a slot for the number that came out the other end. That's the part where I want to spend the rest of this episode on, not whether the lawsuit wins, whether the number itself was ever capable of being fair in the first place.

Okay, before I get into what's actually wrong with this, I want to steelman it because if you ever sat in a room where layoffs were being planned, and I have, you know, the thing nobody says out loud, but everyone is thinking. Nobody wants to be the manager who decides who loses their jobs. Trust me, that is a terrible position to be in. You're looking at a spreadsheet of people that you know personally and you have to rank them. And whatever you pick, someone is going to say it was personal. Someone's going to say you kept your favorites. Someone's going to say you didn't like them or you didn't get them or you played favorites with the people who reminded you of yourself at 25.

So the pitch for something like this practically writes itself. What if we didn't have to trust a manager's gut? What if we pulled real data? Commits, activity, output, usage, numbers, things that you can point to, things that don't have a bad day or a grudge or a blind spot. And I get why that's appealing because managers have blind spots. I have blind spots. Performance reviews have been unfair to people for as long as performance reviews have existed. And a lot of that unfairness runs along exactly the lines you expect. Who the manager likes, who looks busy, who's good at hallway visibility instead of actual work. So an engineering org looks at the mess and says, "Let's build something more defensible." Something we can point to later and say, "We didn't choose. The data chose." That's not a cartoonishly evil thought. That's a team trying to remove themselves from a decision they don't want to be making with their gut alone.

The problem is what happens next. Because once you build the system, you have to decide what counts as a signal. And that's the part where this whole thing quietly stops being neutral.

And if this sounds familiar, that's because we've run the exact experiment before, just with worse instrumentation. GE spent about a decade and a half ranking employees on a forced curve. And there's many companies that still do this. I've worked for them. And what they do is they cut the bottom 10% every year, call it a vitality curve, and eventually walk away from it, ranking people against each other on a curve. And that mostly taught people to sabotage each other instead of actual competition. Enron ran something close to that too, twice a year. And we know how that story went. The pitch was identical then to what it is now. Get the subjectivity out. Make the cut defensible. The only thing that actually changed is the tooling got a lot more precise. Nobody's manually stack ranking a spreadsheet in a conference room anymore. The model does it continuously off data. Nobody generated it with a ranking in mind.

Let's sit with that dashboard for a second because I think it's the part everyone skips past on the way to the more dramatic word which is layoff. A token usage dashboard in a company that's rolled AI coding tools out to basically every engineer measures one thing directly. How many tokens got spent talking to the model? That's a raw signal. Everything else stacked on top of it is inference. The inference runs like this. More tokens spent means more prompts sent. More prompts sent means more engineering work attempted. More work attempted, especially if it lines up with commits or tickets closed, means more output. And output in a performance review has always been treated as a reasonable stand-in for contribution.

Look at what the chain quietly requires to be true, though. It requires the person is at the keyboard. It requires that they're at the keyboard something like 8 hours a day, 5 days a week for the entire measurement window. It requires that a season of a person's life where they're not generating tokens because they're recovering from a surgery or feeding a newborn at 3:00 a.m. or sitting in a hospital room with a parent gets read by the system as identical to a season where that same person just decided not to bother. The dashboard doesn't have a field for not present because pregnant. It has a field for zero and zero sitting inside a ranking algorithm next to 40 other people's very much not zero numbers doesn't get read as context. It gets read as a rank.

Think about a baby monitor for a second. It's a weirdly good comparison. Okay, a baby monitor only picks up sound. If the baby's asleep and quiet for 6 hours, the monitor doesn't say anything is happy in the room. It just doesn't have anything to report. Nobody looks at a silent baby monitor and concludes the baby stopped existing. But that's exactly the leap this dashboard makes with a zero. Silence gets read as absence of a person, not absence of activity. And every ranking downstream inherits that mistake.

Objective in this framing doesn't mean fair. It means the humans who built the system got to stop personally weighing anyone's circumstances and handed that weighing job to a number that was never actually built to weigh anything. It just counts.

And this isn't the first time this exact shape has shown up this year. About a month before this filing, Workday was already in its own AI hiring discrimination case, over screening tools that filtered candidates before a human ever looked at them. I'm not going to tell you how that case should resolve. I don't know enough to say. I'm just telling you it's twice in a few weeks that the shape is identical. A company points at a model and says the model decided and somebody gets asked to prove that's actually true.

And that's the part I keep circling back to. Somebody somewhere built this scoring system. And I'd bet they built it with the same intentions as that steelman a few minutes ago I talked about. Get the messy, biased, gut feeling part of layoffs out of human hands. Replace it with something you can point to and defend. Whoever built that system got a real win out of it. They got to walk into a room and say, "We have a defensible data-driven process." Now they got to take themselves out of the accusation of playing favorites. That's an actual efficiency game. It's the same trade every system on the show ends up chasing eventually. Less judgment required, more throughput, a cleaner story for whoever has to stand in front of the board and explain why the layoffs were fair.

But that win doesn't come from nowhere. It gets captured by whoever built the system and whoever signed off on it. The cost of everything the system couldn't account for gets paid by somebody else entirely. In this case, 26 people now in arbitration trying to explain after the fact that the zero on their record wasn't a performance problem. It was a leave date. Nobody sat in a room and decided to target people on leave. Nobody had to. The system did that on its own quietly as a side effect of being built to remove exactly the kind of judgment that would have caught it.

Now, this is normally where I tell you the fix is complicated, but it isn't. Or at least the first step isn't. Every one of these scoring systems is built on a window, a start date and an end date for the period being measured. That window already exists somewhere in a database because payroll and HR systems track leave down to the day. That's literally how they calculate benefits and return dates. The information needed to solve this isn't missing. It's just not connected to the thing making the ranking decision.

A system that actually accounted for this wouldn't wait for someone to appeal after the list gets made. It would exclude the leave window from the denominator before the ranking ever runs, not flag it for a human to reconsider later. Actually remove those weeks from the math. The same way you exclude a server that was intentionally taken down for maintenance from the uptime calculation. Nobody calculates five nines of uptime by counting the scheduled downtime against the system. We already know how to do this. We just haven't decided that it applies to people.

I want to be specific about why the after the fact version doesn't count as a fix. If the correction only shows up in an appeals process, the burden of proof has quietly shifted onto the person who was on leave. They have to notice they got flagged, understand why, and make a case that their zero doesn't mean what the system assumed it means. That's not a technical gap anymore. That's the same accountability question from earlier in this episode, just moved one level down into how the exception gets handled instead of who built the rule.

None of this requires new technology. It requires somebody deciding before the ranking runs. That a person's protected leave isn't a data quality problem to be handled on appeal. It's an input the model was never supposed to see as a zero in the first place. That's a design decision. Somebody makes it or somebody doesn't. Either way, it's a choice, not a limitation of tooling.

That's this episode of Chaotic Commits. If you build anything that scores people, even loosely, even as one input among a bunch of others, go find out what actually happens to that score when someone's on leave. Don't assume the field exists just because it should. Go look. I'm Joanne. New episodes every Friday, wherever you get podcasts. I'll see you there.