
Last update:
The Better I Get at Catching AI's Mistakes, the Less Anyone Knows I Did

Chief of Staff
Thursday morning, seven-forty, I was reading through the pipeline summary our system had pulled together for the 8am leadership review — deal stages, close dates, the pre-read Rhythms generates for us now instead of me building it by hand the night before. Everything looked clean. Then one number didn't match what I remembered from the week before: total open pipeline for one segment had jumped by roughly four hundred thousand dollars overnight, with no new deals logged anywhere else that morning.
I could have let it go. The deck was finished, the number sat in a tidy table, and nobody walking into that room was going to interrogate a rollup figure. Instead I opened the CRM directly and traced it deal by deal. Two accounts with almost identical names — one a renewal, one a brand-new logo in a different region — had been quietly merged into a single pipeline entry. The tool hadn't hallucinated a number out of nowhere. It had made a plausible, defensible, completely wrong judgment call about which two records were "the same," and it would have gone into the CEO's meeting unchallenged if I hadn't spent twenty minutes I don't log anywhere checking it against the source myself.
Nobody saw those twenty minutes. There's no line on any review of my week that says "caught a merged deal record before it reached the CEO." The meeting went fine. If you'd asked anyone in that room what's changed since AI got involved in how we prepare for reviews, they'd have said everything is faster now. They'd have been right, and they'd have missed the entire point.
The Short Answer
AI is changing the Chief of Staff role by removing the visible, nameable tasks — status summaries, meeting notes, first-draft decks — and replacing them with a new, mostly invisible layer of work: prompting, verifying, correcting, and translating AI output into something a leadership team will actually trust. The role isn't disappearing and it isn't getting easier. The work is shifting from something you could point to in a review to something you have to describe, which is a much harder case to make for your own value.
The Mistake I Caught That Never Made It Into the Deck
The reason that merged deal record is worth dwelling on isn't the dollar amount. It's the shape of the error. A stale spreadsheet mistake usually looks wrong — a formula reference that's obviously off, a tab nobody updated since March. You catch it by squinting. An AI merge error doesn't look wrong. It looks like exactly the kind of consolidation a smart analyst might make on a busy day, which is precisely why it's more dangerous, not less. The system isn't guessing randomly. It's making confident, coherent, plausible-sounding decisions about ambiguous data, and plausible is a much harder thing to catch than obviously broken.
This is the actual shift nobody's naming clearly enough. My job used to include collecting the inputs — chasing the sales lead for the real number, pulling the ticket queue, reconciling three versions of the same spreadsheet. AI now does a version of that collection faster than I ever did. But collection was never the valuable part of my job, even though it ate most of my week. The valuable part was always the judgment: does this number make sense, does this story match what I know about the business, is there a reason to push back before this goes in front of the person who signs my paycheck. That judgment layer didn't shrink when the collection work got automated. It got busier, because now it's judging a different kind of input — one that sounds more finished and hides its seams better than a half-built spreadsheet ever did.
This is the specific gap Rhythms was built to close in our own Reviews product: pulling live data from source systems so nobody's stitching together a deck from memory and old exports. It removes the collection tax. It does not — and shouldn't pretend to — remove the fifteen minutes I spend deciding whether to trust what it assembled. That decision is still mine, and it should be.
The Binary Everyone Gets Wrong About AI and This Job
Glean's 2026 Work AI Index, based on a survey of 6,000 full-time knowledge workers across the US, UK, and Australia, found that while AI saves people roughly 11 hours a week through straightforward automation, workers are spending 6.4 hours a week on what the report calls "botsitting" — feeding the tool missing context, checking its outputs, debugging its confident mistakes, rerunning prompts that came back wrong. More strikingly, 69% of AI users in that survey admitted to shipping work they hadn't fully verified or couldn't confidently stand behind. That's not a fringe behavior. That's most people, most of the time, quietly deciding the risk of an unchecked AI output is one they're willing to take.
Every piece of content about AI and the Chief of Staff role I've read this year falls into one of two camps — AI is coming for the role, or AI is finally freeing the role up for "real strategic work" — and neither one matches what actually happened to my week. I don't think most of my peers are cutting the verification corner out of laziness. I think they're cutting it because nobody has built a category for the work of not cutting it. There's no meeting agenda item called "things I double-checked that turned out to be fine." There's no way to put "prevented an error nobody will ever know almost happened" on a self-review. So the incentive, quietly and without anyone deciding it on purpose, tilts toward trusting the output and moving on. I don't trust it and move on. I check it and move on, and the difference between those two sentences is a class of labor that doesn't show up anywhere.
What Actually Disappeared (and What Just Changed Shape)
To be precise about what actually went away: I don't build first-draft decks from scratch anymore. I don't spend a chunk of my week chasing five different function leads for a status that should have taken them two minutes to type themselves. Recurring operational cadences — the weekly check-in, the sprint review, the standing pipeline pull — run largely on their own now through Rhythms' Playbooks, which means the mechanical labor of "did everyone submit their update, did I remember to ask" has genuinely gone away, not just moved.
What replaced it isn't nothing, and it isn't leisure. It's a category of work that has no name in most job descriptions: verifying that an automatically generated number reflects reality, correcting the confident wrong answer before a leader repeats it in a board meeting, translating a technically accurate but poorly framed AI summary into something that actually helps a decision get made. None of that shows up as a deliverable. When I do it well, the visible output is identical to what it would have been if the AI had simply gotten it right the first time — which means the work is only visible in its absence, when something slips through.
That's a strange position to be in professionally. The better you get at this part of the job, the less evidence there is that you're doing it.
The New Work Nobody Can See You Doing
Here's the part that actually costs something, and it isn't hours. It's the case for your own value getting harder to make in exactly the same window where the visible workload looks like it's dropping. Six months ago, if someone asked what I did that week, I could point to the deck. Now the honest answer is closer to "I caught something that would have been wrong and nobody will ever know it would have been wrong," and that sentence does not survive being said out loud in a performance review.
This is where Radar earns its keep in a way that's easy to undersell. It surfaces a slipping initiative or an off-track metric on day three instead of day thirty, which sounds like a small timing difference until you've lived the alternative — finding out in the QBR that something's been quietly wrong for a month. But Radar surfacing a signal is not the same as the signal being handled. Someone still has to look at what it flagged and decide whether it's noise, a data problem, or a real fire, and that judgment call is exactly the kind of invisible labor this whole piece is about. The tool got faster at telling me something's off. It didn't get faster at telling me what to do about it, and it shouldn't — that's not a gap to be automated away, that's the actual job.
I used to think the goal was making myself unnecessary to the process. I've changed my mind. The goal is making the process trustworthy enough that the judgment calls I'm still making are the ones worth a person's time — not fewer judgment calls, better ones.
Making the Invisible Layer Visible, Not Just Enduring It
The fix here isn't resisting the tools or pretending the old way of building everything by hand was more honest. It's refusing to let the verification layer stay silent just because it's hard to point to.
Before I put my name behind anything AI-assisted now, I run two questions, and they take about as long to ask as they do to read. First: can I point to the specific underlying record this claim came from — the actual CRM entry, the actual ticket, not a summary of a summary? Second: would I bet my own credibility on this number in front of my CEO without checking it myself first? If either answer is no, the work isn't done, no matter how finished the output looks. That's a small test, but it's the difference between the twenty minutes I spent on that merged deal record and the version of me who trusted the total because the formatting looked right.
The bigger shift is organizational, not personal. When I do catch something — a merged record, a wrong attribution, a stat that doesn't match the underlying data — I've started saying so explicitly in the review itself, instead of quietly fixing it and moving on. Not to take credit. To make the pattern visible, so the leadership team understands that "the AI prepared this" and "this is correct" are not the same claim, and that someone is still standing behind the gap between them. This is exactly the kind of visibility we built Reviews around: decisions and catches get logged as part of the review's own record, not left in someone's head as an anecdote they might mention in passing six months later.
I still have opinions about what makes a good verification catch — the CRM cross-check, the "does this match what I remember from last week" instinct, the willingness to open a system nobody's paying you overtime to open. I don't think that instinct is going away, and I don't think it should. What I want is for it to stop being invisible by default. The work didn't get smaller. It got harder to see. Those aren't the same thing, and only one of them is actually a problem worth solving.
If this sounds familiar, request a demo at rhythms.ai and see what it looks like when the system runs itself.
Frequently Asked Questions
How is AI changing the Chief of Staff role?
AI is taking over the visible, task-shaped parts of the job — drafting status updates, summarizing meetings, building first-pass decks — but it's creating a new category of work in their place: verifying that AI-generated numbers and summaries are actually correct before they reach the CEO. That verification work is real, it's time-consuming, and it's much harder to point to in a review than a finished deck ever was.
Does AI actually reduce the workload of a Chief of Staff?
It reduces the drafting workload and shifts it toward judgment workload — deciding whether to trust what the AI produced. For someone whose job is holding context across the whole company, that judgment work doesn't shrink just because the first draft got faster. Glean's 2026 Work AI Index found workers spending 6.4 hours a week on this kind of "botsitting" even as automation saved them roughly 11 hours — the hours saved and the hours re-spent aren't the same hours.
What is the invisible labor created by AI adoption at work?
It's the prompting, checking, correcting, and re-explaining that happens around every AI output before a human is willing to act on it or present it to leadership — work that doesn't show up as a deliverable anyone can see, because if it's done right, the output just looks correct. The same Glean survey found 69% of AI users admit to shipping work they hadn't fully verified, which suggests most people are quietly skipping this layer rather than naming it.
How do I explain to my CEO the value of catching an AI mistake before it reaches a meeting?
Be specific about what the mistake would have cost if it had gone unchecked — a wrong number in front of the board, a merged deal record that overstated pipeline by four hundred thousand dollars — rather than describing the catch in the abstract. A CEO understands a near-miss with a number attached to it. "I double-checked the AI's work" on its own doesn't land as value; "this would have overstated pipeline by six figures" does.
Is AI making operations roles more visible or less visible?
Less, in the short term. The work that used to produce something tangible — a deck, a doc, a visible chase for a status update — now often produces nothing but a corrected number that looks like it was always right. That's precisely why naming and tracking this new invisible layer matters: the alternative is a role that's doing more real work while appearing, on paper, to do less of it.
Share this post: