Last update:

The AI ROI Everyone Promised Isn't Showing Up in the P&L. Here's Where It Went.

Rhythms

Rhythms

Rhythms

Our FP&A lead put two numbers on the same slide in last month's leadership review, and the room went quiet in a way that had nothing to do with the numbers being bad. The first number was our AI tooling spend for the quarter — a forecasting copilot, a drafting assistant for the GTM team, a couple of smaller tools stitched into support workflows. Real money, approved, tracked against a business case we'd all signed off on. The second number was supposed to be the payoff: the efficiency gain we'd modeled when we bought the tools. It was flat. Not down. Not dramatically up. Flat, the way a line looks when nothing happened.

Nobody in that room thought the tools didn't work. Every person using them would tell you, unprompted, that their individual week got easier. The forecasting copilot really does draft a passable first pass in minutes instead of an hour. The drafting assistant really does clear a blank-page problem faster than a person staring at one. And yet here was a slide showing that none of it had moved the number my CEO actually cares about. Someone finally said the thing everyone was thinking: "So where did it go?"

I didn't have a clean answer that day. I have one now, and it isn't the answer I expected to find.

The Short Answer

AI's promised operational ROI isn't missing — it's being reabsorbed by the same coordination layer it was supposed to eliminate. Tools got faster at producing information, but the verification, reconciliation, and re-explaining that happens around every review and report didn't shrink to match, so the hours AI saved quietly get spent again downstream. Closing the gap requires redesigning the review cadence itself, not adopting a better model.

The ROI Isn't Missing. It's Being Reabsorbed.

The instinct, staring at a flat line like that, is to conclude the tools underdeliver. That instinct has real data behind it. PwC's 2026 Global CEO Survey found that 56% of CEOs report no significant financial benefit from their AI investment to date, and only 12% say they've seen gains in both cost and revenue. A separate, widely cited MIT Project NANDA study — "The GenAI Divide: State of AI in Business" — found that 95% of the generative AI pilots it examined delivered no measurable P&L impact. Two different research teams, two different methodologies, landing in almost the same place: the gap between "the tool works" and "the P&L moved" isn't a rounding error. It's the norm.

But underdelivering and disappearing are not the same failure, and conflating them is how you end up cutting a tool that was never actually the problem. When I traced our own numbers — not the aggregate spend-versus-savings slide, but where the hours the copilot freed up that week actually went — they hadn't vanished. They'd gone to a person double-checking the copilot's draft against last week's actuals. They'd gone to a Slack thread where someone asked "wait, is this number from the new tool or the old process?" and three people had to weigh in before anyone trusted it enough to put it in a deck.

That reconciliation loop is exactly the overhead we built Rhythms' Reviews around — not a faster deck, but a review where every number traces back to the same connected, live source, so the "wait, where did this come from" question doesn't have an opening to happen in the first place. The coordination tax nobody puts on the AI business case existed before the AI tool arrived and will exist after it, unless something about how the review runs actually changes.

Why "Hours Saved" Is the Wrong Number to Chase

Most ROI conversations start with the wrong question. "How many hours did the tool save on the task" is measurable, satisfying to put in a slide, and almost useless — because a saved hour on a task that still has to be verified, reconciled against three other sources, and re-explained to a stakeholder who doesn't yet trust the output isn't a saved hour at all. It's a relocated one.

Here's what that looked like on our revenue team specifically. Their forecasting assistant cut individual prep time for the weekly pipeline call from ninety minutes to twenty. Nobody measured what happened next: the sales leader still spent the same forty minutes in the meeting asking "why does this number disagree with what you told me Tuesday," because the assistant pulled from a data source that updated an hour after the leader's last manual check. The saved seventy minutes showed up nowhere on any dashboard, because the thing that got slower — trust reconciliation — was never being measured in the first place.

This is the actual reason "hours saved" fails as a metric. It measures the task, not the decision the task was supposed to feed. A better question, and a harder one to answer cleanly, is whether a decision gets made faster or with less back-and-forth than it used to. Almost no company I've talked to is tracking that. They're tracking adoption logs and task-completion timestamps, which tell you the tool ran. They don't tell you whether the organization moved any faster because it ran.

The Meeting Where the Saved Hours Go to Die

Here's the pattern once you go looking for it: the leak isn't in the task the AI tool touches. It's in the review that happens after. Every operating rhythm — the weekly pipeline call, the monthly business review, the quarterly board prep — has an unspoken verification step built into it, because for years the person walking in with numbers had assembled them by hand and everyone knew, implicitly, how much to trust that process. AI changed the assembly. It didn't change the verification step, because nobody redesigned the review to account for a different, faster way of producing the input.

So the review still runs on the old rhythm: someone presents a number, someone else asks where it came from, a third person mentions they heard something different last week, and the meeting spends its first fifteen minutes re-litigating data provenance before it gets to the actual decision. That's the same dynamic as the pipeline call above, just wearing a different hat — a tool that turned ninety minutes into twenty upstream, and a room that still spends forty downstream deciding whether to believe what it produced.

I used to think of this as a trust problem with the tool. It isn't. It's a design problem with the review. We built Radar specifically to shrink this gap — surfacing a data conflict or a stalled metric quietly, to the person who owns it, on day three instead of in front of the whole leadership team on day thirty. A leader who walks into a review already knowing where the soft spots are doesn't spend the first fifteen minutes finding them live.

What Changes When You Redesign the Cadence, Not Just the Tool

The fix isn't a better model or a second tool layered on top of the first one that isn't paying off yet. It's redesigning the review itself so the coordination step shrinks along with the assembly step. That's a structural decision, not a technology purchase, and it's the part almost nobody's business case accounts for.

The lever that makes it durable is the cadence, not any single review. This is where Playbooks comes in for us — the recurring pull of the same structure, from the same sources, on the same schedule, so a team isn't rebuilding trust in the process from scratch every single week. Trust compounds when the mechanism is identical every time. It resets to zero every time the mechanism looks different, which is what happens when review prep stays one-off and ad hoc even after the underlying tooling gets faster.

None of these levers show up as "hours saved" in a spend-versus-return slide. They show up as a review that takes forty-five minutes instead of ninety, with a decision at the end of it instead of a data argument. That's the number that actually reaches the P&L, and it's not the number most AI business cases are built to measure.

What I'd Tell the Room If We Had That Meeting Again

I still don't think our forecasting copilot or our drafting assistant did anything wrong. I think we bought speed and asked it to show up as savings, without touching the part of the organization that was actually consuming the time — the part where people check each other's work because the process taught them, for years, that they had to. That habit doesn't retire on its own just because the input got faster.

If your AI tooling spend and your efficiency line are sitting on the same flat slide right now, I wouldn't start by asking whether the tool is good enough. I'd start by finding the specific meeting where the saved hours go to die — the one where someone asks "where did this number come from" and three people have to answer before anyone moves on. That's not a symptom of a bad tool. It's the actual location of the ROI everyone's been looking for.

Try Rhythms for free at rhythms.ai.

Frequently Asked Questions

Why hasn't our AI investment shown up in our efficiency numbers?

Most likely because the hours AI saves in one stage — drafting, data-pulling, summarizing — get spent again in the next stage: verifying the output, reconciling it against other sources, or re-explaining it to stakeholders who don't yet trust it. The saving is real. It just isn't reaching the P&L because nothing downstream was redesigned to keep it there.

What percentage of companies see no financial benefit from AI?

PwC's 2026 Global CEO Survey found that 56% of CEOs report no significant financial benefit from their AI investment so far, with only 12% reporting gains in both cost and revenue. Separately, MIT Project NANDA's "GenAI Divide" study found 95% of the generative AI pilots it studied delivered no measurable P&L impact. Different studies, same conclusion: the gap between adoption and measurable return is currently the norm, not the exception.

How do I explain a missing AI ROI to my board or CFO?

Reframe the conversation away from "did the tool work" and toward "did we redesign the process around the tool." Most AI deployments succeed at the task level and fail at the process level — the review and reporting cadence around the tool still assumes the old, slower way of producing information, and that's where the reclaimed hours quietly go.

What should we measure instead of hours saved when evaluating AI ROI?

Measure whether a decision gets made faster or with less back-and-forth, not whether a task finished faster. An hour saved on a task that still has to be re-verified, reconciled, and re-explained before anyone acts on it isn't ROI — it's relocated work. Track meeting length, number of reconciliation questions raised, and time from "number presented" to "decision made" instead.

Share this post:

FAQs

What is Rhythms?

Who built Rhythms?

How is Rhythms different from other OKR tools?

What tools does Rhythms integrate with?

How long does it take to set up Rhythms?

Stop managing the process.
Start building the business.

Stop managing the process.
Start building the business.

See how Rhythms replaces your operational overhead with AI that actually runs.