
Last update:
What Is Agent-Washing? The One Question That Exposes It in Any AI Vendor Demo

We sat in on a demo last quarter where the "agent" could reschedule a meeting on command and reprioritize a task list without being asked twice. Then someone in the room gave it a real calendar conflict — not the clean one in the script, the messy one where two attendees had overlapping holds and a third had blocked the whole afternoon. The agent froze. The sales engineer smiled and said, "we'd loop in a human for that," like it was a feature. It wasn't. It was the demo undone in one sentence.
That's agent-washing, and it's showing up in more evaluation cycles than most buying committees realize.
Agent-washing is when a vendor markets a scripted, rules-based workflow as "agentic AI" without the system actually making autonomous decisions or handling situations it wasn't explicitly programmed for. The fastest way to expose it in a demo is to ask what happens when the system hits something it wasn't briefed on. A genuine agent adapts or clearly flags its own uncertainty. A scripted system freezes, errors out, or quietly routes to a human while the demo calls that a "feature."
The Demo Question Nobody Asks (and the One You Should)
Most evaluation calls spend forty-five minutes watching a vendor run their best-case path — every step rehearsed, every input matched to an output. It's a good demo the way a magic trick is good — smooth by design.
The question that breaks it on purpose is simple: what happens when the agent is wrong, and who finds out first, and how? Not "does the system have human oversight" — every vendor says yes to that, and it means nothing on its own. Ask instead for the last time this thing made a bad call in production, and exactly how someone caught it. A vendor with a genuinely agentic product answers in one breath, with a name attached to the catch. A vendor agent-washing a scripted tool reaches for reassurance instead of a mechanism, because there isn't one to describe.
What Agent-Washing Actually Is (and Why It's Everywhere Right Now)
"Agentic" has become a checkbox the way "cloud-native" was a decade ago, when plenty of vendors relabeled the same on-prem software as "cloud-enabled" without changing the architecture underneath. The label moved faster than the substance did.
We're in that same window now, except faking "autonomous decision-making" for 45 minutes is easier than faking "runs in the cloud" ever was — a scripted workflow with a chat interface on top looks agentic if you never push past the happy path. It's only at the edge of the rules that the difference shows up — and according to Kai Waehner's 2026 "Enterprise Agentic AI Landscape" research, 84% of enterprise leaders say they've personally encountered agent-washed products during vendor evaluations. That's most buying committees, right now.
The Tell: What Happens at the Edge of the Rules
A scripted workflow executes predefined steps in response to predefined triggers. It can be genuinely useful and look sophisticated, but it can't handle a situation its rules didn't anticipate — only fail at it gracefully or ungracefully. A real agent makes a judgment call within a scoped decision space, and it can explain or flag that judgment when asked.
The tell isn't in the pitch. It's in the edge case. This is the same reason we built Rhythms' Radar around surfacing what actually happened rather than what a system claims it would do — a slipping initiative gets flagged on day three because the mechanism is watching real signals, not because a script promised it would. If a vendor can't point to an equivalent mechanism for their own agent, the "autonomous" claim is doing more work than the product is.
Five Evaluation Questions to Bring to Your Next AI Demo
You don't need a technical audit to catch agent-washing — five questions, asked out loud before the pilot agreement gets signed, will do it.
1. What's the last decision this agent made that surprised even your own team? A polished success story isn't an answer — push again. Every genuinely agentic system has produced at least one outcome its builders didn't fully anticipate; that's what autonomy means. A vendor with nothing to say here is describing a workflow, not an agent.
2. Show me the exact moment a human gets looped in — not the policy, the mechanism. "Human oversight is maintained" is a sentence, not a system. Ask for the specific trigger: a dollar threshold, a confidence score, a category of action. No named trigger usually means the human gets looped in whenever the model gets stuck — a failure mode wearing a governance label.
3. What happens when two of the agent's own rules conflict? Scripted systems usually resolve this by picking whichever rule fired first, arbitrarily and undisclosed. A real agent should describe, even imperfectly, how it weighs competing priorities. This is one of the sharpest tells available, and almost nobody asks it.
4. Can it tell me it doesn't know, or does it just proceed? A system that always produces a confident answer, never an "I'm not certain, here's why," is either extremely well-scoped or quietly hiding its failure cases. Ask it to just say "I don't know" out loud, in the room. If it can't do that, that's the answer.
5. How does this get evaluated after the pilot, and by whom? A recurring cadence matters more than a one-time bake-off. "We'll check back in at renewal" is a pilot with no feedback loop. We run this internally as a standing check inside Rhythms' Playbooks — a recurring line in a cadence that already exists, not a special project someone has to remember — because a governance question asked once a year gets forgotten.
What a Genuinely Agentic Vendor Sounds Like When You Ask
The honest answer to "what happens when the agent is wrong" is never smooth. It involves a specific example, a specific person, and usually a redacted dollar figure — a little uncomfortable to say in a sales call, because it's an admission the system doesn't work perfectly. Paradoxically, that's the strongest signal it works at all.
We'd rather a vendor tell us about the deal record their agent misread and the reviewer who caught it than watch another flawless script run its happy path. This is exactly what we try to build into how Rhythms' Reviews surfaces agent activity: not a highlight reel, but the actual decision trail, including the ones that needed a second look.
Trust in the category is already strained. The same broad Waehner survey found 87.5% say agent-washed products have damaged their trust in AI vendor claims generally, not just the specific product evaluated. That's the real cost: it doesn't just burn one vendor relationship, it makes the next genuinely agentic pitch harder to believe for everyone.
The fix isn't cynicism about agentic AI. It's one sharper question, asked out loud, before the contract gets signed: what happens when this is wrong, and who finds out first? A vendor with a real answer will tell you. A vendor without one will tell you not to worry about it.
If you're evaluating vendors and want to see what an honest answer to that question actually looks like, request a demo at rhythms.ai.
Frequently Asked Questions
What is agent-washing in AI marketing?
Agent-washing is when a vendor labels a scripted, rules-based automation as "agentic AI" to ride the current wave of enterprise interest in autonomous systems, even though the underlying product can't adapt to situations outside its predefined rules. It's the same move as "cloud-washing" a decade ago, applied to a new buzzword.
How do I know if an AI tool is actually agentic?
Ask to see it handle something it wasn't specifically set up for during the demo — an edge case, a conflicting input, a scenario the sales engineer didn't rehearse. A genuinely agentic system reasons through it or clearly flags its own uncertainty; a scripted system fails visibly or reveals a human was quietly going to handle that case all along.
What questions should I ask in an AI vendor demo to test for real agentic behavior?
The single most useful one: what happens when the agent is wrong, and who finds out first, and how? A vendor with a genuine agentic product has a specific, concrete answer involving monitoring and escalation. A vendor agent-washing a scripted tool usually reframes the question as a reassurance about "human oversight" without describing an actual mechanism.
Why has trust in AI vendors declined recently?
Because buyers have increasingly encountered a gap between what's marketed as "agentic" and what the product actually does once deployed. Kai Waehner's 2026 Enterprise Agentic AI Landscape research found 84% of enterprise leaders had personally encountered agent-washed products during vendor evaluations, and 87.5% said it had damaged their trust in AI vendor claims broadly, not just the specific product involved.
What's the difference between a scripted workflow and a real AI agent?
A scripted workflow executes predefined steps in response to predefined triggers — it can look sophisticated but can't handle a situation its rules didn't anticipate. A real agent makes a judgment call within a scoped decision space and can explain or flag that judgment. The demo tell is what happens at the edge of the rules, not in the middle of them.
Share this post: