What "AI-Powered" Actually Means, and What It Often Doesn't

AI in Dealerships

Monty Wanless

Numa reads live DMS data to answer, route, and act on customer conversations, a specific, verifiable capability, across 1,300+ dealerships, and that specificity matters more than it might seem when almost every vendor pitch sounds the same. Nearly every vendor this year will claim to be "AI-powered." Research on how often that claim actually reflects meaningful capability suggests a GM should treat the label itself as close to meaningless, and focus instead on a handful of specific questions that separate genuine AI from a relabeled version of what already existed.

Why Almost Every Vendor Now Claims to Be "AI-Powered"

The label has become close to universal. McKinsey's 2025 Global Survey of nearly 2,000 executives across 105 countries found 88% of organizations now claim to use AI somewhere in their business, a number that's climbed sharply in a short span of time. That widespread adoption of the label isn't matched by widespread depth: the same survey found only about 6% of organizations qualify as what McKinsey calls "AI high performers," genuinely attributing meaningful business impact to it, and nearly two-thirds of companies claiming AI use haven't actually scaled it beyond a pilot or a narrow use case.

Key takeaway: Claiming to use AI and actually delivering meaningful AI capability have become two very different things. McKinsey's own research found the vast majority of organizations claiming AI use haven't scaled it into anything with measurable impact.

The Gap Between the Claim and the Underlying Product

This gap has a name in industry analysis. Gartner calls the pattern "agent washing": vendors relabeling existing chatbots, rule-based scripts, and basic automation as autonomous AI agents without the underlying technology to back up the label. Gartner estimates only a small fraction of the vendors currently claiming agentic AI capabilities are actually building systems with genuine autonomous reasoning behind them, and the firm projects the pattern will contribute to more than 40% of agentic AI projects being scrapped by the end of 2027, once the gap between the marketing and the actual product becomes impossible to ignore.

This isn't a reason to distrust AI as a category. It's a reason to stop trusting the word itself as evidence of anything, and to ask more specific questions instead.

Key takeaway: "Agent washing" isn't a fringe problem. Gartner's own estimate is that only a small fraction of vendors currently claiming agentic AI capability are actually building systems with genuine autonomous reasoning behind the label.

Even Genuine AI Fails Often, for a Reason Worth Asking About Directly

Passing the "is this actually AI" test isn't the end of the diligence either. RAND Corporation's 2024 research, based on interviews with 65 data scientists and engineers across government and industry, found more than 80% of AI projects fail to reach meaningful production, roughly double the failure rate of comparable non-AI technology projects. The research is clear that this usually isn't a story about the underlying models being inadequate. It's a story about the conditions the AI gets deployed into, and one specific, recurring cause stands out: fragmented data spread across disconnected systems that the AI can't actually read cleanly.

That finding lands directly on a question worth asking before signing anything. A system that reads live, connected DMS data is working with exactly the kind of data foundation RAND's research says separates the AI projects that actually work from the 80% that don't. A vendor with genuinely capable AI, sitting on top of the same fragmented, siloed data every other tool in the building also can't fully see, is still likely to land in that failure statistic, regardless of how real the underlying technology is.

Numa perspective: Genuine AI capability isn't sufficient on its own. RAND's own research found data fragmentation is one of the most common reasons even real AI projects fail to deliver value, which makes "what data does this actually connect to" as important a question as "is this actually AI" in the first place.

What Actually Distinguishes Real AI Capability From a Relabeled System

A few concrete questions separate a genuinely capable system from a script wearing new marketing. Does it read live, current data, an actual account status, an actual schedule, an actual customer record, or does it work from a fixed set of pre-written responses regardless of what's actually true right now? A system limited to the second is automation with a modern label, not a system capable of the kind of judgment "AI-powered" implies. What happens when a conversation goes somewhere the vendor didn't specifically script for? A genuinely capable system handles the unexpected case reasonably. A relabeled script either breaks, loops, or defaults to a human handoff immediately, which is a legitimate design choice for some situations, but not the same thing as the autonomous capability being advertised.

A broader framework for evaluating any AI vendor's claims goes deeper into the specific due-diligence questions worth bringing into any vendor conversation, but the two above are the fastest filter when time with any single vendor is limited.

Numa perspective: The word "AI" in a vendor's pitch tells you what they want you to assume. It tells you nothing about whether the system reads real data, handles the unexpected, or does anything a well-built script from five years ago couldn't already do.

Why a Live Demo Is a Better Test Than a Spec Sheet, If You Ask the Right Way

A scripted demo is built to succeed, which means it's the weakest possible test of a system's actual range. The more useful test is asking a vendor to run a scenario they didn't prepare for, ideally something specific and slightly unusual pulled from an actual situation your store deals with regularly, rather than a generic example. A genuinely capable system will engage with it, even imperfectly. A relabeled script tends to reveal its edges quickly, either by producing a generic non-answer or by routing to a human immediately regardless of how simple the actual question was.

The Bottom Line: Ask What It Reads and What It Does When It's Wrong

The word "AI" has stopped functioning as a useful signal on its own, and the research on how widely the label gets applied relative to how rarely it reflects genuine, scaled capability backs that up directly. Numa's own approach is built around specifics that hold up to exactly this kind of scrutiny: live DMS data, real escalation logic, and a system that keeps improving from actual interactions rather than staying fixed to whatever it shipped with. Walking any vendor conversation with two questions in hand, what data does this actually read, and what happens when it hits something it wasn't specifically prepared for, filters out more noise in five minutes than an hour of comparing marketing claims ever will.

Frequently Asked Questions

Why do so many vendors now claim their product is "AI-powered"?

The label has become close to universal in marketing, but research shows adoption of the term has far outpaced actual scaled capability. McKinsey's 2025 survey found 88% of organizations claim to use AI, while only about 6% qualify as genuinely deriving measurable business impact from it, meaning the claim itself has become a weak signal of anything specific.

What is "agent washing"?

Agent washing is a term Gartner uses to describe vendors rebranding existing chatbots, scripts, or rule-based automation as autonomous AI agents without the underlying technology to support that label. Gartner estimates the pattern will contribute to more than 40% of agentic AI projects being canceled by the end of 2027 once the gap between the marketing and the actual capability becomes clear.

Does confirming a vendor's AI is genuine mean the project will actually succeed?

Not necessarily. RAND Corporation's research, based on interviews with 65 data scientists and engineers, found more than 80% of AI projects fail to reach meaningful production even when the underlying technology is real, with fragmented data across disconnected systems as one of the most common causes. Confirming genuine AI capability is a necessary check, not a sufficient one; the data foundation the AI is connected to matters just as much.

What questions actually reveal whether a system has genuine AI capability?

Two are especially useful in a short conversation: does the system read live, current data or work from a fixed set of pre-written responses, and what happens when a conversation goes somewhere the vendor didn't specifically script for. A system that only handles anticipated scenarios and defaults immediately to a human or a generic response outside of them is closer to automation with a new label than genuine AI capability.

Is a live demo a reliable way to evaluate an AI vendor?

A standard demo is built to succeed and therefore isn't a strong test on its own. A more revealing approach is asking the vendor to handle a specific, slightly unusual scenario pulled from an actual situation your business deals with, rather than a generic example, since that's where the difference between genuine capability and a relabeled script tends to show up quickly.

Does this mean AI claims from vendors shouldn't be trusted at all?

Not entirely, but the word alone shouldn't be treated as evidence of anything specific. The more reliable approach is asking concrete questions about what data the system actually reads and how it behaves outside of prepared scenarios, rather than accepting "AI-powered" as a meaningful claim on its own.

See what Numa's AI actually reads, and how it handles what it wasn't specifically prepared for. Talk to Numa.