Skip to content.
Back to Walnut
The Sales Insider
Brought to you by

Key Takeaways

  • Nearly every sales tool now claims AI, which makes the claim itself useless as a filter and puts the burden of verification on the buyer.
  • The most reliable test is mechanism, not features: in a genuinely AI-native product, AI is how the value gets produced, so turning it off leaves you with very little.
  • Ask what the AI is grounded in. A system drawing on your CRM and your approved content behaves differently from one drawing on a general model with your logo on it.
  • Ask who is accountable when the output is wrong. If the answer is the model, there is no answer.
  • A system that says it does not know is more trustworthy than one that always produces something, and that behavior is worth testing directly during an evaluation.
  • Gartner expects 80% of sales leaders to treat AI integration as a critical competitive factor by 2030, which means these evaluations are about to matter considerably more than they do today.

Every sales tool on your shortlist claims AI. That is not a cynical observation, it is a practical problem: when a claim is universal, it carries no information, and the work of separating real capability from positioning moves onto you.

The stakes are rising fast. Gartner projects that 70% of routine sales tasks will be automated by 2030, and that 80% of sales leaders will consider AI integration a critical competitive factor by that point (Gartner, The Future of Sales 2030). Meanwhile adoption is shallower than the marketing suggests. Forrester found in 2025 that only one in five marketing organizations have actually embedded generative AI into their workflows, a figure cited in Walnut’s State of Generative AI in B2B Marketing 2025. The gap between how many vendors claim AI and how many buyers have made it work is where bad purchases happen.

What follows is a framework for evaluating AI sales tools that does not rely on the vendor’s own vocabulary. Five questions, each designed to be answered in a live demo rather than in a security questionnaire, plus the red flags that should slow a deal down and what a good answer actually sounds like.

Why the AI-native claim became worth checking

There is a useful distinction underneath all of this, and it is about mechanism rather than feature count.

In a product where AI is a lane bolted onto an existing core, the software still works the old way and AI helps you do the work faster. In a product where AI is the mechanism, the AI produces the outcome and there is no old way underneath it. Both can be legitimate purchases. They are different products with different risk profiles, and vendors in the first category have a strong incentive to describe themselves as if they were in the second.

The reason this matters commercially is roadmap. A product built with AI as the mechanism improves as models improve, because the mechanism gets better. A product with an AI lane attached improves when the vendor builds more lanes. Over a three-year contract that difference compounds, and it is largely invisible at the point of purchase unless you go looking for it.

None of this means AI-native is automatically the right choice. It means the claim should be verified rather than accepted, the same way you would verify an uptime number.

The five questions to evaluate AI sales tools

Ask these in a working session with the product open, not in a written RFP response. Written answers are drafted by marketing. Live answers are not.

What happens if you turn the AI off?

This is the fastest test there is. Ask the vendor to describe what remains if every AI capability is disabled.

If the answer is a fully functional product that you would still buy, AI is an enhancement layer. That may be exactly what you want, but price and evaluate it accordingly. If the answer is that there is not much of a product left, AI is the mechanism. The tell is how quickly and comfortably they answer, because vendors in the first category tend to hear this question as an attack and reach for a longer answer than it needs.

Can they show you the mechanism, not just the output?

Any vendor can show you a polished result. Ask instead to watch the thing being produced, live, with an input you supply.

Bring your own scenario. Ask them to build against it in the session rather than showing a prepared example. The difference between a system that generates and a system that retrieves a pre-built artifact becomes obvious within about ninety seconds, and no amount of preparation hides it when the input is genuinely new.

This single move eliminates more false claims than any other question on the list, because a prepared demo can survive almost any interrogation except an unfamiliar input.

What is the AI grounded in?

Grounding is what separates output that is about your business from output that is merely plausible. Ask specifically: which of our systems does it read, what content does it draw on, and what happens to that content when we update it.

A tool grounded in your CRM knows the account, the stage, and the stakeholders. A tool grounded in your approved content produces claims your product marketing team has already signed off on. A tool grounded in neither is producing confident text from a general model, which is fine for a first draft and dangerous in front of a buyer. The connection between demo systems and CRM data is worth understanding in detail, and we covered the mechanics in How Do Interactive Demos Integrate With CRM Systems?.

Follow-up worth asking: when the source changes, does the output change automatically, or does someone have to rebuild it? That answer tells you whether you are buying a living system or a generator that produces artifacts which immediately start aging.

Who is accountable when the output is wrong?

Every AI system produces wrong output sometimes. The question is what the product does about it structurally.

Look for an approval step that a named person passes through, a record of what was generated and who signed off, and controls over what the system is permitted to say. If a vendor treats the absence of human review as a feature, understand what they are actually offering: they have removed the only point in the process where your organization was answerable for what a customer received.

This is not a theoretical concern once these systems touch buyers directly. When a prospect challenges something an AI system told them, your team needs to be able to reconstruct what happened. If the product cannot show you that, you will be reconstructing it from memory during a deal review.

What does it do when it does not know?

This is the question that reveals the most about engineering seriousness, and almost nobody asks it.

A system that always produces an answer is not more capable than one that sometimes declines. It is less honest about its limits, and in a buyer-facing context that is a liability rather than a feature. Ask the vendor to show you the behavior at the edge: pose a question the system has no grounds to answer and watch what happens.

The good outcome is that it says so plainly and routes to a human with the context attached. The bad outcome is a confident, fluent, wrong answer delivered to your prospect at the exact moment they were deciding whether to trust you.

Red flags worth slowing a deal down for

Four patterns recur, and each is checkable in a single session.

Vague capability language is the most common. “AI-driven insights” describes nothing. A vendor who knows what their system does can tell you what it takes as input, what it produces, and where it fails. Ask them to complete the sentence “our AI takes X and produces Y” and see how long it takes.

Demos that only run on the vendor’s data are the second. If your scenario cannot be used, the system may not generalize as well as the prepared example suggests.

Refusal to discuss failure modes is the third. Every real system has them. A vendor who claims none is either not being straight with you or does not know their own product well enough.

The fourth is the subtlest: an AI story that does not match the product’s structure. If the AI capabilities live in a separate menu, are priced as a separate add-on, and are described in separate marketing, the product is organized around AI as a lane. That is worth knowing regardless of what the homepage says.

A useful cross-check on all four is to ask the same questions of two different people at the vendor, ideally an account executive and a solutions engineer, in separate conversations. Where a product genuinely works the way the marketing describes, those two answers match closely. Where the story was assembled for the market rather than from the product, they diverge quickly, and the solutions engineer is almost always the one telling you how it actually works.

How to run this in a live evaluation

Compress it into one ninety-minute working session with the vendor and two people from your side, one commercial and one technical.

Bring a real scenario from an open opportunity, sanitized if needed. Ask them to produce against it live. While that runs, work through the five questions in order, and take notes on how they answer rather than only what they answer, because hesitation on the mechanism question is itself data.

Then do one thing most evaluations skip: ask to see the administrative view. Who has access, what has been generated, what controls exist over what the system may say. Products that took governance seriously have a real answer here and are usually pleased to be asked. Products that did not will show you a settings page with three toggles.

At Walnut, that is the shape the platform is built around: AI Mode generates and adapts demos from a described change rather than manual screen-by-screen editing, StoryCaptureAI assembles a demo from a real workflow and captures the live product as interactive HTML so it keeps reflecting the product as it ships, and InsightsAI answers questions about engagement in plain language. The human sets the story and approves the output. That division is deliberate, and it is the thing to probe with any vendor, including us.

For a related evaluation angle specific to demo platforms, Why AI Mode Is Now a Must-Have When Evaluating Interactive Demo Platforms covers what to require, and Interactive Demo Platform vs Custom Code covers the build-versus-buy decision underneath it.

What a good answer sounds like

Good answers share a texture. They are specific about inputs and outputs, comfortable naming limits, and unbothered by the harder questions.

A vendor being straight with you will tell you what their system is bad at without being asked twice. They will let you drive. They will describe the human’s role in the process rather than pretending the human has been eliminated, because a product designed for real revenue teams assumes a person is accountable at the end.

Notably, 78% of heavy AI users report confidence that their output is unique, per Walnut’s generative AI research, which is a reminder that confidence and verification are not the same thing. That applies to the vendors you are evaluating and to your own team’s assessment of the tools you already bought.

The single best question remains the first one. Turn the AI off and see what is left. Everything else is a refinement of that.

Frequently asked questions

What does AI-native actually mean?

AI-native describes a product where AI produces the value rather than assisting a user who produces it. The practical test is subtraction: disable the AI and see what remains. In an AI-native product there is little left, because the AI is the mechanism. In a product with AI added, the original software still works and the AI made parts of it faster.

What is AI washing?

AI washing is presenting a product as AI-driven when the underlying capability is conventional automation, a thin layer over a general model, or simply marketing language applied to existing features. It is detected by asking what the system takes as input and produces as output, and by watching it run against a scenario the vendor did not prepare for.

How do I test an AI sales tool during a demo?

Bring your own scenario from a real opportunity and ask the vendor to work against it live rather than showing a prepared example. Ask what happens when the AI is disabled, what data the AI is grounded in, who approves output before it reaches a customer, and what the system does when it lacks the information to answer. Ask to see the administrative and governance view, not just the end-user experience.

Is a tool with AI features worse than an AI-native tool?

Not necessarily, and treating it as a hierarchy leads to bad decisions. They are different products. A tool with AI features is often the right buy when your team already has a working process and wants specific steps accelerated. An AI-native tool is the better buy when you want the outcome produced rather than assisted. The mistake is paying AI-native prices for AI-assisted capability because the marketing did not distinguish them.

What questions should I ask a vendor about AI safety and control?

Ask what the system is permitted to say and who configures those limits, whether output passes a human approval step before reaching a customer, what record exists of what was generated and who approved it, how access is controlled by role, and what the system does when it does not know an answer. Also ask what data the AI reads and where it is processed, which your security team will need regardless.

Why does it matter what the AI is grounded in?

Grounding determines whether output is about your business or merely plausible. A system reading your CRM and your approved content produces claims that match your positioning and your accounts. A system running on a general model produces fluent text that may contradict your messaging or invent capability. For anything a buyer sees directly, grounding is the difference between a useful tool and a risk.

How much should AI capability influence a buying decision in 2026?

It should influence it heavily, but as a question of mechanism and control rather than of feature count. Gartner expects 80% of sales leaders to treat AI integration as a critical competitive factor by 2030, so the direction is clear. The practical advice is to weight how the AI produces value, what it is grounded in, and what oversight exists, well above how many capabilities carry an AI label.

Ready to see what personalized demos can do for your pipeline? Start for free with Walnut.

You may also like...

AI

AI Demo Creation Is the Easy Part. Knowing What to Show Isn’t.

Key Takeaways Twenty-nine percent of B2B marketing teams now produce more than half their content with generative AI, and among…
13 min read
Keep reading
AI

Interactive Demo Platform vs Custom Code: Why B2B Teams Choose a Platform

KEY TAKEAWAYS Engineering can build a demo. The question is whether they should. The pitch for custom code is straightforward:…
11 min read
Keep reading
Your Buyer Met an AI Demo Agent Before Your Rep
AI

Your Buyer Met an AI Demo Agent Before Your Rep

Key Takeaways Your buyer just had a demo. Your rep wasn’t there. Neither were you. They opened ChatGPT, or Claude,…
10 min read
Keep reading

You sell the best product.
You deserve the best demos.

Halftone purple background Halftone green background
Never miss a sales hack
Subscribe to our blog to get notified about our latest sales articles.

Book a Demo

Are you nuts?!

Walnut squirrel mascot illustration

Appreciate the intention, friend! We're all good. We make a business out of our tech. We don't do this for the money - only for glory. But if you want to keep in touch, we'll be glad to!


Let's keep in touch, you generous philanthropist!

Sign up here!

Fill out the short form below to join the waiting list.

Let's get started

Enter your email to get started

Nice to meet you

Share a bit about yourself

Company Info

Introduce your company by filling in your company details below

Let's get you started in Walnut…

Set your password and start building interactive demos in no time.

Continuing in 4 seconds...