AI Keeps Recommending Your Brand. How Strong Is the Case Behind It?
We tested three evidence conditions across four answer engines. The same brand won all 36 runs. The reasons behind that choice were far less stable.
Michelle Perkins, Founder of ValueTempo

TL;DR
In 36 controlled VAT-compliance comparisons, the same brand won every time. But the evidence supporting that choice was not equally relevant to every buyer question. Specific responses inferred ongoing finance-team workload from setup time, partner involvement or product structure without direct workload evidence. PROVE gives AEO and GTM teams a way to audit that claim-to-evidence gap.
The top-line result from phase 3 of our AEO research looked almost boring: 36 runs. Four AI answer engines. One winner.
Then we examined the explanations and found the real story: the evidence trail changed even when the selection did not.
One answer saw managed services and account support, then inferred less work for a lean finance team. Another saw partners in the delivery model, then inferred more coordination overhead. Both conclusions sounded plausible. Neither was supported by comparable evidence of the work the buyer would actually perform.
That gap is easy to miss when an AI recommendation is favorable, well written and full of citations. It is also where the PROVE stage begins.
ValueTempo's AI-mediated buying framework separates four decisions. APPEAR asks whether your brand enters consideration. PREFER asks which credible option becomes the first choice. PROVE asks whether the evidence behind that preference addresses the buyer's actual decision criteria or merely sits close enough to sound relevant. PERFORM asks whether that preference turns into buyer action and, ultimately, revenue.

PROVE examines whether the evidence behind an AI preference supports the buyer's actual decision criteria.
We continued to use the VAT-compliance software category because it forces a buyer—and an answer engine—to evaluate more than a feature list. A growing international SaaS company must understand who remains merchant of record, which tax obligations are covered, how the product connects to its billing stack and how much responsibility stays with its finance team. A brand can look capable at the category level while remaining a poor fit for that operating model.
Our earlier comparison fixed the buyer and candidate set to examine brand preference. It produced a consistent winner, but it could not show whether the evidence supporting that preference was equally relevant to every buyer requirement.
This follow-on experiment moved from which brand was preferred to how the case for that preference was constructed. It focused on evidence-role architecture: which sources were available, what each source could credibly establish and where the answer engine crossed from evidence into inference.
We investigated three questions:
- Does adding external corroboration strengthen the case beyond vendor claims?
- Does adding limitations and delivery boundaries change how engines qualify the recommendation?
- Can the evidence change the explanation before it changes the winner?
The third question produced the clearest answer. Across the three evidence conditions, parts of what the engines cited, verified and qualified changed while the selected brand remained the same. Recommendation share alone would have concealed that movement in the decision story being told about the product.
How we tested evidence architecture
We fixed one buyer scenario and three brand profiles: a SaaS-oriented tax platform, a broader enterprise tax platform and a managed-compliance brand.
The buyer was a growing international SaaS company that wanted to remain merchant of record, keep control of billing, support B2B and B2C sales, integrate with Stripe or Chargebee, cover the compliance lifecycle and operate with a lean finance team.
We then changed the evidence available to the answer engines:
| Condition | Evidence supplied | What it tested |
|---|---|---|
| A: Owned facts | Product capabilities, integrations, service scope and implementation information | What vendor-controlled evidence could establish |
| B: Added corroboration | A plus ecosystem, institutional, marketplace and customer evidence | What outside confirmation added |
| C: Added boundaries | B plus limitations, delivery dependencies and implementation or customer-risk evidence | Whether fuller context changed the judgment |
ChatGPT, Claude, Gemini and Perplexity each evaluated all three conditions three times. We used fresh sessions, rotated brand order and recorded browsing as disabled.
4 engines × 3 conditions × 3 repetitions = 36 responses.
This was a supplied-evidence test, not a live-search test. Because each condition added information as well as source types, it does not isolate provenance as a single causal variable.

The selected provider remained stable, but sources, some requirement ratings and qualifications changed across the evidence conditions.
The winner did not move. The evidence trail did.
The SaaS-oriented platform was selected in all 36 responses.
That tells us the preference was stable in this scenario. It does not tell us that the evidence treatments had no effect.
Condition B added different forms of corroboration for each brand because that is what the public evidence supported. The SaaS and enterprise platforms gained Stripe and Chargebee documentation. The enterprise brand also had certification for a defined U.S. tax program. The managed-compliance brand gained an institutional marketplace listing and customer-service reviews, but no evidence of direct Stripe or Chargebee integration.
This was not an equal dose of “more proof.” It preserved real differences in what could be corroborated.
Those differences exposed two patterns.
Corroboration can strengthen an evidence trail without changing a score
For both the SaaS-oriented and enterprise platforms, billing integration was already rated Verified in every Condition A response. It remained Verified in Condition B.
The added Stripe and Chargebee documentation still appeared in the explanations. The engines used it to support claims that had previously rested on vendor documentation alone. Our three-level rating simply had no room to record that stronger provenance.
For AEO teams, this is an important measurement problem. A flat requirement score can hide a meaningful change from self-asserted capability to ecosystem-confirmed capability.
A credible source can still be indirect evidence for the buyer's question
In two ChatGPT repetition pairs and one Gemini pair, the SaaS platform's lean-team rating moved from Partial under owned facts to Verified after customer and implementation evidence was added.
Perplexity did not make that jump. It kept the rating Partial across all three A/B pairs and noted that the packet did not directly measure finance-team workload.
The important distinction was not source credibility alone. It was what those facts were sufficient to prove about this buyer.
Implementation timing can help answer, “How long might setup take?” It does not directly answer, “How much work will our finance team own after launch?” Reviews from software companies can make the customer sample more comparable. They still may not describe responsibility for exceptions, partner handoffs or recurring compliance work.
This is the evidence-role problem: a source can be credible and relevant to the category while still being insufficient for the buyer conclusion an engine draws from it.
More boundary evidence did not guarantee a more disciplined answer
Condition C supplied more of the awkward facts: supported-scope limits, delivery dependencies, implementation risks and customer-reported friction.
The confidence scores do not support a simple conclusion that more boundary evidence makes engines uniformly more cautious.
| Engine | Owned facts | Owned + corroborated evidence | Owned + corroborated + boundary evidence |
|---|---|---|---|
| ChatGPT | 84 | 87 | 84 |
| Claude | 62 | 70 | 70 |
| Gemini | 90 | 90 | 85 |
| Perplexity | 78 | 78 | 82 |
Median self-reported confidence across three repetitions per condition.
Confidence fell for some engines, held for one and rose for another. These numbers describe the engines' stated confidence; they do not measure correctness or buyer trust.
The more useful signal appeared in the written reasoning. In one Condition C response, the packet explicitly said the available evidence did not establish how the managed model affected buyer effort or human handoffs. The answer still penalized that brand for handoff complexity.
In that response, supplying the limitation did not ensure that the answer respected it.
That is not proof that boundary evidence fails. It shows that supplying more evidence is only one component. Teams also need to inspect the relationship between the supplied fact and the conclusion produced from it.
The real failure mode: a plausible bridge with a missing span
In the audited examples, the problem began with a documented fact and then crossed into an unsupported buyer outcome.
| Documented information | Buyer conclusion in an audited response | Missing evidence |
|---|---|---|
| Managed functions and an implementation estimate | The platform will reduce ongoing manual burden | Work remaining after implementation |
| Partners support calculation or payment functions | The managed-compliance brand creates the most coordination overhead | Who owns coordination, exceptions and handoffs |
| Several products or services may be required | The buyer will need dedicated resources | Staffing requirements for this configuration |
These distinctions matter commercially.
Setup time is not operating workload. Partner involvement is not buyer-owned coordination. Product breadth is not a staffing model.
An answer engine can connect those ideas in a sentence that reads well and still skip the evidence needed to make the connection useful for this buyer.
Some responses preserved the gap. They described a potential benefit, then said staffing or ongoing effort was not quantified. That is the behavior PROVE should reward: distinguish what is documented, what is a reasonable hypothesis and what remains unknown.
We made a similar mistake in our first analysis
Our original read of the experiment moved too quickly from changed confidence and explanations to “stronger evidence improved the decision.” The data did not establish that.
We also translated some grouped lifecycle judgments into finer requirement scores without a consistent rule. We rebuilt the explicit-rating record, preserved ambiguous items as unscored and withdrew aggregate error counts that had not passed a complete claim-level review.
The corrected conclusion is narrower and more useful:
In this test, added evidence changed parts of the support and qualification around a stable recommendation. It did not establish a systematic improvement in reasoning, calibration or commercial outcome.
The same standard applies to AEO reporting. If a dashboard says visibility, citations or confidence improved, ask what decision claim became better supported—not just what metric moved.
A PROVE workflow for AEO and Marketing Engineering
The operating question is no longer simply, “How do we earn more mentions?” It is, “How do we make the buyer-relevant claims in those answers more defensible?”
A practical review loop has five steps:
- Capture the exact claim. Preserve the answer, cited passage, prompt, engine, date and settings.
- Map it to a buyer requirement. Name the decision question the claim is supposed to answer.
- Grade the evidence relationship. Is the claim directly supported, partially supported or inferred beyond the source?
- Resolve the product truth. Ask product, implementation, customer success or commercial teams what actually happens in the workflow.
- Retest the claim. Hold the prompt and environment as constant as practical, change the evidence, and inspect whether the answer becomes more accurate or appropriately qualified.

AEO systems can automate collection and comparison. Buyer relevance, source sufficiency and product truth still require reviewed judgment.
The division of labor matters. Current AEO platforms can automate much of the observation layer; the buyer-decision layer still needs a reviewed standard.
| Work | AEO tools (examples only) | What the system can handle | What remains manual or custom |
|---|---|---|---|
| Collection | Profound Answer Engine Insights, Peec AI, OtterlyAI | Run tracked prompts, retain responses, extract mentions and citations, and build history across engines | Define prompts around real buyer decisions; preserve controlled settings and evidence packets when causal learning matters |
| Comparison | Profound, Peec AI, Scrunch Monitoring | Compare visibility, position, sentiment, citations, competitors and changes by prompt, topic, model or time | Create buyer-requirement codes; decide whether two claims are equivalent; separate run variation from intervention effects |
| Evidence review | Profound FactCheck, Peec Sources and Chats, Scrunch Signals | Surface claim clusters, cited URLs, source gaps, content issues and potentially inaccurate brand statements | Judge whether a passage directly proves the buyer outcome; verify responsibilities, limitations and product truth with subject-matter owners |
| GTM action | Profound Agents, Peec Actions, Scrunch Signals and API | Generate prioritized content actions, route reports, create briefs, track changes and send data into downstream workflows | Decide whether the gap constrains real evaluations; assign commercial priority; connect changes to CRM, call, pipeline or experiment data |
Feature availability varies by plan, and some workflow or API capabilities are enterprise-only or in beta. More importantly, these products are strongest at monitoring and operationalizing AI-search signals. Teams still need a governed claim record—inside an existing platform, a spreadsheet or a custom system—to evaluate buyer requirement × answer claim × supporting passage × unresolved evidence together.
For the workload example, the content response may be a responsibility table rather than another feature page. It could show the task, who performs it, when a partner enters, what remains with the customer and how exceptions are handled.
But only publish what the operating model can substantiate. If ongoing workload has not been measured, preserve the uncertainty. Do not convert a convenient inference into positioning.
Success is not a higher confidence score or one more citation. At the PROVE stage, success means the answer makes a more supportable claim about a buyer-relevant question. Whether that better answer influences evaluation or revenue belongs to the next stage.
What PROVE changes about evidence architecture
Evidence architecture is often treated as a source-diversity exercise: add third-party mentions, reviews and authoritative citations.
Our test points to a stricter definition:
Evidence-role architecture is the fit between a buyer requirement, the claim made about it, the source used to support that claim and the limitations that bound it.
That definition changes the content plan. The gap may not be a missing article. It may be a missing proof object: a responsibility map, an implementation boundary, a workflow comparison, a customer example with the right operating context or an explicit statement of what is not yet known.
It also changes the measurement plan. Track recommendation share, but do not stop there. Track which buyer requirements are addressed, which sources carry each claim and where the answer crosses from documented fact into plausible inference.
The central PROVE question is:
Which buyer requirements does the evidence behind this recommendation actually address—and where is the answer filling in the gaps?
Part 4, PERFORM, moves from the quality of the recommendation to its commercial effect: what buyers do next, which outcomes can be observed and what would justify attributing those outcomes to an AEO intervention.
Frequently asked questions
What is PROVE in AEO?
PROVE is the stage where a team evaluates whether the evidence behind an AI recommendation directly supports the buyer-relevant claims being made. It follows visibility and preference analysis by examining source fit, limitations and unsupported inference.
Did more evidence change the recommended brand?
No. The same brand was selected in all 36 responses. Added evidence did change some cited support, requirement ratings and qualifications, but this experiment did not establish a systematic improvement in reasoning or confidence.
How should AEO teams test evidence quality?
Start with the exact claim in the answer, map it to a buyer requirement and compare it with the cited passage. Then classify the relationship as direct support, partial support or unsupported inference, verify the underlying product truth and retest with comparable prompts and settings.
What should Marketing Engineering automate?
Automate repeated collection, source extraction, versioning, claim comparison and issue routing. Keep buyer relevance, source sufficiency, product truth and commercial priority under reviewed human judgment.
Research scope: one buyer scenario, three fixed brands and three repetitions per engine-condition combination. Brand order changed with repetition. Perplexity's free interface did not expose its model. Disabling browsing did not remove learned knowledge. No buyers participated, and no purchase progression, pipeline or revenue was measured. The article uses verified selection and confidence results plus individually audited examples; a full frequency analysis of qualitative errors is not claimed.
Also read:
Part 1 - APPEAR: Being Cited Is Not Being Chosen: How AI Answer Engines Build B2B Consideration Sets
Part 2 - PREFER: From Consideration to Preference: How AI Answer Engines Choose Among Credible B2B Brands
