From Consideration to Preference: How AI Answer Engines Choose Among Credible B2B Brands
The second decision state in AI-mediated buying.
Michelle Perkins, Founder of ValueTempo

TL;DR
- Brand A made the final three in 11 of 12 runs and was selected #1 in 10. Brand B made the final three in 10 of 12 and was selected #1 zero times.
- The same buying question produced different buyer priorities and evidence paths across the four answer engines. ChatGPT, Gemini, and Claude selected Brand A; Perplexity selected Brand C.
- When we held the compared brands constant and made the buyer requirements explicit, all four selected Brand A. In this experiment, clearer buyer context reduced recommendation variance.
- For GTM teams, the working model is to support five evidence jobs: fit, capability, corroboration, tradeoffs, and decision risk, then test and refine the evidence system over time.
In Part 1 of this four-part series on AI visibility-to-revenue, APPEAR, we used the same high-intent B2B buying question across ChatGPT, Claude, Gemini, and Perplexity to examine whether a brand reliably entered the right shortlist.
PREFER starts once several credible brands are already in consideration. The question gets narrower: what makes one become the #1 choice?
One result from the VAT compliance category experiment made that question hard to ignore.
Brand A made the final three in 11 of 12 runs and was selected as the #1 choice 10 times.
Brand B also made the final three in 10 of 12 runs. It had a broader institutional footprint, more independent review evidence, deeper documentation, major ecosystem integrations, and substantial market presence. It was selected #1 zero times.
That contrast gave us a cleaner research question: why did one credible brand keep winning the #1 spot while another never did?

PREFER is the second decision state: among credible options, which brand becomes the #1 choice?
Want the category-level findings behind this research?
Explore the August AI Search & AEO Platform Benchmark Executive Brief, including the 12 companies analyzed, category patterns, and the buyer-confidence gaps we found.
Explore the August BenchmarkAEO research is moving closer to the buying decision
Recent AEO research is starting to distinguish whether a brand is named, cited, or actually recommended.
Heft recently tested 12,400 purchase-intent questions across 41 B2B categories and classified whether AI answers named a brand, cited a source, or made an outright recommendation. Only 31% of the questions produced agreement on the top recommendation across the six engines they tested.
For GTM teams, that creates a more useful progression: mention → citation → recommendation. Each step moves the research closer to the buying decision.
Source: Heft, AEO Report
A recent demand-side GEO research paper on arXiv pushes further into buyer context. Its framework combines what buyers ask, what information they need, which sources they trust, and recommendation data. Controlled empirical testing remains future work.
Source: Demand-side GEO research paper
Taken together, these studies point to the question we wanted to examine next: once several credible brands are available, how does the buying context shape which one becomes preferred?
The same question produced different buying frames
We went back through the answer structures and evidence traces from the original experiment to look beyond the final rankings.
The same question produced somewhat different buyer priorities and evidence paths before the #1 choice was selected.
ChatGPT
Buyer priority: retain control of billing and the customer relationship while modernizing global SaaS tax operations.
Evidence emphasis: SaaS-specific fit, capability evidence, customer evidence, and boundary evidence.
Selected #1: Brand A.
Perplexity
Buyer priority: reduce international tax and compliance operations, including consideration of a Merchant of Record architecture.
Evidence emphasis: category and operating-model framing, policy context, and ecosystem/editorial evidence.
Selected #1: Brand C.
Gemini
Buyer priority: digital and subscription SaaS with specialized cross-border tax requirements.
Evidence emphasis: SaaS fit, review evidence, and broader product knowledge.
Selected #1: Brand A.
Claude
Buyer priority: growth-stage SaaS operating concerns across credible VAT/GST providers.
Evidence emphasis: review depth, setup, support, usability, and operating fit.
Selected #1: Brand A.
These are observable differences in answer structure and evidence traces. They do not reveal hidden reasoning or establish permanent weighting rules for any engine.
Perplexity provides the clearest example of why the buying frame matters. It treated minimizing international tax operations as the priority and introduced a Merchant of Record architecture. That widened the comparison beyond the direct-tax-compliance model implied by the buyer's wording and led to a different selected brand.
Gemini surfaced a related issue inside the same session. It produced two plausible comparison paths for the identical buyer question: one centered on direct tax automation and another that introduced a Merchant of Record alternative.

Similar shortlists, different buying frames. In this experiment, preference varied with buyer context and evidence emphasis.
This is where decision architecture enters the AEO problem. After retrieval, an answer engine still has to frame the buyer problem, decide which solution types belong in the comparison, and determine which public evidence is relevant.
What happened when we narrowed the buyer context
To test whether clearer buyer context would reduce that variation, we ran a controlled comparison. All four answer engines compared the same three brands against the same explicit requirements. The buyer wanted to:
- remain the merchant of record
- retain control of billing and the customer relationship
- serve B2B and B2C customers internationally
- use a modern billing stack such as Stripe or Chargebee
- handle obligation monitoring, registration, calculation, filing, remittance, and ongoing compliance
- operate with a lean finance team
ChatGPT, Claude, Gemini, and Perplexity all selected the same brand as the #1 choice.
The result is specific to this buyer profile. Under these conditions, Brand A fit the specified operating model best. A broader ranking of VAT compliance providers would require different buyer profiles and tests.
Brand A's public evidence aligned closely with the direct-seller SaaS buyer we had specified: modern billing infrastructure, retained customer ownership, international tax automation, and a lean finance team. Brand B remained a credible enterprise alternative with broader institutional proof. Brand C was more legible as an expert-led compliance partner focused on registration, filing, authority communication, and ongoing support.
With the buying context narrowed, recommendation variance fell. In this VAT compliance experiment, preference aligned closely with how well each brand's public evidence mapped to the buyer's operating model and material requirements. Whether that pattern holds in other categories is now a question we can test.
How to architect evidence for AI-mediated preference
In this test, raw evidence volume did not explain preference well. The better content question is whether public evidence performs the decision jobs the buyer context requires.
For PREFER, our working model has five evidence jobs:
- Fit. Who is this product designed for, and what operating model does it support?
- Capability. Can the product satisfy the requirements that matter to this buyer?
- Corroboration. Do ecosystem partners, customers, reviews, and other credible sources support the same buyer-fit story?
- Tradeoffs. Where are the meaningful limits, exceptions, costs, or conditions that narrow the choice?
- Decision risk. Are implementation requirements, policies, constraints, and commercial edges documented well enough to act?
Different source types can support different jobs. Product pages can establish intended use cases and capabilities. Ecosystem evidence can corroborate compatibility. Customer and review evidence can add operating context. Documentation can verify workflows, conditions, and implementation details.
Boundary and tradeoff evidence deserves particular attention. Buyers need enough public evidence to understand where a product fits best, where its limits begin, and when another operating model may be a better match. In our testing, that limiting evidence sometimes helped define the conditions under which a recommendation remained credible.
This is the pattern we want to test longitudinally: different credible sources performing different jobs around the same buyer-fit story. A product page can establish capability. An ecosystem partner can verify compatibility. A customer can describe operating experience. Documentation can show implementation constraints and limitations.
Peec's recent work helps explain the upstream side of this system. It separates discovery, passage selection, source selection, and answer generation, showing where useful content can fall out before it reaches an AI answer.
Source: Peec, Rerankers for GEO/AEO
Our experiment picks up further downstream by asking what happens once candidate brands and evidence are available: which buying frame forms, which evidence becomes decision-relevant, and which brand becomes preferred.

A GTM operating model connects evidence inputs to five decision jobs, then tests, audits, closes gaps, and retests as answer engines and buying contexts change.
Because answer-engine behavior changes, the durable strategy is to strengthen the evidence system across all five jobs and learn from repeated buying-decision tests.
That creates the closed feedback loop: define the buyer context, architect the evidence, test the buying decision, record the gaps, strengthen the evidence environment, and retest. The learning compounds when the team remembers what changed, how the observed decision responded, which interpretation was tested, and what should be tested next.
Two longitudinal questions now matter to this research:
Which relationships between buyer context, evidence type, and preference remain durable as answer engines change?
How quickly can a GTM team detect that the buying logic has changed and update the evidence system accordingly?
Preference changes with the buying motion
The VAT experiment above reflects a conventional B2B SaaS buying scenario. Developer products introduce another preference question that we want to test across future categories.
For a conventional B2B SaaS buyer, preference may mean which vendor best fits the operating model. For a developer evaluating an API, infrastructure product, or developer platform, technical preference can form earlier around docs, API fit, implementation speed, performance, and developer effort. Security, procurement, pricing, reliability, and organizational approval may come later.
That creates a different path to preference: technical choice → production confidence → internal champion → vendor decision.

A developer product can win the technical decision before it wins the vendor decision. Those two decisions require different evidence.
We will treat this as a research lens for future categories. The VAT experiment did not test developer-product buying motions.
Preference still needs proof
Scrunch gives us one reason to care about preference commercially. Its analysis of linked AI and web behavior found that recommendations were associated with stronger subsequent branded search, site visits, and product-page activity than simple brand mentions, with differences by platform, placement, and framing.
Source: Scrunch, Prompt-to-Purchase Pipeline
A #1 recommendation may be commercially meaningful. The buyer still has to decide whether the evidence is strong enough to act on.
Our own evidence audit gave us a concrete reason to ask that next. In the four initial answer-engine tests, each engine produced a recommendation before its follow-up audit had verified every material buyer requirement.
Those audits surfaced requirements marked Verified by one engine and Partial by another, different thresholds for sufficient evidence, source-provenance problems, and claims that went beyond what the cited source established.
The next buying question is:
Can the evidence behind an AI recommendation actually withstand scrutiny?
That is the next decision state in the series: PROVE.
Where is your GTM system losing buyer confidence?
A full ValueTempo Buyability Assessment examines the public evidence around your product to identify where buyers and AI answer engines may struggle to understand your fit, evaluate your claims, compare you with alternatives, or justify the decision.
We look across positioning, buyer and operating-model fit, pricing and packaging, product evidence, ecosystem signals, customer proof, risk, and the handoffs between them.
Contact us for a full Buyability Assessment to identify the gaps in your GTM system that may be making your product harder to understand, evaluate, justify, or buy.
Request a Buyability AssessmentFrequently asked questions
What is the PREFER decision state in AI-mediated buying?
PREFER is the second of four decision states ValueTempo uses to study AI-mediated B2B buying: APPEAR, PREFER, PROVE, and PERFORM. PREFER begins once several credible brands are already in consideration and asks which one becomes the #1 choice for a specific buyer context.
What did the August VAT compliance experiment show about preference?
Two brands repeatedly made the final three, but their preference outcomes were very different. Brand A made the final three in 11 of 12 runs and was selected #1 in 10. Brand B made the final three in 10 of 12 runs and was selected #1 zero times, despite having broader institutional proof, more independent evidence, deeper documentation, and major ecosystem integrations.
Can clearer buyer context change which brand an AI answer engine prefers?
It did in this experiment. The original buyer question produced somewhat different buying frames and evidence paths across ChatGPT, Claude, Gemini, and Perplexity. When ValueTempo held the three compared brands constant and made the buyer's operating model and requirements explicit, all four answer engines selected the same brand as the #1 choice. This result is specific to that buyer profile and does not establish a universal provider ranking.
What kinds of public evidence can help shape preference?
ValueTempo's current PREFER working model identifies five evidence jobs: fit, capability, corroboration, tradeoffs, and decision risk. Different sources can perform different jobs. Product pages can establish fit and capability, ecosystem and customer evidence can corroborate claims, and documentation can clarify implementation requirements, limitations, and commercial boundaries.
Does PREFER work the same way for developer products?
That remains a research question. The VAT experiment tested a conventional B2B SaaS buying scenario, not a developer-product buying motion. For developer products, technical preference may form earlier through documentation, API fit, implementation speed, performance, and developer effort, while security, procurement, pricing, reliability, and organizational approval may enter later. ValueTempo will test this distinction in future categories.
Method note
This article uses the same staged August VAT compliance experiment introduced in Part 1, but analyzes the research through the PREFER decision state: what happens after several credible brands have already entered consideration.
The research included:
- a recommendation-stability test across four AI answer engines × three fresh-session runs = 12 recommendation observations;
- a review of observable answer structures and evidence traces to compare the buyer priorities, comparison frames, and public evidence emphasized in each answer;
- a public-evidence audit of the leading brand archetypes to compare institutional proof, SaaS operating-model fit, capability evidence, corroboration, and boundaries;
- and a controlled comparison in which all four answer engines evaluated the same three brands against the same explicit buyer requirements.
The controlled buyer profile specified that the company wanted to remain the merchant of record, retain its billing and customer relationship, serve B2B and B2C customers internationally, use a modern billing stack such as Stripe or Chargebee, manage the full VAT/GST compliance lifecycle, and operate with a lean finance team.
The goal was to examine the relationship between buyer context, public evidence, and observed preference without claiming access to hidden reasoning, proprietary retrieval systems, or internal model weights.
These tests support findings about the decision patterns observed in this experiment. They do not establish population-level behavior, universal evidence weights, permanent characteristics of individual AI answer engines, or a universal ranking of VAT compliance brands.
The B2B SaaS versus developer-product comparison in this article is a research lens for future testing, not a finding from the VAT compliance experiment.
Brand identities remain anonymized because the purpose of the research is to study the buying mechanism rather than publish a provider ranking.
ValueTempo will continue testing how buyer context, evidence type, preference, and decision confidence change across categories, buying motions, answer engines, and repeated tests.
