Back to Blog

Being Cited Is Not Being Chosen: How AI Answer Engines Build B2B Consideration Sets

The first decision state in AI-mediated buying.

Michelle Perkins, Founder of ValueTempo

Being cited is not being chosen. APPEAR: how AI answer engines build B2B consideration sets, from public evidence through the decision frame to the consideration set.

TL;DR

  • AI visibility is only the first gate. In 12 fresh-session runs across four AI answer engines, the same brand appeared in the final three in 11 of 12 runs and was selected as the #1 choice in 10 of 12. Yet all four answer engines produced more than one final-three consideration set across their three repetitions, and two also changed which brand they selected as the #1 choice. In some runs, the underlying buyer framing, solution category, evidence path, and tradeoffs changed as well. A stable shortlist can hide unstable decision logic.
  • Before comparing brands, an AI answer engine has to construct the buying problem. It may infer who the buyer is, which operating model fits, what category of solution belongs in the comparison, and which brands are eligible. ValueTempo calls this first decision state APPEAR: does the brand reliably enter the right consideration set for the right buying context?
  • Citation optimization helps a brand get seen. Evidence architecture helps it get understood correctly. The longer-term GTM challenge is to make buyer fit, operating-model fit, capabilities, corroboration, and credible boundaries legible across the different sources AI answer engines use.

Getting cited positively in AI answers is what many AEO programs optimize for today. But visibility only tells you that a brand was surfaced. It does not tell you whether the answer engine understood the brand well enough to place it in the right buying context, include it in the final consideration set, or select it as the #1 choice.

That is the gap we wanted to study.

What the August benchmark revealed

Our August benchmark of 12 AI-search visibility platforms started with the customer problem the category is built to solve: improving visibility in AI answers.

Ten of 12 platforms primarily meter observation-oriented units such as prompts, queries, audits, credits, seats, or similar activity. At the same time, we saw products adding diagnosis, recommendation, and execution capabilities designed to help customers improve that visibility.

That progression raised a downstream question:

What happens after a brand becomes visible?

Does it enter the right consideration set? Does the answer engine understand what problem the product solves, what kind of solution it is, and which buyer it fits? And when several credible alternatives are present, what determines which brand gets selected as the #1 choice?

Across the 12 platforms, we also found insufficient public evidence for a buyer to independently evaluate recommendation quality through a stated validation methodology, baseline or holdout, accuracy or lift, or comparable outcome proof.

So we moved downstream from visibility and tested the buying decision itself.

Want the category-level findings behind this research?

Explore the August AI Search & AEO Platform Benchmark Executive Brief, including the 12 companies analyzed, category patterns, and the buyer-confidence gaps we found.

Explore the August Benchmark

A working model for AI-mediated buying

By AI-mediated buying, I mean a buying process in which AI answer engines help frame the problem, identify candidate brands, compare evidence, and influence the shortlist before or alongside human evaluation.

To make that buying process easier to study, ValueTempo organizes it into four cumulative decision states:

APPEAR: Does the brand reliably enter the right consideration set for the right buying context?

PREFER: Among credible alternatives, does the brand become the #1 choice for the right reasons?

PROVE: Is the evidence behind that preference credible, traceable, and strong enough to justify the decision?

PERFORM: Does that preference translate into buyer action and commercial impact?

These states are cumulative, not linear. Evidence can influence more than one state at once.

The 4 answer-engine decision states: APPEAR, PREFER, PROVE, PERFORM, what AI answer engines are really deciding during AI-mediated buying

This first article focuses on APPEAR.


What does APPEAR mean?

APPEAR asks whether a brand reliably enters the right consideration set for the right buyer, problem, and solution context.

Retrieval alone does not establish consideration. A brand can be cited, mentioned, or even described positively and still remain peripheral to the buying decision.

Before proof can influence preference, the answer engine first needs enough evidence to understand:

What problem does this brand solve?

What kind of solution is it?

Which buyer and operating model does it fit?

Which alternatives belong in the same comparison set?

That is why APPEAR is the first decision state we tested.


How we tested APPEAR

We started with a high-intent B2B buying question in a category where buyers have multiple credible alternatives and where public evidence plays a meaningful role in buying decisions.

Rather than treat one answer as representative, we used a staged research design:

1. Exploratory progression

We first varied the buyer question from broad category discovery toward final brand selection. The goal was to see where brands moved from being visible or mentioned to being seriously considered.

2. Controlled recommendation replication

We then held one high-intent buying question constant and repeated it three times in fresh sessions across ChatGPT, Perplexity, Gemini, and Claude.

4 AI answer engines × 3 fresh-session runs = 12 recommendation observations.

We tracked the final-three consideration set, the #1 choice, and how the buying problem and recommendation rationale changed across runs.

3. Evidence audit

When the repeated runs exposed an important mismatch, we examined the public evidence available around contrasting brand archetypes rather than assuming why one brand appeared and another did not.

4. Controlled mechanism test

Finally, we made the buyer requirements explicit and asked all four answer engines to compare the same three brands. This let us separate one question from another: was a brand missing because it never entered consideration, or because its public evidence was better aligned to a different buying context?

The purpose was not to estimate a population-level recommendation rate or infer hidden answer-engine reasoning.

It was to distinguish several observable decision states that are often collapsed into one AI-visibility metric:

visibility → consideration → preference → verification

That distinction matters because a brand can perform well at one state without reliably advancing to the next.


What the 12 runs showed

The same brand appeared in the final three in 11 of 12 runs and was selected as the #1 choice in 10 of 12.

At first glance, that looks remarkably stable.

But the consideration set underneath it was not.

Across the three repetitions on each answer engine:

What the experiment actually revealed: five findings from the August benchmark and follow-on recommendation tests, including that top 3 is not the same as selected #1
  • All four AI answer engines produced more than one final-three consideration set.
  • Two of the four changed which brand they selected as the #1 choice.
  • In some runs, the underlying buyer framing, solution category, evidence path, and tradeoffs changed as well.

So the outcome could look stable while the path to that outcome moved.

A stable shortlist can hide unstable decision logic.

That matters for APPEAR because shortlist stability is not only about whether one strong brand keeps showing up. It is also about whether the answer engine consistently understands which brands belong in the buying decision at all.

One case in the VAT compliance category made that distinction especially clear.

One brand had previously been featured in a public AEO case study reporting a substantial increase in AI visibility. In our earlier exploratory testing, that brand appeared in a high-intent final-three shortlist.

When we repeated the same high-intent buying question across 12 fresh-session runs, however, the brand appeared in 0 of 12 final-three consideration sets.

That does not mean the visibility improvement was ineffective. The two observations measure different decision states.

A brand can be visible in a category yet fail to enter consideration for a specific buyer context.

That is why evidence architecture has to make buyer fit, operating-model fit, and requirement coverage legible, not just the category association.

That is the APPEAR problem.

Three forms of stability help separate these outcomes:

Consideration stability: Does the brand reliably enter the final consideration set for the intended buyer and buying context?

Preference stability: Once considered, does the brand reliably become the #1 choice?

Rationale stability: Is the brand being considered and preferred for the right reasons?

“Right reasons” means the answer engine accurately understands the brand's problem fit, buyer fit, category or operating-model fit, differentiated value, and supporting evidence.

The next question was why those consideration sets moved. That took us back to how the buying problem itself was being constructed.


Why positioning can affect eligibility before proof is weighed

Before comparing brands, an AI answer engine has to construct the buying problem.

We saw this directly in the VAT compliance research. Our evidence audit found that three credible brands were publicly legible for quite different operating models: one primarily as an expert-led compliance partner, one as broad enterprise tax infrastructure, and one as a SaaS-focused tax operating layer integrated with modern billing systems.

We then replayed the same high-intent buyer question across four answer engines. The buying frame changed:

  • ChatGPT treated the problem as finding the best tax-compliance fit for a modern SaaS company that retains control of its billing and customer relationship.
  • Perplexity prioritized reducing international tax operations and introduced a Merchant of Record architecture.
  • Gemini emphasized digital/SaaS specialization, subscription billing, and cross-border tax requirements.
  • Claude constructed its shortlist partly around review depth and evidence of SaaS operating fit.

Those different frames produced different consideration sets and, in one replay, a different #1 choice.

We then changed the test. Instead of letting each answer engine infer as much of the buyer context on its own, we made the buyer requirements explicit and asked all four to compare the same three brands.

The buyer was defined as a growing international SaaS company that wanted to remain the merchant of record, retain control of billing and the customer relationship, support B2B and B2C sales, integrate with a modern billing stack, automate tax compliance, and operate with a lean finance team.

Under those more explicit conditions, all four answer engines selected the same brand as the #1 choice.

This does not establish a universal sequence for every AI buying decision. But in this experiment, it gave us observable evidence that how the buying problem and buyer requirements are framed can shape which brands become eligible before detailed product proof is compared.

That means positioning can affect eligibility before preference.


How public evidence shapes consideration

Public evidence does more than prove what a product can do. It also helps define what kind of solution the brand is understood to be.

In the VAT compliance audit, the three anonymized brands were all credible. But their public evidence made them legible in different ways:

  • one was strongest around expert-led registration, filing, advisory, and ongoing compliance;
  • one had the broadest enterprise tax infrastructure, review depth, ecosystem validation, and institutional authority;
  • one had the densest evidence around SaaS specialization, modern billing-stack integration, lifecycle automation, and lean-finance operations.

The difference was not simply how much evidence each brand had. It was what buying context that evidence supported.

That is why evidence architecture matters before preference is even formed.

Different public sources can contribute different parts of that understanding:

  • Owned product and solution evidence can establish what the product does, who it is built for, and how it operates.
  • Ecosystem evidence can show whether the product fits the systems the buyer already uses.
  • Customer and review evidence can make implementation experience, usability, support, and operating fit more visible.
  • Institutional or regulatory evidence can establish credibility for claims where external authority matters.
  • Competitor, editorial, analyst, and community sources can influence how the category, alternatives, and tradeoffs are framed.

None of these sources is automatically “better” in every situation.

The more useful question is:

Is this source credible for the decision job it is being asked to perform?

For APPEAR, the first job of evidence architecture is not to prove that your brand is universally better.

It is to make the intended buyer, problem, operating model, category, and comparable alternatives clear enough that the brand consistently enters the right consideration set.

That requires both owned evidence and earned validation.

GTM implication: architect owned proof first, then earn external validation. Foundation of owned signals, validation of earned signals, and decision impact to monitor

Owned evidence defines the story. Earned evidence helps confirm, contextualize, or challenge it.

The strongest architecture is not every source repeating the same claim. It is different sources reinforcing the parts of the buying story they are most credible to support.

The brand selected as the #1 choice by all four answer engines in our controlled comparison is a useful example. Its public evidence worked across several layers:

  • its own product and solution pages made the SaaS use case, tax automation workflow, and lean-finance positioning explicit;
  • Stripe and Chargebee independently confirmed that the product integrates with the billing systems named in the buyer requirements;
  • customer and review evidence added signals about implementation, usability, support, and day-to-day operating fit;
  • its documentation also surfaced limitations and operating boundaries rather than presenting the product as universally right for every buyer.

No single source carried the whole recommendation. Together, they made the same buyer-fit story legible from different angles.

Citation optimization helps a brand get seen. Evidence architecture helps it get understood in the right buying context.

For developer-centric products, that buying context often changes as the decision moves from technical adoption to organizational selection.


What changes for developer-centric products?

For developer-centric products, the consideration set can form at more than one level.

A developer may first evaluate whether the product can solve the technical problem: API fit, documentation quality, integration effort, performance, reliability, or ease of experimentation.

But the organizational buying decision can introduce a different evidence set: security, governance, pricing predictability, procurement fit, operational risk, and confidence that the product can support production use at scale.

A developer product can win the technical decision before it wins the vendor decision. Those two decisions require different evidence.

That matters for APPEAR because the brand has to be legible in both contexts.

A product may appear in an answer about “the easiest API to prototype with” but disappear when the question becomes “which platform should our company standardize on for production?”

The evidence architecture therefore has to make clear not only what developers can build with the product, but also what makes it a credible organizational choice when the buying context expands.

For developer-centric GTM teams, a useful APPEAR question is:

Are we entering the right consideration set when the buyer moves from technical adoption to organizational selection?

B2B SaaS and developer-product SaaS are weighted differently: positioning fit comes first for both, but the evidence mix shifts by product type and decision after eligibility

The GTM implication: evidence architecture should follow decision architecture

Getting cited is an AEO tactic. The longer-term GTM challenge is making sure the public evidence around your brand supports the buying decision you want to enter.

For APPEAR, that means asking:

  • Which buyer and problem do we want to be considered for?
  • Which operating model should an answer engine understand us to support?
  • Which alternatives should we be compared against?
  • What owned evidence makes that positioning explicit?
  • What ecosystem, customer, review, or other external evidence reinforces it?
  • Where might public evidence be placing us in a different buying context?

The goal is not simply more citations or more evidence. It is a public evidence environment that makes the brand accurately legible for the intended buying decision.

So instead of asking only:

“Are we visible?”

ask:

“When an AI answer engine constructs the buying problem, does it understand us well enough to put us in the right consideration set for the buyer and requirements we are trying to win?”

“What evidence is creating that understanding, and where is it coming from?”

That is APPEAR.

The next decision state is PREFER: once several credible brands make the right consideration set, what makes one consistently become the #1 choice, and for the right reasons?

Where is your GTM system losing buyer confidence?

A full ValueTempo Buyability Assessment examines the public evidence around your product to identify where buyers and AI answer engines may struggle to understand your fit, evaluate your claims, compare you with alternatives, or justify the decision.

We look across positioning, buyer and operating-model fit, pricing and packaging, product evidence, ecosystem signals, customer proof, risk, and the handoffs between them.

Contact us for a full Buyability Assessment to identify the gaps in your GTM system that may be making your product harder to understand, evaluate, justify, or buy.

Request a Buyability Assessment

Frequently asked questions

What is the APPEAR decision state in AI-mediated buying?

APPEAR is the first of four cumulative decision states ValueTempo uses to describe AI-mediated B2B buying: APPEAR, PREFER, PROVE, and PERFORM. APPEAR asks whether a brand reliably enters the right consideration set for the right buyer, problem, and solution context.

Does being cited by ChatGPT, Perplexity, Gemini, or Claude mean a brand will be selected as the buyer's top choice?

Not necessarily. In ValueTempo's testing, a brand appeared in the final three consideration set in 11 of 12 runs and was selected as the #1 choice in 10 of 12, but all four AI answer engines produced more than one final-three consideration set across their three repetitions, and two changed which brand they selected as the #1 choice. A stable shortlist can hide unstable decision logic.

Can how a buying problem is framed affect which brands become eligible for consideration?

Yes. In ValueTempo's research, the same high-intent buyer question produced different consideration sets, and in one replay a different #1 choice, depending on how each AI answer engine interpreted the buyer's operating model and requirements. When the buyer requirements were made explicit and all four answer engines compared the same three brands, all four selected the same brand as the #1 choice.

What changes when a developer-centric product moves from technical adoption to organizational selection?

The evidence buyers weigh most heavily changes. Technical evaluators often weigh API documentation, integration effort, and reliability most heavily, which can be enough to win the technical decision. Organizational selection typically introduces a second evidence layer: security, governance, pricing predictability, and procurement fit. A developer product can win the technical decision before it wins the vendor decision.


Method note

This research used a staged exploratory design around one high-intent B2B buying problem in the VAT compliance category.

The research sequence included:

  • an exploratory progression of buyer questions from category discovery toward final selection;
  • a recommendation-stability test across four AI answer engines × three fresh-session runs = 12 recommendation observations;
  • a public-evidence audit comparing contrasting brand archetypes;
  • a natural retrieval replay to observe differences in candidate sets, buying frames, and stated evidence use;
  • and a controlled comparison in which the buyer requirements were made explicit and all four answer engines evaluated the same three brands.

The goal was to distinguish visibility, consideration eligibility, preference, and verification without claiming access to hidden reasoning or internal retrieval telemetry.

These tests support findings about the decision patterns observed in this experiment. They do not establish population-level behavior, universal evidence weights, or general recommendation rates across AI answer engines.

Brand identities are anonymized in this article because the purpose is to study the buying mechanism, not publish a provider ranking.

ValueTempo will continue testing these mechanisms across future benchmark categories, buyer problems, answer engines, and practitioner interviews.