Back to Blog

AI Search Visibility Platforms Are Moving From Measurement to Action. Commercial Evidence Hasn't Caught Up Yet.

What 12 AI search visibility and AEO companies reveal about pricing, recommendation proof, agent readiness, and the next layer of the category.

Michelle Perkins, Founder of ValueTempo

TL;DR

The AI search visibility category is expanding from measurement into diagnosis, recommendation, agent readiness, and execution, but the commercial evidence is not evolving at the same pace.

  • 10 of 12 companies still meter primarily on observation-based units, even as products move closer to recommending and executing work.
  • Recommendation capability is ahead of recommendation proof. None of the 12 publishes enough evidence for a buyer to independently evaluate recommendation quality at scale.
  • The category increasingly looks like three buyer jobs: Visibility Measurement → Diagnosis, Recommendation & Optimization → Agent Readiness & Execution. The next competitive layer may be connecting those jobs into a continuous loop that observes, recommends, executes, measures, and learns.

The GTM implication: as AI products do more of the work, pricing, packaging, proof, and buyer-facing evidence need to make that expanded value easier to understand and trust.

The full August 2026 AI Search Visibility & AEO Benchmark covers all 12 companies, individual scores by dimension, company-level evidence, and the GTM questions we think operators should watch next.

Download the August 2026 Benchmark Executive Brief →

AI search visibility started with a fairly simple buyer question:

Are we showing up in AI-generated answers?

That question is getting bigger.

Across the 12 companies in ValueTempo's August 2026 AI Search Visibility & AEO Benchmark, products increasingly diagnose why visibility changes, recommend what teams should do next, and in some cases help execute those actions.

The commercial evidence buyers can see is moving more slowly.

10 of the 12 companies we analyzed still meter primarily on observation-based units such as prompts, audits, credits, seats, or similar forms of consumption. At the same time, none of the 12 publishes enough evidence for a buyer to independently evaluate recommendation quality.

So the category has an interesting GTM tension:

Products are moving toward recommendation and action, while much of the public commercial evidence still explains measurement and consumption.

That does not tell us whether the products are good or bad. The benchmark measures something narrower: what a buyer can independently understand and verify before speaking with sales.

And right now, that gap is where the interesting questions are.


What we analyzed

For the August benchmark, ValueTempo evaluated 12 AI search visibility and AEO companies across eight dimensions of publicly observable commercial evidence:

Ahrefs, AthenaHQ, Botify, Conductor, Goodie AI, HubSpot, Otterly.AI, Peec.ai, Profound, Scrunch AI, Semrush, Similarweb.

The analysis asks one question:

What can a buyer independently understand and verify before speaking with sales?

The benchmark does not measure product quality, customer satisfaction, internal execution, brand strength, commercial strategy, or market leadership.

Across the eight evidence dimensions, the companies averaged 51% of the available evidence score. ICP and Job Clarity was the strongest dimension. Safety Rails & Trust Surfaces was the weakest, followed by Value Unit clarity.

But the more useful story sits underneath the aggregate score.

A note on the sample: AirOps is an important company in this category, but it was not part of the fixed 12-company scored roster. We are not adding it retroactively because doing so would require recalculating the benchmark. We do reference AirOps research separately as category context, and it belongs on the list for a future update.


One category is beginning to separate into three buyer jobs

The products we studied increasingly span three distinct jobs.

One Category, Three Emerging Buyer Jobs — Visibility Measurement, Diagnosis Recommendation and Optimization, Agent Readiness and Execution

1. Visibility Measurement

Buyer question: Are we visible, and how do we compare?

This is the most established layer.

Products monitor brand mentions, prompts, citations, visibility, competitors, and changes in how AI systems represent a company or category.

The commercial model is familiar too: prompts tracked, seats, projects, audits, credits, queries, and monitoring volume.

2. Diagnosis, Recommendation & Optimization

Buyer question: Why did visibility change, and what should we do next?

Visibility data alone does not tell an operator what to change.

Increasingly, platforms are expected to diagnose the cause of a visibility gap and recommend actions that might improve it.

That is a meaningful step. The product is no longer only reporting what happened. It is beginning to exercise judgment about what should happen next.

3. Agent Readiness & Execution

Buyer question: Can agents understand, trust, and act on our content?

The third layer goes further.

Now the question becomes whether a company's content, data, and systems are structured so agents can reliably interpret and act on them.

The benchmark found real technical signals here, including MCP servers, agent-experience platforms, and other agent-facing infrastructure. But packaging, pricing, and buyer-facing value explanation are much less settled.

These three layers are an emerging model, not a fixed taxonomy. Still, they help explain why two companies can look similar at the score level while making very different commercial choices.


The product may be moving toward action. The meter often isn't.

10 of 12 companies primarily meter observation-based units — the product may be moving toward action, the meter often isn't

We found at least four metering philosophies across the category:

  1. Seat or platform access
  2. Volume or consumption
  3. Outcome proxies
  4. Action units

Most companies remain closer to the observation side.

That means tracking prompts, running audits, counting credits consumed, provisioning seats or users, and metering API calls.

Those units are easy to count and familiar to buy. But as the product moves toward recommending and executing work, a harder question appears:

Should the commercial unit move closer to the work delivered too?

Goodie AI and Profound stand out in the current sample because their public evidence moves closer to action-oriented units. The other ten primarily meter observation or consumption.

That does not make action-based pricing automatically better. It simply exposes a familiar monetization problem:

The easiest thing to count is not always the thing buyers are willing to pay for.

For product and GTM leaders, that puts more pressure on the value unit. A useful unit should help the buyer understand what is counted, what triggers consumption, how usage relates to value, and how cost can be estimated and reconciled.

A credit, after all, is only friendly if the buyer can tell where it went.


Recommendation is a stronger promise than measurement

Measurement asks:

What happened?

Recommendation asks:

What should I do about it?

That second question raises the evidence bar.

Across all 12 companies, we did not find enough public evidence for a buyer to independently assess recommendation quality using signals such as:

  • accuracy, precision, or lift
  • baseline or holdout comparisons
  • documented outcome evidence
  • a stated validation methodology
Recommendation is a stronger promise than measurement — Measure, Diagnose, Recommend, Validate, with Validate flagged as the public evidence gap

That does not mean the recommendations are ineffective.

It means their quality is not yet independently evaluable from the public evidence in the benchmark.

For product leaders, this is partly a validation problem. For GTM leaders, it is an evidence-design problem.

An operator increasingly wants to know: Why this recommendation? What evidence supports it? What should change if it works? And how will the system learn from what happens next?

That matters because a recommendation without a feedback loop is advice. It is not yet learning.


Technical capability does not automatically make customer value legible

Goodie AI and Conductor both publish MCP server access, but their public commercial surfaces explain that capability very differently.

Goodie AI ties MCP access to published tiers and action-oriented commercial units.

Conductor publishes MCP capability, but the benchmark found substantially less public evidence around unit definitions and usage pricing.

Ahrefs also publishes an MCP server, paired with one of the most decomposed unit structures in the sample.

Technical capability does not automatically make customer value legible — Goodie AI, Conductor, and Ahrefs compared on agent-facing signal, public unit, and what's legible

The point is not that one implementation is inherently right.

It is that MCP access is a capability, not a business model.

The same pattern appears with other agent-facing infrastructure. Product teams may understand exactly why a capability matters. Buyers still need help connecting it to a job, a value change, a package, and eventually an economic model.


What the scores hide

A benchmark score is useful because it compresses evidence.

But compression is also the problem.

Ahrefs and Profound, for example, both scored 10/16, but they expose different strengths.

Ahrefs provides unusually granular unit decomposition.

Profound was the only company to reach the benchmark's highest standard for the Value Unit dimension because its public evidence gives buyers stronger visibility into the unit across the usage lifecycle.

Same score, different evidence pattern — Ahrefs and Profound both scored 10/16 but expose different strengths across value unit, technical signal, visibility, and pattern

Those are different GTM design choices.

That is why this benchmark is best read as an evidence maturity diagnostic, not a ranking of company strategy.

The scores tell us where evidence is strong or weak.

The patterns underneath tell us where the category may be heading.


Three GTM questions to watch next

The research leaves us with three questions that matter more than the leaderboard.

1. How should pricing evolve as products move from observation to action?

Most companies in the benchmark still charge primarily for observation or consumption. But as products increasingly recommend and execute work, buyers need a clearer connection between the unit they pay for and the work they receive.

GTM Question 1: How should pricing evolve as products move from observation to action

Profound provides a useful example of stronger value-unit visibility.

Credits are used for Agent consumption, with usage varying according to Agent complexity. Before an Agent runs, the customer sees an estimated credit cost. After the run, actual credits consumed are reflected in the balance. An admin panel aggregates usage across agents, and customers are notified as they approach their credit limits.

That creates a chain a buyer can follow:

unit definition → pre-use estimate → actual consumption → account-level visibility → limit notification

Goodie AI represents another direction: its published pricing moves the meter closer to work delivered by including defined numbers of optimization actions.

What good looks like

A buyer should be able to answer, before and after use:

  • What are we paying for?
  • What causes that unit to be consumed?
  • How much will this work probably cost?
  • What was actually consumed?
  • Can we reconcile usage with the value delivered?
  • What happens when we approach or exceed the limit?

The question is not whether every AI search platform should adopt action-based pricing.

The harder question is why, when, and under what conditions moving the commercial unit closer to the work delivered makes sense.

2. How should agent readiness be commercialized before the value model has settled?

Agent-facing infrastructure is becoming real before the category has settled on how to package or price it.

GTM Question 2: How should agent readiness be commercialized before the value model has settled

Scrunch AI is a useful example.

Its Agent Experience Platform detects agentic traffic and serves optimized content to AI agents. Its Agent Analytics surface shows which agents retrieve content, which pages they rely on, and how AI traffic changes over time.

That is a substantive new product layer.

But it is not yet commercialized as a distinct agent-readiness unit. Scrunch's confirmed public billing still revolves around tracked prompt volume, site audits, and seats.

The capability has changed faster than the commercial unit around it.

What good looks like

A buyer encountering an agent-readiness capability should be able to answer four questions:

Who is it for? Which buyer, operator, developer, or agent workflow needs it?

What does it enable? What can an AI agent now retrieve, understand, or execute that it could not before?

What customer value changes? Does it improve discoverability, reduce manual work, increase successful agent interactions, accelerate execution, or something else?

How is that value commercialized? Is the customer paying for access, traffic, consumption, successful actions, or another unit?

A technically impressive capability should become legible as customer value, rather than asking the buyer to do the translation.

3. How do you make recommendation value legible before recommendation quality can be validated at scale?

This is where the category could become more than a measurement market.

GTM Question 3: closed AEO learning loop — observe, measure, diagnose, recommend, execute, observe again

AEO systems increasingly contain the pieces needed for a closed learning loop:

Observe → Measure → Diagnose → Recommend → Execute → Observe again

The recommendation should not be the end of the workflow.

Once an action is executed, the system can observe what changed, measure the result, compare it with the prior state, and use that evidence to inform the next recommendation.

And because answer engines and competitive landscapes keep changing, the loop cannot be static.

A well-designed system should ideally stay aware of:

  • changes in how answer engines respond
  • shifts in citations and source selection
  • competitor visibility and positioning
  • changes in the customer's own content and authority
  • whether previously recommended actions actually changed observed outcomes

That makes AEO less like a reporting dashboard and more like a continuous learning system.

AthenaHQ shows both the progress and the remaining gap

AthenaHQ's public evidence makes parts of the recommendation chain unusually visible.

Its recommendation engine says recommendations are mapped to the passages and sources AI models actually use within a category. Its Ask Athena copilot similarly says answers are grounded in the customer's data.

That gives the buyer some visibility into why a recommendation exists.

AthenaHQ also publishes named customer outcome examples, including changes in Share of Voice, AI-search-driven demos, and AI Overview impressions.

But the public evidence does not close the entire chain.

We did not find a published precision or accuracy measure for the recommendation engine, a holdout comparison, or a stated methodology connecting a specific recommendation to an observed outcome.

So the current tension looks like this:

reasoning is becoming visible → outcomes are being reported → causal validation remains difficult to evaluate independently

What good looks like

Imagine the system observes that a brand is losing citation share for an important topic.

A well-designed AEO system could make the chain visible:

Observe — Citation share for a priority query cluster falls.

Diagnose — Competing sources are now being cited more frequently, and they contain evidence, structure, or authority signals the customer's content lacks.

Recommend — Add specific evidence, restructure a section, strengthen a source, or create a supporting asset addressing the gap.

Explain — Show which passages, sources, competitors, or answer-engine behaviors triggered the recommendation.

Execute — Make or facilitate the approved change.

Measure — Track what happened to citations, visibility, qualified traffic, or another relevant outcome.

Learn — Use the result to influence the next recommendation as the answer-engine and competitive environment continue to change.

That changes the customer experience from:

“The AI told me what to do.”

to:

“The system can show me what changed, why it recommended this action, what happened after we acted, and what it learned from the result.”

This flywheel is our ValueTempo interpretation of where the category could evolve, not a benchmark finding that any of the 12 companies has already closed the loop publicly.

The benchmark shows pieces of that loop emerging. The next competitive layer may be the ability to connect them.


What GTM teams can audit now

You do not have to wait for the category to settle.

Five places AI-native GTM teams can pressure-test their commercial evidence today

The benchmark suggests five places AI-native teams can pressure-test today.

1. Define the billable unit

Ask: Can a prospect clearly identify what they are paying for?

What good looks like: The unit has a clear name, boundary, consumption rule, and relationship to customer value. A buyer can tell what counts, what does not, and why the unit is economically meaningful.

2. Make spend easier to estimate

Ask: Can a buyer anticipate likely cost before committing?

What good looks like: The buyer can estimate usage before purchase or before a meaningful action occurs, then compare estimated consumption with actual consumption afterward.

3. Make recommendation logic inspectable

Ask: Can buyers understand how recommendations are produced and what evidence supports them?

What good looks like: Recommendations show the conditions that triggered them, the rationale behind the proposed action, and what observable outcome should change if the recommendation works.

4. Document usage limits and controls

Ask: Are caps, guardrails, and override mechanisms visible before the sales conversation?

What good looks like: Customers can see included usage, thresholds, overage behavior, alerts, administrative controls, and mechanisms to constrain cost or autonomous activity.

5. Clarify ownership and rights

Ask: Can buyers easily find the terms governing data, outputs, and IP?

What good looks like: The commercial surface clearly explains who owns customer inputs, generated outputs, and derived data; what may be retained or reused; and where important legal or model-training boundaries apply.

These are buyer-facing evidence problems.

Increasingly, they are also GTM design problems, whether the company operates through PLG, enterprise sales, or a hybrid motion.

In a PLG motion, the product and commercial surface have to explain value, usage, limits, cost, and next steps without a salesperson filling in the gaps.

In an enterprise sales motion, the same logic has to hold across sales, solution engineering, security, procurement, and executive sponsors. Not every detail has to be public, but the value story cannot change every time the buyer enters a new meeting.

In a hybrid motion, continuity matters most. The value logic a buyer encounters during self-service exploration should still make sense when the account moves into sales-assisted evaluation, procurement, implementation, and expansion.

The motion changes where and how the explanation happens. It does not remove the need for the underlying commercial logic to hold together.


The category is moving faster than its commercial language

AI search visibility is no longer only about visibility.

Products are beginning to diagnose, recommend, optimize, and act. Agent-facing capabilities are appearing. Execution is becoming part of the product promise.

Now the commercial system has to catch up.

For CEOs, that is a business-model question.

For product leaders, it is a product-system and validation question.

For GTM leaders, it is a positioning, packaging, evidence, and buyer-journey question.

The three functions increasingly meet at the same point:

Can the customer understand what the system does, why it matters, how value compounds, and what evidence should make them trust it?

As AI makes execution easier to build, durable advantage may increasingly depend on something harder:

making the value understandable, inspectable, measurable where possible, and trustworthy before the buyer speaks with sales, then keeping that logic intact after they become a customer.


The AVS Benchmark is produced by ValueTempo. Scores reflect publicly observable, buyer-facing evidence as of August 2026. This report does not constitute an endorsement or recommendation of any scored company.

Get the full August 2026 AI Search Visibility & AEO Benchmark

All 12 companies. Individual scores by dimension. Company-level evidence and the GTM questions we think operators should watch next.