Back to Blog

Why AI Speech Is the Most Buyer-Ready Category We've Benchmarked (And What That Reveals)

Michelle Perkins, Founder of ValueTempo

TL;DR

  • The June 2026 AVS Benchmark scored 12 AI speech companies across STT, TTS, and voice agent sub-types on published commercial evidence.
  • The category averaged 86% — 17 points above May's highest-scoring category — and every company scored Exemplary, a first in the benchmark.
  • Developer-first, usage-based GTM forces commercial transparency. When buyers self-qualify on documentation, the documentation has to be complete enough to close them.
  • Four of eight dimensions scored perfectly across all 12 companies: ICP clarity, buyer alignment, value unit definition, and packaging.
  • The shared gaps sit in D1 (verifiable performance claims, avg 1.33/2) and D7 (production-edge behavior, avg 1.25/2).
  • Consumption-model companies averaged 92.5% vs. 83.0% for hybrid, a 9.5-point gap concentrated in D1 and D7.
  • Even the two 16/16 scorers have named gaps in spend controls, billing disputes, and refund policies the rubric doesn't currently capture.

The full June 2026 AI Speech Platform Benchmark covers all 12 companies, individual scores by dimension, category-specific scoring rules, and sub-segment analysis.

Download the full report →

Every month, the AVS Benchmark scores a category of AI companies on what buyers can independently verify before they engage sales: pricing structure, billing units, packaging clarity, usage controls, and operational trust signals. All drawn from public surfaces (pricing pages, documentation, trust pages) — the same evidence a buyer or AI agent would encounter.

The June benchmark covered AI Speech Platform. Twelve companies. STT providers, TTS platforms, voice agent builders.

The result: 86% category average. Every company scored Exemplary. That has never happened before.

This post explains why, and why the score alone misses the more important finding.


What the AVS Rubric actually measures

Before the findings: a clarification that matters.

The AVS (Adaptive Value System) Rubric measures published commercial evidence: what buyers can verify on pricing pages, documentation, and other public surfaces before speaking with sales. Product quality, feature depth, and customer satisfaction are outside its scope. A company with a 16/16 score may have worse support than a company with a 12/16 score. A company with a lower score may simply have a sales-led motion where pricing is disclosed after qualification.

Higher score = stronger published evidence a buyer can use to independently evaluate, budget for, and justify a purchase.

That's the lens for everything that follows.


Why AI Speech scored 17 points above May's top category

May's benchmark covered five categories: AI Customer Support, AI Agent Platform, AI Coding Assistant, AI Sales Intelligence, and AI Revenue Intelligence. The highest average was 69%. AI Speech landed at 86%.

That gap has a structural explanation.

In AI Speech, the primary buyer reads the documentation before talking to anyone. A developer evaluating Deepgram or AssemblyAI doesn't schedule a discovery call to learn what a credit is worth or what happens when they hit a rate limit. They read the pricing page. They read the docs. They make a judgment and start building.

That buying motion forces commercial transparency. When your buyer self-qualifies using public documentation, the documentation has to be good enough to close them. Lose them there, and they never show up in your CRM.

Compare that to enterprise software categories where pricing is gated, value units are vague, and packaging is explained in a deck on call two. Both approaches can work as a business model. But only one scores well on a rubric measuring what buyers can independently verify.

Developer-first, usage-based GTM structurally requires buyability. That's why AI Speech landed where it did.


Four dimensions that scored perfectly

Across all 12 companies, four of eight dimensions scored 2.0 (the maximum):

  • D2 (ICP & Job Clarity): Every company clearly states who the product is for and what job it serves.
  • D3 (Buyer & Budget Alignment): Plans map to buyer types; budget tier is legible.
  • D4 (Value Unit): The billable unit is named, predictable, and auditable. You know what you're paying for.
  • D6 (Pools & Packaging): Tiers distinguish between exploration and production workloads.

These are the dimensions May's five categories struggled with most. The benchmarks converge on the same point: usage-based, developer-first products solve the basics because they have to.


The two dimensions that didn't

D1 — Product North Star (avg 1.33/2): Eight of 12 companies make performance claims (accuracy rates, latency benchmarks, quality comparisons) that buyers can't independently verify without running a proof of concept. The claim is published. The evidence to verify it is not.

This matters because benchmark claims are becoming a standard part of AI product marketing. When every vendor claims 95%+ accuracy and sub-200ms latency, the signal is noise. The four companies that scored full marks on D1 are all consumption-model companies, and the pattern holds: a single, defined billing unit makes it easier to anchor a verifiable benchmark claim to a specific condition.

D7 — Overages & Risk Allocation (avg 1.25/2): Nine of 12 companies don't document what happens when a production deployment hits a rate limit, exhausts credits, or runs into an overage event. For a developer building on this infrastructure, that's an unquantified production risk. The question isn't whether it will happen. It's whether the vendor has told you what to expect when it does.

These two gaps share a common root: they're the things vendors can defer until after a buyer is already committed. D1 gaps become visible when you're deep enough into evaluation to run a POC. D7 gaps surface post-deployment. Both are the vendor's responsibility to document. Most haven't.


What a 16/16 score still misses

Deepgram and Retell AI both scored 16/16 (the maximum). Both have exceptionally clear pricing, named billing units, explicit packaging, and published overage behavior.

Both still have specific, named gaps.

No published spend caps. No documented billing dispute process. No refund policy for platform errors on the vendor's side. These are the commercial terms an informed buyer will ask about before committing a production workload to a platform.

The rubric is an 8-dimension, 16-point system. It captures what buyers can verify on public surfaces against a defined framework, not every commercial term a sophisticated buyer will surface. The rubric has defined boundaries. A 16/16 score means the rubric's criteria are fully met. The commercial experience may still have gaps.

A perfect score means a buyer can get in. It says nothing about whether the vendor has solved every commercial question a buyer at scale will eventually ask.


The sub-type gap within the Exemplary band

All 12 companies scored Exemplary. But variation within the band is real:

Sub-typeAverage
Full Speech Platform93.8%
Voice Agent Platform89.6%
TTS84.4%
STT81.3%

The 12 companies scored, by sub-type and pricing model:

CompanySub-typePricing model
AssemblyAISTTConsumption
SpeechmaticsFull Speech PlatformConsumption
DeepgramFull Speech PlatformConsumption
Retell AIVoice Agent PlatformConsumption
RimeTTSConsumption
Bland.aiVoice Agent PlatformHybrid (access + consumption)
CartesiaTTSHybrid (access + consumption)
ElevenLabsTTSHybrid (access + consumption)
LMNTTTSHybrid (access + consumption)
Murf AITTSHybrid (access + consumption)
VapiVoice Agent PlatformHybrid (access + consumption)
Wellsaid LabsTTSHybrid (access + consumption)

The gap between Full Speech Platform and STT is 12.5 points, meaningful within a single band. The driver is D5 and D7: documentation depth around cost drivers and production-edge behavior. Full speech platforms serve more complex deployment scenarios and document them accordingly. STT providers, more often pure-API, leave more of that documentation to inference from pricing tables.

Pricing model also predicts score within the category. Consumption-only companies averaged 92.5%. Hybrid (access fee + consumption) companies averaged 83.0%, a 9.5-point gap that concentrates in D1 and D7. A single billing unit makes it easier to anchor a verifiable benchmark claim. A single exhaustion event is easier to document than the two-layer failure modes of a hybrid model.


What this means for GTM operators

If you're building on AI speech infrastructure, the benchmark is a due-diligence tool. D7 scores tell you which vendors have documented production-edge behavior. D1 scores tell you which performance claims have documented evidence behind them. That's relevant before you commit a production workload.

If you're selling AI speech infrastructure, the benchmark identifies which gaps are worth closing now. D1 and D7 are where buyers will push on the next round of evaluation. Vendors who document rate limit behavior and clarify benchmark methodology before buyers ask will have an advantage. Buyers in this category already know to ask.

If you're a GTM operator in a different category, this benchmark shows you the floor. Developer-first, usage-based products have solved the basics because their buying motion required it. If your buying motion doesn't require that, you can defer. But AI agents evaluating vendors don't call sales. They read the documentation. That buyer is coming to every category.


The AVS Benchmark is produced by ValueTempo. Scores reflect publicly observable, buyer-facing evidence as of June 30, 2026. This report does not constitute an endorsement or recommendation of any scored company.

Get the full June 2026 AI Speech Platform Benchmark

All 12 companies. Individual scores by dimension. Category-specific scoring rules and sub-segment analysis.