Show HN: Nari Qwen3-TTS 和 Qwen3-ASR —— 高精度、低延迟且低成本
Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost

原始链接: https://narilabs.com/blog/nari-labs-leads-coval-voice-ai-benchmarks/

截至2026年9月,Nari Labs 已正式登顶 Coval 语音 AI 基准测试榜单,在语音转文字(STT)和文字转语音(TTS)模型的质量、延迟与成本平衡方面均处于行业领先地位。 在 STT 领域,Nari Labs 的 Qwen3-ASR Fast 模型延迟排名第 1(首字响应时间 44ms),准确率排名第 2(词错误率 3.6%),且价格远低于 AssemblyAI 和 Deepgram 等竞争对手。 在 TTS 领域,Qwen3-TTS Fast 模型延迟排名第 2(首字音频响应时间 63ms),准确率排名第 1(词错误率 3.8%)。尽管性能卓越,它仍是市场上最具成本效益的解决方案,价格比 ElevenLabs 和 Cartesia 等竞品便宜 6.5 倍。 Nari Labs 通过优化“帕累托前沿”(Pareto Frontier),以行业标准价格的一小部分提供卓越性能,持续超越标准模型部署和专用端点。该公司目前正将其公测 API 转为正式商用,并为新用户提供免费试用额度。

Nari Labs 发布了一款针对 Qwen3-TTS 和 Qwen3-ASR 优化的开源推理引擎,旨在证明开源语音模型在速度、成本和准确性方面可以超越闭源方案。 针对现有推理系统(如 vLLM)在处理多模态数据时的局限性,Nari Labs 开发了一款专用引擎,实现了低于 50 毫秒的延迟。根据 Coval 语音 AI 基准测试,其 Qwen3-TTS 端点目前在准确性方面排名第一,延迟排名第二,同时保持了业内最低成本。同样,其 Qwen3-ASR 模型也提供了业内领先的低延迟和高水准准确性。 通过对推理工程的优化,Nari Labs 力求使语音技术普及化,让开发者无需承担高昂的单次使用成本,即可集成高质量的语音合成(TTS)和语音识别(STT)功能。该团队计划在不久的将来将重点扩展至说话人日志(diarization)、视频推理和世界模型领域。
相关文章

原文

Research

By the Nari Labs Team · Sep 14, 2026

TL;DR

Nari Labs leads Coval’s voice AI benchmark by sitting on the quality-latency Pareto Frontier for both Text-to-Speech and Speech-to-Text. We also lead the latency-cost and quality-cost Pareto Frontier out of all publicly available models on the benchmark.

Coval is a leading provider of voice AI evaluation and benchmarks. They help speech AI agents perform better in production and publish one of the most widely cited benchmarks in the industry.

The Text-to-Speech (TTS) benchmark evaluates latency from text input to first audible chunk of audio (time-to-first-audio or TTFA) and Word Error Rate (WER). The Speech-to-Text (STT) benchmark evaluates latency from user’s finalize request to the final text output (time-to-final-segment or TTFS) and Word Error Rate (WER).

TTFA and TTFS are critical for voice agents, where latency can make a voice AI agent feel unresponsive. Low WER is an obvious key factor for model performance as well.

As of mid September 2026, Nari Labs tops both the Speech-to-Text and Text-to-Speech benchmarks. STT: #1 Latency, #2 WER. TTS: #2 Latency, #1 WER. Note that Coval’s benchmarks can fluctuate every 30 minutes*. We only include publicly available endpoints in our rankings and charts.

Speech-to-Text

Our Qwen3-ASR Fast model is ranked #1 in Time-to-Final-Segment (TTFS), at p50 of 44 ms and WER of 3.6%, placing #2 behind AssemblyAI’s Universal 3.5 Pro at 3.5%.

The pricing makes it even better. At $0.12 / hour, our Fast endpoint ties for the 2nd-lowest price among models with known public rates in Coval’s pricing directory. Universal 3.5 Pro costs 3.75× more, and Deepgram Nova 3 costs 2.4× more. Our Standard endpoint would be the cheapest at $0.06 / hour.

Coval STT latency and accuracy: Nari Qwen3-ASR Fast on the Pareto frontier at 44 ms median TTFS and 3.6% WER.Nari Labs
Coval STT word error rates: Nari Qwen3-ASR Fast ranks second at 3.6%, behind AssemblyAI Universal 3.5 Pro at 3.5%.Nari Labs

Text-to-Speech

Our Qwen3-TTS Fast model is ranked #2 in Time-to-First-Audio (TTFA), at p50 of 63 ms and WER of 3.8%, coming in at #1.

The only model with a lower median TTFA than ours is vui from Fluxions, at 49 ms. It is a 300M parameter model, compared to the 1.7B Qwen3-TTS that we serve.

At $10 per 1M characters, our Fast endpoint is tied for the #1 cheapest model on Coval’s pricing directory. ElevenLabs Eleven v3 Conversational costs 5x more, and Cartesia Sonic 3.6 costs 6.5x more. Our Standard endpoint would be the cheapest at $5 per 1M characters.

Coval TTS latency and accuracy: Nari Qwen3-TTS Fast on the Pareto frontier at 63 ms median TTFA and 3.8% WER.Nari Labs
Coval TTS word error rates: Nari Qwen3-TTS Fast leads at 3.8%.Nari Labs

Interestingly, the official Qwen3 TTS Flash Realtime endpoint sits at 8.8% WER and 692 ms median TTFA. We both serve the same model.

We also surpass Baseten’s dedicated Qwen3-TTS endpoint, which records 6.0% WER and 101 ms median TTFA.

Get Started

Try both our Speech-to-Text and Text-to-Speech models for free for a limited period of time. We are moving our Public Beta APIs to a paid GA within this week and will provide $20 in credits for everyone who has created an account when the switch happens.

Try STT and TTS

Need help meeting the latency and capacity requirements of your voice application? Talk to our engineers

联系我们 contact @ memedata.com