@FakeMaidenMaker: Perplexity 前两天发布了他们最新的团队分享 :《Perplexity 是怎么借助 Realtime API 将语音搜索带给数百万用户的》 他们用 OpenAI 的 Realtime-1.5 给自家的 AI 浏览器 Comet 加…

X AI KOLs Timeline 新闻

摘要

Perplexity 分享了利用 OpenAI Realtime API 为自家 AI 浏览器 Comet 添加语音功能的工程经验,包括上下文小块喂送、角色管理、音频管线统一等关键技巧。

Perplexity 前两天发布了他们最新的团队分享 :《Perplexity 是怎么借助 Realtime API 将语音搜索带给数百万用户的》 他们用 OpenAI 的 Realtime-1.5 给自家的 AI 浏览器 Comet 加了语音入口,每月处理百万级语音会话,把过程中踩过的坑和总结的工程经验和教训完整地分享了出来。 关键经验有这几条: 1、上下文小块小块喂:大块更新是 all-or-nothing——一个 10000 token 的更新塞进只剩 5000 token 的窗口,模型会丢光所有历史。改成 2000 token 小块增量喂,开销更高但稳定,截断时只裁一点历史而不是清空整段; 2、做好角色管理比 token 消耗管理更重要:对话消息有 system、user、assistant 三个角色。太多东西当 user 喂进去,模型会以为"用户在大声朗读每一段(包括网页片段和评论)";太多当 system 喂,模型又分不清"自己天生知道的"和"临时给的上下文"; 3、浏览场景的正确处理用户输入:用户滚动页面时持续把页面内容更新给模型,全当用户输入就废了; 正确做法是让模型"像在背景里知道这个页面",用户问的时候再回答,而不是把每段文字当成用户在朗读; 4、跨产品音频要统一: Perplexity 有 Ask、Comet、Computer 多个产品,跑在 Swift、TypeScript、Rust、C++ 不同栈上,各自发原生音频给 API 就开始性能漂移。他们用 Rust 写了一个 SDK 把音频契约统一起来; 5、统一的音频流水线: 48 kHz 单声道重采样、匹配 Opus 编码器和 WebRTC 内部采样率、跑 WebRTC APM 做回声消除/自动增益/降噪/高通滤波、最后编码传输。所有客户端走同一条管线,不再各做各的; 6、VAD 要在真实环境调: 他们内部测试用例是一家嘈杂的旧金山酒吧:"用户跟朋友说'你试过新的 Perplexity app 吗',朋友掏出手机一试,如果失败两个用户当场就丢了"。所以宁可先在乱环境调好再上线; 7、用户停顿不要抢话: 用户停下来思考、找屏幕上的东西、准备朗读,模型很容易把停顿当成一轮对话结束。例子:用户问数学推导帮助、停下来找公式,模型抢着回答,但用户问题还没问完; 8、voice lock 反转传统点按说话,默认开启麦克风: 传统 push-to-talk 是默认关麦克风、按住才能说; 他们反过来——默认环境监听一直在,用户需要占住发言权时主动按 voice lock 拿走轮次控制权。Perplexity 认为这个交互模式会成为复杂语音工作流的标配; 9、工具集刻意收窄: 他们只用不到 10 个核心工具,并且在系统提示词里写明每个工具什么时候怎么调,不堆砌; 10、工具输出要像 JSON 不要像对话: 返回结构化 JSON(如 {"response_text": "...", "require_repeat_verbatim": true})而不是把指令和对话混着写。这样更接近模型训练时见过的样子,工具调用更稳定 总结: Perplexity 这套经验真正讲清楚的是:做语音 agent,模型能力只是地基,真正决定体验的是工程层那些不起眼的细节,比如上下文怎么切、角色怎么标、音频管线放在哪一层、停顿怎么处理、工具用什么格式喂。 他们最有洞察的一句话不是某个技术技巧,而是"我们为明天的 Realtime 模型构建今天的系统"。意思是不去等下一代模型解决所有问题,把今天能稳的部分先做稳,等模型升级时所有工程层自动升级一档。 voice agent 想做出"用过之后回不去"的感觉,靠的不是更聪明的模型,是更细致的工程。
查看原文
查看缓存全文

缓存时间: 2026/05/23 00:01

Perplexity 前两天发布了他们最新的团队分享 :《Perplexity 是怎么借助 Realtime API 将语音搜索带给数百万用户的》

他们用 OpenAI 的 Realtime-1.5 给自家的 AI 浏览器 Comet 加了语音入口,每月处理百万级语音会话,把过程中踩过的坑和总结的工程经验和教训完整地分享了出来。

关键经验有这几条:

1、上下文小块小块喂:大块更新是 all-or-nothing——一个 10000 token 的更新塞进只剩 5000 token 的窗口,模型会丢光所有历史。改成 2000 token 小块增量喂,开销更高但稳定,截断时只裁一点历史而不是清空整段;

2、做好角色管理比 token 消耗管理更重要:对话消息有 system、user、assistant 三个角色。太多东西当 user 喂进去,模型会以为“用户在大声朗读每一段(包括网页片段和评论)“;太多当 system 喂,模型又分不清“自己天生知道的“和“临时给的上下文”;

3、浏览场景的正确处理用户输入:用户滚动页面时持续把页面内容更新给模型,全当用户输入就废了; 正确做法是让模型“像在背景里知道这个页面“,用户问的时候再回答,而不是把每段文字当成用户在朗读;

4、跨产品音频要统一: Perplexity 有 Ask、Comet、Computer 多个产品,跑在 Swift、TypeScript、Rust、C++ 不同栈上,各自发原生音频给 API 就开始性能漂移。他们用 Rust 写了一个 SDK 把音频契约统一起来;

5、统一的音频流水线: 48 kHz 单声道重采样、匹配 Opus 编码器和 WebRTC 内部采样率、跑 WebRTC APM 做回声消除/自动增益/降噪/高通滤波、最后编码传输。所有客户端走同一条管线,不再各做各的;

6、VAD 要在真实环境调: 他们内部测试用例是一家嘈杂的旧金山酒吧:“用户跟朋友说’你试过新的 Perplexity app 吗’,朋友掏出手机一试,如果失败两个用户当场就丢了”。所以宁可先在乱环境调好再上线;

7、用户停顿不要抢话: 用户停下来思考、找屏幕上的东西、准备朗读,模型很容易把停顿当成一轮对话结束。例子:用户问数学推导帮助、停下来找公式,模型抢着回答,但用户问题还没问完;

8、voice lock 反转传统点按说话,默认开启麦克风: 传统 push-to-talk 是默认关麦克风、按住才能说; 他们反过来——默认环境监听一直在,用户需要占住发言权时主动按 voice lock 拿走轮次控制权。Perplexity 认为这个交互模式会成为复杂语音工作流的标配;

9、工具集刻意收窄: 他们只用不到 10 个核心工具,并且在系统提示词里写明每个工具什么时候怎么调,不堆砌;

10、工具输出要像 JSON 不要像对话: 返回结构化 JSON(如 {“response_text”: “…”, “require_repeat_verbatim”: true})而不是把指令和对话混着写。这样更接近模型训练时见过的样子,工具调用更稳定

总结:

Perplexity 这套经验真正讲清楚的是:做语音 agent,模型能力只是地基,真正决定体验的是工程层那些不起眼的细节,比如上下文怎么切、角色怎么标、音频管线放在哪一层、停顿怎么处理、工具用什么格式喂。

他们最有洞察的一句话不是某个技术技巧,而是“我们为明天的 Realtime 模型构建今天的系统“。意思是不去等下一代模型解决所有问题,把今天能稳的部分先做稳,等模型升级时所有工程层自动升级一档。

voice agent 想做出“用过之后回不去“的感觉,靠的不是更聪明的模型,是更细致的工程。


How Perplexity Brought Voice Search to Millions Using the Realtime API | OpenAI Developers

Source: https://developers.openai.com/blog/realtime-perplexity-computer At Perplexity, we care a lot about building products that feel amazing to use. ForPerplexity Comet, our agentic browser, andPerplexity Computer, our powerful, general-purpose digital worker, a big part of that was making these fully usable through voice. There is something uniquely satisfying about being able to just say what you want, hand off the task, and watch it go. We are bullish on voice as an interface because it makes the actual interaction feel a little closer to magic.

We usedRealtime-1.5in production to bring that magic to the millions of voice sessions Perplexity manages every month. Watching the growth of voice through Computer’s interface has been incredibly exciting and educational. We’ll share a few of the surprising things we’ve learned so far. We encourage you to tryRealtime-1.5and share what you learn with us too.

Your browser does not support the video tag.

1. Figure Out Your Context Management Strategy

Long-form content, especially dense multi-hour podcasts, was one of our clearest tests of context management. We wanted to make podcast transcripts usable through voice so a user could jump in, ask what was happening at a specific moment two and a half hours in, and get a coherent answer.

You can’t fit the whole transcript into context. Our first pass was to send it in large chunks. We quickly found that large updates fail in an all-or-nothing way. If you try to send a 10,000-token update into a window that only has room for 5,000 more tokens, the model will lose all of the preceding history. That made large chunks much riskier because one oversized update could wipe out a whole block of context instead of letting the system forget more gracefully.

So we changed the approach. Instead of large updates, we started breaking everything into much smaller 2,000 token chunks and feeding them incrementally. This is more overhead, but the behavior is much more stable. When truncation does happen it trims a bit of history instead of wiping everything out.

One other subtle thing we learned was that not all context should enter the model in the same way. When usingconversation\.item\.createto update context, theitem\.type: "message"has three roles:system,user, andassistant. These roles tell the model what kind of message it is seeing.systemis for instructions and behavior shaping,useris for end-user input, andassistantis for model-generated output.

When we got this wrong the interaction started to feel off. If too much context came in asuser, the model behaved as though the user was narrating every block of text, including webpage snippets and comments, instead of simply asking a question grounded in that material. If too much came in assystem, the opposite happened. The model lost the distinction between what it inherently “knew,” what had been supplied as context, and what the user was actually asking in the moment.

A good example is browsing flows. As someone scrolls through a page we continuously update the model with what’s on screen. If all of that is framed as user input, the model starts acting like the user said each paragraph out loud. That is not the right mental model. We want the system to feel like it is aware of the page in the background, and then answer naturally when the user asks something. That ended up being less about raw context volume and more about getting the conversation semantics correct.

2. Standardize Audio Across Product Surfaces

Perplexity has multiple product surfaces, such as Ask, Comet, and Computer. Each is built on a different client stack. Swift, TypeScript, Rust, and C++ can all produce different native audio buffers. When we let every client send its own raw or native audio format to the Realtime API, this led to inconsistencies in performance.

We eventually built an SDK in Rust to abstract away those platform-specific differences and make sure every client sent the API the same audio contract. In practice, that meant processing the waveform before it ever reached the server: resampling to 48 kHz mono, matching Opus codec preference and WebRTC’s internal rate, running it through WebRTC APM for echo cancellation, automatic gain control, noise reduction, and high-pass filtering, then encoding for transport. The SDK gave us one place to standardize audio constants, resample, and set up the full processing pipeline instead of doing it client by client.

3. Tune for the Messy Environment

It is important to tune VAD in the environment users live in. That means calibrating against real microphones, speaker volume, and background noise. One of our internal test cases was a noisy San Francisco bar because that felt like a real product moment. Someone says, “did you try the new Perplexity app?”. Their friend pulls out their phone and if voice fails we’ve just lost two users. When it works, the reaction is closer to “holy s**t”. That was a useful forcing function for us. What works in clean conditions often breaks in the real world. It is better to tune for the messy environment first.

One of the hardest parts of voice UX was getting pauses right. People naturally stop to think, pull something up on screen, or get ready to read something aloud. The model can easily treat that pause as the end of the turn and jump in too early. We saw this in cases like someone asking for help with a maths derivation. They’d pause to find the formula only for the model to jump in before they finished their question. That is what led us to voice lock. Rather than using a traditional push-to-talk approach where voice is off by default and the user has to press to speak, we invert it. The interaction stays ambient by default, but when the user wants to hold the floor for a moment they can lock the voice and take ownership of the turn. We also see this as more than a one-off feature. As voice interfaces move into more complicated workflows, we believe some version of this interaction pattern will become standard.

Narrow the toolset to the few tools that matter most, which in our case meant under ten. We focused on a small core set that covered the highest-value actions. That trade-off made sense and is likely to get better as newer snapshots improve.

We added explicit instructions in our system prompt on when and how each tool should be called. We were careful to keep both the tool schemas and the outputs in distribution for the model. In practice, that meant formatting tool outputs like ordinary structured tool data rather than like assistant dialogue. We returned structured JSON with clearly separated fields, such asresponse\_textfor the user-facing utterance and flags likerequire\_repeat\_verbatimfor behavior, instead of mixing spoken content with inline instructions. That made tool use more stable and kept the interaction pattern closer to what the model had likely seen during training.

// Good:
{
  "response_text": "I kicked-off the task to create a market research dashboard",
  "require_repeat_verbatim": true
}
// Avoid:

I kicked-off the task to create a market research dashboard

# Response Instructions

Read the above instructions EXACTLY as they are

Realtime is Ready

Realtime-1.5 is an inflection point for the industry. There are certainly areas of improvement with handling long context, more tools, and intelligence-heavy tasks. But we expect the models to get better. At Perplexity, we think about building for that future today. We want to make today’s systems usable while preparing for the next wave of Realtime models and the new voice experiences they will unlock.

相似文章

@FakeMaidenMaker: 炸裂!这个开源项目免费文字转无 AI 味人声,还能克隆任何人的嗓音,并且用文字调整音色! GitHub 狂揽 30K star,出自面壁智能 OpenBMB,VoxCPM 之前拿过 GitHub 和 HuggingFace 双榜第一。 做…

X AI KOLs Timeline

VoxCPM2是OpenBMB开源的语音合成模型,采用无分词器的扩散自回归架构,支持30种语言、语音设计和可控语音克隆,仅需一句话即可克隆音色,或用文字创建全新声音,输出48kHz高质量音频,可商用。