audio-understanding

标签

Cards List
#audio-understanding

thinkingmachines/Inkling

Hugging Face Models Trending · 2026-07-14 缓存

Inkling is a large open-weights multimodal model (975B total, 41B active parameters) using a sparse MoE architecture, accepting text, image, and audio inputs and generating text outputs, intended for agentic systems, coding assistants, and chatbots.

0 人收藏 0 人点赞
#audio-understanding

面向大型音频语言模型的连续音频思考

arXiv cs.AI · 2026-06-18 缓存

该论文引入了连续音频思考(CoAT)框架,为大型音频语言模型配备了一个连续的潜在工作空间,用于在生成文本响应之前组织声学信息,从而在音频推理、理解和转录任务中提升性能,且不增加额外的解码成本。

0 人收藏 0 人点赞
← 返回首页

提交意见反馈