streaming-inference

Tag

Cards List
#streaming-inference

@UnTalNixon_exe: Forget about GPUs and million-dollar clusters. They just made the world's largest open model (Kimi K3 – 2.78 trillion p…

X AI KOLs Timeline · 2026-08-04 Cached

A new tool called kimi-k3-in-c runs the 2.78T-parameter Kimi K3 open model on a single CPU with as little as 8.24 GB RAM, streaming experts from disk and achieving deterministic output at 10-32 seconds per token.

0 favorites 0 likes
#streaming-inference

@svpino: Here is a new open-weight audio model you can integrate with your app. I'm a huge sucker for open models that you can h…

X AI KOLs Following · 2026-07-28 Cached

Fish Audio S2 is a new open-weight audio model available on Hugging Face, offering two models for timing and acoustic details, with fast inference and a hosted version S2.1 Pro supporting 83 languages at lower cost than ElevenLabs.

0 favorites 0 likes
#streaming-inference

Incremental BPE Tokenization

arXiv cs.CL · 2026-06-01 Cached

This paper introduces an incremental algorithm for Byte Pair Encoding (BPE) tokenization that processes each byte in O(log^2 t) time, enabling efficient partial tokenization in streaming settings and achieving speedups over existing implementations.

0 favorites 0 likes
#streaming-inference

Dolphin-CN-Dialect: Where Chinese Dialects Matter

arXiv cs.CL · 2026-05-12 Cached

Dolphin-CN-Dialect is a streaming-capable ASR model that improves dialect recognition through temperature-based sampling and redesigned tokenization, achieving competitive performance with a smaller model size.

0 favorites 0 likes
#streaming-inference

robbyant/lingbot-map

Hugging Face Models Trending · 2026-04-16 Cached

LingBot-Map is a feed-forward 3D foundation model for streaming 3D reconstruction that uses a Geometric Context Transformer architecture, achieving state-of-the-art performance with efficient ~20 FPS inference on long sequences exceeding 10,000 frames.

0 favorites 0 likes
← Back to home

Submit Feedback