Tag
Tacit-TTS 将 IndexTTS2 的自回归文本到语义解码替换为掩码非自回归生成,并结合免训练声学长度估计与 ReFlow 蒸馏,实现无需参考文本的零样本语音克隆,在超过 5 秒的语音上比 IndexTTS2 快 10 倍以上,同时支持跨语言及非语音(如婴儿咿呀学语、合成乱码)的参考音频克隆。
Laya is a 322M-parameter non-autoregressive decision engine designed to replace LLM-as-a-judge, providing fast and deterministic typed decisions without token generation, making it cost-effective for classification tasks.
The author describes their earlier work on non-autoregressive decision models and criticizes a frontier lab for calling a similar concept a breakthrough, while announcing their own faster, open-source RL Agent model.
Laya is an open-source, fast multilingual decision engine that offers non-autoregressive, calibrated probabilities for structured schemas, claiming to be 6-8 times faster than Jev with full openness.
An individual claims to have developed and open-sourced a non-autoregressive AI architecture similar to Jev a year before its public announcement, expressing frustration over the lack of recognition and support from a frontier lab.
This paper proposes a discrete diffusion framework for efficient one-to-many machine translation, achieving sublinear latency scaling with the number of target languages and supporting zero-shot transfer to unseen source languages.
ACache introduces a caching mechanism for Diffusion Large Language Models that selectively recomputes critical tokens to improve inference efficiency without losing accuracy.
This paper proposes a non-autoregressive CTC-based approach for speech-to-text diacritic restoration in Arabic, incorporating hard constraints during decoding to improve efficiency and reduce error rates.
FreyaTTS is a compact, tokenizer-free Turkish-first text-to-speech model based on a non-autoregressive conditional flow-matching Diffusion Transformer, achieving state-of-the-art performance with a fraction of the parameters of larger systems and released under Apache-2.0.
iLLaDA is an 8B parameter masked diffusion language model with fully bidirectional attention, trained from scratch on 12T tokens. It shows broad improvements over LLaDA and remains competitive with Qwen2.5 7B on several benchmarks. The model and code are open-sourced.
A systematic experimental analysis evaluating eight state-of-the-art Diffusion Language Models across multiple benchmarks, analyzing trade-offs between generation quality and computational efficiency.
Proposes a non-autoregressive scoring method for punctuation restoration in streaming ASR that preserves the input transcript and outperforms prompt-based and fine-tuned baselines under a limited lookahead budget.
Researchers propose a training-free method called Suffix-Anchored Confidence Modulation to improve confidence-based decoding in diffusion language models by addressing issues with EOT tokens and premature decoding.
This paper introduces a diffusion language model that treats text as a continuous process over binary bitstreams, using entropy-gated stochastic sampling to close the performance gap with autoregressive models. It achieves state-of-the-art results on LM1B and OWT benchmarks while reducing memory footprint.
Meta's FAIR team released the code for Flowception, a CVPR 2026 paper presenting a non-autoregressive video generation framework that interleaves frame insertion with continuous denoising to reduce error accumulation and computational cost.
Cola DLM is a hierarchical latent diffusion language model that uses text-to-latent mapping and conditional decoding to achieve efficient, non-autoregressive text generation.
CRoCoDiL proposes a continuous and robust conditioned diffusion approach for language that shifts masked diffusion models into a continuous semantic space, achieving superior generation quality and 10x faster sampling speeds compared to discrete methods like LLaDA.