Tag
Poolside released Laguna S 2.1, their most capable model for long-horizon tasks, along with multiple quantized variants (FP8, NVFP4, INT4, DFlash, GGUF) on Hugging Face.
Proposes AdaLook, an adaptive multi-step lookahead decoding framework for masked diffusion language models that dynamically determines rollout depth and branch expansion based on candidate-score variance, achieving better accuracy-decoding steps trade-off compared to existing one-step lookahead decoding methods.
Introduces Self-Correcting Coupled Markov Jump Processes (SC-CMJP) and a training-free sampler CO2Jump for concurrent image understanding and generation, achieving state-of-the-art joint performance on editing, maze, and nonogram tasks.
Inkling is a large open-weights multimodal model (975B total, 41B active parameters) using a sparse MoE architecture, accepting text, image, and audio inputs and generating text outputs, intended for agentic systems, coding assistants, and chatbots.
A curated list of over 35 LLM providers offering free, replenishable text-generation API quotas without requiring credit cards, tested and maintained on GitHub.
The paper introduces Telescope Perplexity, a metric that measures token repetition probability to detect LLM-generated text in a zero-shot manner, achieving state-of-the-art or competitive performance across diverse datasets.
Fable 5 demonstrates impressive ability to decode shorthand where users type the first letter of each word, accurately reconstructing the intended message.
A quiz from UnslopAI shows readers can detect AI-generated text 75% of the time, highlighting the need for tools that make AI writing sound more human to avoid search engine penalties and maintain authenticity.
This paper systematically studies the limitations of steering vectors for controlled text generation, finding that their effectiveness varies across traits, degrades on task transfer, and suffers from composition tradeoffs.
Introduces fixed-point flows, a self-conditioned flow language model that treats self-conditioning as a fixed-point iteration, enabling distillation into a few-step flow map language model (FMLM⋆) that outperforms prior work on OpenWebText.
This paper reveals that the low generative perplexity (Gen-PPL) reported by continuous diffusion language models like ELF is misleading, as it rewards repetition; the authors identify a one-dimensional attractor in the self-conditioning loop as the cause and propose ACE, a simple fix that subtracts this direction to reduce repetition without sacrificing quality.
This article presents a technique to improve LLM creative writing by modifying the sampling process using entropy, aiming to reduce the generic 'LLM feel' in generated text.
A tweet compares outputs from GLM-5.2, Fugu Ultra, and Fable 5 using the same one-shot prompt, asking who did it best.
The paper identifies why deterministic few-step generation fails for text while succeeding for images: the sharp categorical readout in text decoders amplifies small errors, causing token flips, whereas continuous image decoders are smooth. It proposes diagnostics (DABI, CCI) and escape mechanisms such as categorical commitment and stochastic re-injection.
This paper proposes Multi-Block Diffusion Language Models (MBD-LMs), extending single-block diffusion to concurrent multi-block decoding with improved training strategies like Multi-block Teacher Forcing and an optimized Block Buffer decoding algorithm. Experiments show increased tokens per forward pass and improved accuracy on benchmarks.
Brain2Qwerty is a non-invasive brain-computer interface that decodes brain waves into text, enabling communication without surgery.
DeepSeek released DSpark, a system where the main model rapidly generates a sentence while a tiny editor fixes coherence before verification, pushing LLM systems engineering beyond new architecture.
A quantized GGUF version of the abliterated GLM-5.2 model is released on Hugging Face, enabling local inference with various tools like Transformers, llama.cpp, and vLLM.
This paper investigates prompt-based learning for automatically generating highlights of academic papers, using models like GPT-2, T5, and ChatGPT, and shows that ChatGPT with few-shot prompts achieves performance comparable to or better than supervised methods without requiring task-specific training data.
Nvidia claims a 15x speedup in text generation using a diffusion model, generating entire blocks at once.