Tag
The paper investigates how periodic subject changes in generated text affect judged surprise and connection in base language models, finding that interruptions raise these metrics but do not compose into integrated documents.
This paper proposes a statistical model to efficiently estimate uncertainty dynamics in text generation, smoothing noisy resampling data to significantly reduce computational costs while maintaining accuracy in analyzing LLM reasoning chains.
An interactive quiz tests users' ability to identify watermarked outputs from large language models, exploring AI text watermarking techniques.
The tweet compares Qwen 3.8 27B and Ornith-1.5-35B models on a prompt for generating a bioluminescent abyssal temple animation, noting that Qwen performs better visually while Ornith is faster in build speed and completion.
Ornith-1.5 releases a family of open-source LLMs from 9B to 397B parameters, achieving state-of-the-art performance among comparable models and offering multiple deployment-friendly formats.
An AI tool that writes marketing copy and strips out AI-written signs to produce human-like text, based on communication research.
A user reports that the Qwen 3.8 27B model achieves 50-60 tokens per second on dual 5060 TI cards, showing unexpected speed improvements over previous versions like Qwen 3.6.
A developer created a skill to reduce AI-sounding phrases in generated text, using insights from a research paper and planning server scraping for more human-like tech speech.
Anthropic's development of text watermarking technology marks a new frontier in AI-generated content detection.
Introduces Simplax, an exact Dirichlet-categorical augmentation for uniform discrete diffusion that improves reverse sampling and generative quality on text and Sudoku tasks.
A community build crushes DeepSeek-V4-Flash down to a 54GB IQ2_XXS GGUF variant with aggressive 2-bit quantization, achieving ~20.5 tokens/s on local hardware while drastically reducing memory footprint.
DiffusionGemma is an experimental open-weight language model that generates text via discrete diffusion rather than token-by-token decoding, enabling exceptionally high-speed generation.
AURORA-LM introduces a continuous-latent diffusion language model that separates decodable text representation from distribution modeling, achieving strong performance on OpenWebText and XSum while scaling to 1B parameters.
The paper introduces Latent-Kernel Discrete Flow Maps (LKF), a flow-map kernel for discrete diffusion models that captures correlations between positions via a shared latent, enabling few-step generation without distillation and improving text generation perplexity.
Built and released BetterGPT-150M, a compact 150M parameter causal language model that outperforms GPT-2 Small with low resource footprint. Includes live Hugging Face Space demo for text completion.
Claude Opus 5 (Max reasoning) surpasses Kimi K3 in Frontend Code Arena and Text Arena, claiming first place in both, while the default high-reasoning version also performs well.
Ant Group releases LLaDA2.X series of diffusion language models, scaling to 100B parameters with MoE architecture, open-sourcing weights and training code.
A user comments that you could hook Codex up to the Pangram API to rewrite text until it passes a human evaluation.
A developer demonstrates running a 28.9 million parameter language model on an $8 ESP32-S3 microcontroller using Google's Per-Layer Embeddings to store most parameters in flash, achieving around 9.5 tokens per second on-device text generation.
This paper proposes Merge-Adversarial Training to make text watermarks in open-source LLMs survive model merging, outperforming baselines while preserving downstream capabilities.