language-model

Tag

Cards List
#language-model

@BottleCapAI: Same model size. Same GPU. ~4.7x more work done! ThinkingCap: Qwen 3.8 reaches answers with about half the reasoning. O…

X AI KOLs Timeline ↗ · 17h ago Cached

ThinkingCap is a finetuned model based on Qwen 3.6 27B that reduces reasoning tokens by about 50% while maintaining performance, leading to significant efficiency gains in inference.

0 favorites 0 likes
#language-model

Sharp Limits for Honest Uncertainty in Hard-Budget Repeated Evaluation

arXiv cs.AI ↗ · yesterday Cached

This paper derives sharp theoretical limits for honest uncertainty in repeated evaluation under a hard budget, with applications to language model and agent benchmarking, demonstrating practical improvements in interval width and MSE.

0 favorites 0 likes
#language-model

An Exploratory Ablation of a Small MLA--SSM Hybrid Language Model

arXiv cs.CL ↗ · yesterday Cached

This paper presents an exploratory ablation study of TALH, a hybrid language model combining MLA and SSM, showing that SSM integration is more critical for validation performance than MLA in the tested setup, with insights on memory usage and timing on consumer hardware.

0 favorites 0 likes
#language-model

Claude Opus 5.5 tops SimpleBench with its 88.4% score.

Reddit r/singularity ↗ · yesterday

Claude Opus 5.5 achieves a top score of 88.4% on the SimpleBench benchmark, indicating significant performance in AI evaluation.

0 favorites 0 likes
#language-model

MWE-ECL: Recoverable Long-Range Context Does Not Always Override Local Lexical Priors

arXiv cs.CL ↗ · 2d ago Cached

The paper introduces MWE-ECL, a diagnostic tool to test whether AI models can use long-range context to override local semantic priors in multiword expressions, highlighting gaps between recoverability and behavioral influence.

0 favorites 0 likes
#language-model

@_akhaliq: Contrastive Language Models HF: https://huggingface.co/Contrastive-LM

X AI KOLs Timeline ↗ · 2d ago Cached

The tweet shares links to Contrastive Language Models on Hugging Face, highlighting recent updates to models like CLM-v0.1-8B for text ranking and deepswe-clm-heads-8k.

0 favorites 0 likes
#language-model

PUBG Ally: A Conversational Embodied Agent as an AI Teammate

Hugging Face Daily Papers ↗ · 2d ago Cached

The paper introduces PUBG Ally, an embodied AI agent that functions as a voice-enabled teammate in PUBG: BATTLEGROUNDS, integrating language model reasoning with real-time game control for interactive play.

0 favorites 0 likes
#language-model

Domain-Adaptive Pretraining Enhances Water Treatment Semantic Representation for Large-Scale Structured Literature Mining

arXiv cs.CL ↗ · 3d ago Cached

This paper presents WaterBERT, a domain-adapted BERT model for water treatment literature mining, enhancing semantic representation and enabling large-scale structured information extraction and knowledge graph construction.

0 favorites 0 likes
#language-model

Training a Language Model End-to-End in Rust: An Experience Report

arXiv cs.CL ↗ · 3d ago Cached

This paper reports on training a language model end-to-end in Rust, detailing failures in Rust ML frameworks like Candle and Burn, and proposing verification methods, with the conclusion that Rust is currently better suited for model serving than training.

0 favorites 0 likes
#language-model

@interjc: GPT 6 Astra/Sol/Luna

X AI KOLs Following ↗ · 3d ago Cached

The tweet mentions GPT 6 with variants Astra, Sol, and Luna, indicating a potential new AI model release or speculation.

0 favorites 0 likes
#language-model

@interjc: Opus 5.5, pretty good value for money

X AI KOLs Timeline ↗ · 3d ago Cached

Anthropic introduces Claude Opus 5.5, a new AI model that performs at the level of Claude Fable 5.1 for most tasks while reducing costs by 40%.

0 favorites 0 likes
#language-model

Sonnet 5.5 and Haiku 5.5 Coming Soon!

Reddit r/singularity ↗ · 3d ago

Anthropic is set to release Sonnet 5.5 and Haiku 5.5, which are upcoming updates to their AI language models.

0 favorites 0 likes
#language-model

A New Chatbot Wants to Unlock the Secrets in Tattered Ancient Greek Records

Wired ↗ · 3d ago Cached

The Austrian Academy of Science, partnering with Mistral and Sail Reply, releases Apollo, the first advanced large language model for Ancient Greek to help scholars restore tattered papyrus fragments by predicting missing words.

0 favorites 0 likes
#language-model

Attention-Aware Routing: Coupling Routing and Attention in MoEs

arXiv cs.AI ↗ · 5d ago Cached

This paper introduces Attention-Aware Routing (AAR), a method that enhances Mixture-of-Experts language models by incorporating attention weights into the router, improving mathematical reasoning performance and revealing coupled dynamics between routing and attention.

0 favorites 0 likes
#language-model

mini-AGI - dynamically grown (530M params currently and growing) continual learning model trained from scratch on 8GB VRAM laptop from batch-1 stream of data.

Reddit r/LocalLLaMA ↗ · 5d ago Cached

mini-AGI is a continual learning byte-level language model that dynamically grows its architecture, trained from scratch on an 8GB VRAM laptop, demonstrating the possibility of personal AI that learns continuously without catastrophic forgetting.

0 favorites 0 likes
#language-model

Hemmingway-1, a 27B open weights model for creative writing (EQ-Bench 4: 1330, Apache-2.0)

Reddit r/ArtificialInteligence ↗ · 5d ago

Hemmingway-1 is a 27B open-weights model specialized for creative writing, achieving a score of 1330 on EQ-Bench 4 and claiming human-like performance in blind tests against frontier models.

0 favorites 0 likes
#language-model

Experimenting with hypersurface-constrained dynamic weight updating [P]

Reddit r/MachineLearning ↗ · 6d ago

An experimental language model architecture uses hypersurfaces for dynamic weight updating to reduce training parameters, achieving better performance than unrolled baselines while using only 16% of the parameters. The approach is tested on the FineWeb-Edu dataset and includes a GitHub implementation.

0 favorites 0 likes
#language-model

Ternary Bonsai is a headless chicken

Reddit r/LocalLLaMA ↗ · 2026-09-18

The user tested the Ternary Bonsai 2 27B AI model with a creative prompt, but it entered an infinite loop, repeating without progress for hours and causing disappointment.

0 favorites 0 likes
#language-model

QVAC Genesis III: A Large-Scale, High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training

arXiv cs.AI ↗ · 2026-09-18 Cached

QVAC Genesis III is an open-source synthetic STEM corpus designed to enhance language model pre-training efficiency, demonstrating significant benchmark improvements over prior datasets.

0 favorites 0 likes
#language-model

prism-ml/Ternary-Bonsai-2-27B-mlx-2bit

Hugging Face Models Trending ↗ · 2026-09-16 Cached

Prism ML released a ternary weight 27B-class AI model optimized for on-device use on Apple laptops, retaining 98.2% of full-precision intelligence with an 8.60 GB footprint and ~47 tok/s performance.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback