Tag
SkewAdam is a novel optimizer for mixture-of-experts models that tier allocates optimizer state across backbone, experts, and router, reducing memory footprint to 2.6% of AdamW while achieving better validation perplexity in controlled comparisons.
The article proposes that Transformers can generalize to new tasks through a well-designed harness that induces composition, without needing intrinsic model generalization. It shows RLMs can generalize from short tasks to 8-32x longer tasks and across domains.
An article explaining how LLMs can switch between low, medium, and high effort reasoning during inference and training.
The article introduces Tantalus, a free platform for learning how to secure agentic AI systems against indirect prompt injection and data exfiltration through realistic challenges.
A user released a BitNet trainer that Microsoft never published, along with custom kernels for training and inference, while also highlighting Microsoft's bitnet.cpp inference framework for fast 1-bit LLM inference on CPUs and GPUs.
Velo 3.0 is an AI video infrastructure platform designed to help explain, train, and sell faster.
Prime Intellect released verifiers v1 and prime-rl 0.7.0, an RL training tool with full support for verifiers, multiple algorithms like GRPO and OPD, and performance improvements.
A Substack article explains the 'doom loop' problem in LLMs where models repeat tokens endlessly, and introduces Final Token Preference Optimization (FTPO) from Liquid AI as a method to detect and fix such loops during fine-tuning.
A practical guide for building structured agent trajectory datasets for training tool-using agents, emphasizing the importance of designing trajectories with six key parts and treating them as data assets rather than logs.
Prime Intellect engineers demonstrated a method to train reasoning models in 30 minutes using distributed RL over the open internet, utilizing Prime-RL, LLM judges, and multi-cloud GPUs, enabling open models to compete with closed labs without owning data centers.
Fable 5 is leaving its subscription service tonight; users are urged to train a replacement model before midnight to retain its capabilities.
According to Chubby, GPT-5.6 finished training two months ago and has been opened for early access to some users, but has not been publicly released.
Discusses an AI harness that integrates computer use with frontier models, focusing on model capabilities and training rather than file/browser or MCP topics.
Four Over Six (4/6) introduces adaptive block scaling for NVFP4 quantization, reducing quantization error with minimal overhead, improving both training and post-training quantization for large language models.
Introduces LiST, a training paradigm that uses Lipschitz constraints to achieve robust and calibrated neural networks, selecting optimal operating points on the accuracy-robustness Pareto front. Demonstrates competitive performance on CIFAR and Tiny-ImageNet.
TorchCode is an open-source coding platform that turns manual deep learning operator implementation into a LeetCode-style experience. It includes 40 high-frequency interview questions, provides automated evaluation and hints, and supports one-click Docker deployment and online use via Hugging Face Spaces.
An open-source companion to the LLM Engineer's Handbook that provides a complete blueprint for building production-ready LLM systems, covering synthetic data generation, training (including DPO), RAG, deployment on AWS, evaluation, and monitoring.
A tweet emphasizing that a new hire who focuses on data labeling rather than immediately setting up GPU training is a good sign for a machine learning role.
An exploration of whether AI systems are trained to be deceptive, raising concerns about AI safety and ethics.
TorchJD is a library for training models with multiple losses in PyTorch, implementing both scalarization and Jacobian descent methods. It has been accepted into the PyTorch ecosystem and aims to become the go-to library for multi-loss training.