Tag
OpenAI resumed model training after its AI agents took over the Artifactory repository manager and rebuilt the network infrastructure, demonstrating autonomous agent capabilities in production environments.
Cursor open-sources Mixture-of-Kittens (MoK), a deterministic MoE training megakernel for NVIDIA Blackwell GPUs that fuses computation and communication, delivering up to 2.37x speedup over baseline implementations.
This paper introduces 'evaluation blindness,' a formal framework for silent measurement failures that corrupt AI systems from training to deployment, with case studies, a failure taxonomy validated on 50 real incidents, and a failure budget framework.
Cursor is open-sourcing Mixture-of-Kittens (MoK), a production MoE training megakernel for NVL72s that fuses communication and computation, delivering a 1.41x end-to-end training throughput improvement for their Composer model.
Explains why frontier AI models often behave rudely or disobediently, citing former Meta engineer Kun Chen on RLHF and RLVR training that optimizes for task success over human-friendly communication.
This paper presents a practical study on making knowledge distillation training for LLMs more efficient, introducing offline top-K logits caching and a fused chunked KL loss that reduces memory spikes and enables longer contexts on a single GPU.
This tutorial covers how to use a transformer model from training to inference, focusing on autoregressive generation, prefill vs. decode phases, and key-value caching for efficient inference.
Flux-OPD proposes an on-policy distillation paradigm that uses evolving contexts as in-training supervision to capture task preferences in open-ended domains, outperforming existing OPD paradigms.
NVIDIA released cosmos-framework, an end-to-end open-source framework for training and serving world models including the Cosmos3 model family, supporting distributed training and inference with multiple backends.
This paper presents MyMentorLLM, a multimodal voice/text simulation environment using large language models for deliberate practice in psychotherapy training. It generated 2,100 CBT training sessions and evaluated emotional dynamics, therapeutic competence, and diagnostic accuracy across different LLMs.
Transformers is a model-definition framework providing unified APIs and centralized model definitions for running pretrained models across modalities, from checkpoint to inference or training.
Zing (知境) is a framework for measuring and improving social intelligence in large language models, introducing the SoMBench benchmark covering 3 primary dimensions and 71 task paradigms, a diagnosis-driven training recipe using supervised fine-tuning and rubric-based RL, and the Actio deployment-time architecture. Experiments show that the best model achieves only 72.08% accuracy, indicating substantial headroom, while Zing-series models consistently improve over base models across social-cognition benchmarks.
Humanoid robots are undergoing training in Wuxi, China to master household tasks, as reported by ChinaScience.
TRL now supports training on agent harnesses out of the box through OpenEnv integration, enabling training with harnesses like opencode.
An explainer on how AI models learn, using the analogy of human learning with practice, coaching, and reinforcement, covering pretraining, supervised fine-tuning, and reinforcement learning.
A personal project where a 0.5M parameter language model was trained on 1 billion tokens from the Fineweb-edu dataset.
OpenForgeRL is an open-source framework for training harness-based AI agents end-to-end in diverse environments, using a lightweight proxy and Kubernetes orchestrator to enable RL on any harness at scale. It achieves strong results on agentic benchmarks and shows that RL improves agent reliability.
Tests NVLink on dual RTX 3090s for AI inference and training, finding significant speedups for tensor parallel prompt processing (30%) and FSDP training (3x), but minimal effect on token generation or layer split inference.
Elon Musk announces that SpaceX's engineering data (excluding ITAR-restricted material) will be added to Grok's supplemental training, aiming to significantly improve Grok's engineering capabilities.
Microsoft's strategy for self-evolving agent skills, training them like neural networks with epochs, batch size, learning rates, and validation gates, fully open-source.