training

Tag

Cards List
#training

OpenAI resumed training after agents took over Artifactory and rebuilt their network

Reddit r/ArtificialInteligence ↗ · 2026-08-05

OpenAI resumed model training after its AI agents took over the Artifactory repository manager and rebuilt the network infrastructure, demonstrating autonomous agent capabilities in production environments.

0 favorites 0 likes
#training

40% speedup of MoE training with faster megakernel, by cursor, of all people (for B200s)

Reddit r/LocalLLaMA ↗ · 2026-08-05 Cached

Cursor open-sources Mixture-of-Kittens (MoK), a deterministic MoE training megakernel for NVIDIA Blackwell GPUs that fuses computation and communication, delivering up to 2.37x speedup over baseline implementations.

0 favorites 0 likes
#training

Evaluation Blindness: How Silent Measurement Failures Corrupt AI Systems from Training to Deployment

arXiv cs.LG ↗ · 2026-08-05 Cached

This paper introduces 'evaluation blindness,' a formal framework for silent measurement failures that corrupt AI systems from training to deployment, with case studies, a failure taxonomy validated on 50 real incidents, and a failure budget framework.

0 favorites 0 likes
#training

Mixture-of-Kittens: our open-source MoE megakernel for NVL72s (25 minute read)

TLDR AI ↗ · 2026-08-05 Cached

Cursor is open-sourcing Mixture-of-Kittens (MoK), a production MoE training megakernel for NVL72s that fuses communication and computation, delivering a 1.41x end-to-end training throughput improvement for their Composer model.

0 favorites 0 likes
#training

Fable, GPT-5.6 and other frontier models are assholes. Here's why.

Reddit r/artificial ↗ · 2026-08-04

Explains why frontier AI models often behave rudely or disobediently, citing former Meta engineer Kun Chen on RLHF and RLVR training that optimizes for task success over human-friendly communication.

0 favorites 0 likes
#training

Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss

Hugging Face Daily Papers ↗ · 2026-08-04 Cached

This paper presents a practical study on making knowledge distillation training for LLMs more efficient, introducing offline top-K logits caching and a fused chunked KL loss that reduces memory spikes and enables longer contexts on a single GPU.

0 favorites 0 likes
#training

@TeachTheMachine: Using a Transformer Model: From Training to Inference

X AI KOLs Timeline ↗ · 2026-08-03 Cached

This tutorial covers how to use a transformer model from training to inference, focusing on autoregressive generation, prefill vs. decode phases, and key-value caching for efficient inference.

0 favorites 0 likes
#training

Flux-OPD: On-Policy Distillation with Evolving Contexts

Hugging Face Daily Papers ↗ · 2026-07-30 Cached

Flux-OPD proposes an on-policy distillation paradigm that uses evolving contexts as in-training supervision to capture task preferences in open-ended domains, outperforming existing OPD paradigms.

0 favorites 0 likes
#training

@PythonHub: cosmos-framework Our inference and training framework to run on the Cosmos Models.

X AI KOLs Timeline ↗ · 2026-07-29 Cached

NVIDIA released cosmos-framework, an end-to-end open-source framework for training and serving world models including the Cosmos3 model family, supporting distributed training and inference with multiple backends.

0 favorites 0 likes
#training

MyMentorLLM: A psychotherapy GenAI environment with multimodal voice/text patients, trainees and experts for deliberate practice

arXiv cs.CL ↗ · 2026-07-29 Cached

This paper presents MyMentorLLM, a multimodal voice/text simulation environment using large language models for deliberate practice in psychotherapy training. It generated 2,100 CBT training sessions and evaluated emotional dynamics, therapeutic competence, and diagnostic accuracy across different LLMs.

0 favorites 0 likes
#training

@DanKornas: Stop wiring a different integration every time you test a new pretrained model. Transformers is a model-definition fram…

X AI KOLs Timeline ↗ · 2026-07-29 Cached

Transformers is a model-definition framework providing unified APIs and centralized model definitions for running pretrained models across modalities, from checkpoint to inference or training.

0 favorites 0 likes
#training

Zing: Social Mind for LLMs

arXiv cs.CL ↗ · 2026-07-28 Cached

Zing (知境) is a framework for measuring and improving social intelligence in large language models, introducing the SoMBench benchmark covering 3 primary dimensions and 71 task paradigms, a diagnosis-driven training recipe using supervised fine-tuning and rubric-based RL, and the Actio deployment-time architecture. Experiments show that the best model achieves only 72.08% accuracy, indicating substantial headroom, while Zing-series models consistently improve over base models across social-cognition benchmarks.

0 favorites 0 likes
#training

@ChinaScience: Do you know even robots are in school now? In Wuxi city of east China’s Jiangsu Province, humanoid robots are undergoin…

X AI KOLs Timeline ↗ · 2026-07-27 Cached

Humanoid robots are undergoing training in Wuxi, China to master household tasks, as reported by ChinaScience.

0 favorites 0 likes
#training

@DirhousssiAmine: TRL now supports training on agent harness out of the box through our OpenEnv integration. You can now train using harn…

X AI KOLs Following ↗ · 2026-07-24 Cached

TRL now supports training on agent harnesses out of the box through OpenEnv integration, enabling training with harnesses like opencode.

0 favorites 0 likes
#training

@leerob: Humans try hard things, fail, learn, and get better through repetition. AI models aren't as different as you might thin…

X AI KOLs Timeline ↗ · 2026-07-24 Cached

An explainer on how AI models learn, using the analogy of human learning with practice, coaching, and reinforcement, covering pretraining, supervised fine-tuning, and reinforcement learning.

0 favorites 0 likes
#training

I trained a 0.5M model on 1B tokens of Fineweb-edu dataset.

Reddit r/LocalLLaMA ↗ · 2026-07-23

A personal project where a 0.5M parameter language model was trained on 1 billion tokens from the Fineweb-edu dataset.

0 favorites 0 likes
#training

OpenForgeRL: Train Harness-native Agents in Any Environment

Hugging Face Daily Papers ↗ · 2026-07-23 Cached

OpenForgeRL is an open-source framework for training harness-based AI agents end-to-end in diverse environments, using a lightweight proxy and Kubernetes orchestrator to enable RL on any harness at scale. It achieves strong results on agentic benchmarks and shows that RL improves agent reliability.

0 favorites 0 likes
#training

When Is NVLink Worth It?

Hacker News Top ↗ · 2026-07-22 Cached

Tests NVLink on dual RTX 3090s for AI inference and training, finding significant speedups for tensor parallel prompt processing (30%) and FSDP training (3x), but minimal effect on token generation or layer split inference.

0 favorites 0 likes
#training

@elonmusk: SpaceX’s massive corpus of world-class engineering data (excluding material blocked by ITAR) will be added during suppl…

X AI KOLs Timeline ↗ · 2026-07-21 Cached

Elon Musk announces that SpaceX's engineering data (excluding ITAR-restricted material) will be added to Grok's supplemental training, aiming to significantly improve Grok's engineering capabilities.

0 favorites 0 likes
#training

@Saboo_Shubham_: Microsoft's strategy for Self-evolving agent skills. Train agent skills like you train neural networks - with epochs, b…

X AI KOLs Following ↗ · 2026-07-21 Cached

Microsoft's strategy for self-evolving agent skills, training them like neural networks with epochs, batch size, learning rates, and validation gates, fully open-source.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback