training

Tag

Cards List
#training

@adithya_s_k: You can now finetune models on agent traces directly with TRL Claude Code traces Codex traces OpenClaw traces Pi traces…

X AI KOLs Following ↗ · 2026-06-04 Cached

TRL now supports fine-tuning models on agent traces from various sources like Claude Code, Codex, OpenClaw, and Pi, moving towards a standardized stack for training agentic models.

0 favorites 0 likes
#training

@loganthorneloe: Read this to get started learning ML infra. This is an excellent high-level overview of important considerations in ML …

X AI KOLs Timeline ↗ · 2026-06-03 Cached

CMU Software Engineering Institute publishes an overview of ML training infrastructure, covering hardware considerations like GPU vs CPU and memory requirements.

0 favorites 0 likes
#training

@FinanceYF5: Anthropic is hiring 1000 freelance software engineers to train Claude Code. Each task pays $280. They write prompts, compare code outputs, test the model's follow-up responses, and teach Claude how real developers work. It's like handing...

X AI KOLs Following ↗ · 2026-06-03 Cached

Anthropic is hiring 1000 freelance software engineers to train Claude Code, with each task paying $280. The engineers will write prompts, compare code outputs, test model responses, and teach Claude how real developers work.

0 favorites 0 likes
#training

@FeitengLi: Asynchronous, Sparse, and the Fifth Decimal Place: Engineering Details of Cursor Training Composer 2 https://lattifai.com/zh/podcasts/SequoiaCapital/UDTr9yUnLUI…

X AI KOLs Timeline ↗ · 2026-06-03 Cached

This article delves into the technical details such as asynchronous and sparse methods used in Cursor training Composer 2 model, and provides a comprehensive analysis of the RL infrastructure.

0 favorites 0 likes
#training

@DanKornas: Stop learning LLMs from disconnected tutorials. LLM from Scratch is a hands-on PyTorch curriculum for builders who want…

X AI KOLs Timeline ↗ · 2026-06-02 Cached

A hands-on PyTorch curriculum that teaches LLM training from transformer basics through fine-tuning and alignment, including RLHF and GRPO.

0 favorites 0 likes
#training

@_djdumpling: very exciting work and thrilled to be working on RL this summer at @modal!

X AI KOLs Timeline ↗ · 2026-06-01 Cached

A user expresses excitement about working on reinforcement learning at Modal, referencing Modal's announcement of an open-source library and lessons learned for scaling RL training.

0 favorites 0 likes
#training

@yibie: Training Small Models: The Most Underrated AI Skill in 2026 On May 11, 2026, a person named CJ Zafir posted a tweet. He wanted to teach ordinary people to fine-tune open source models. 2,538 likes, 316 retweets, 178,000 views. This tweet blew up…

X AI KOLs Timeline ↗ · 2026-06-01 Cached

In May 2026, a tweet by CJ Zafir teaching ordinary people to fine-tune open source models gained widespread attention, illustrating the trend of training small models as the most underrated AI skill in 2026.

0 favorites 0 likes
#training

Gradient-Free Training of Spiking Neural Networks via Low-Rank Evolution Strategies

arXiv cs.AI ↗ · 2026-06-01 Cached

Introduces Eggroll, a low-rank evolution strategy for gradient-free training of spiking neural networks, reducing memory and time overhead while achieving competitive accuracy on N-MNIST.

0 favorites 0 likes
#training

Agentic RL: Token-In, Token-Out Done Right (16 minute read)

TLDR AI ↗ · 2026-06-01 Cached

This article explains the 'Token-In, Token-Out' (TITO) invariant in reinforcement learning for LLMs, highlighting a common error when training multi-turn agents with tool calls. It presents two solutions: using per-model renderers or designing training to avoid re-encoding decoded tokens, emphasizing prefix-preserving chat templates.

0 favorites 0 likes
#training

@charles_irl: why use many bytes when few do trick?

X AI KOLs Following ↗ · 2026-05-30 Cached

Nan Jiang of Modal announces their work on open-source RL frameworks to support frontier open-weights models, highlighting delta compression and remaining challenges in weight sync and cross-cluster training.

0 favorites 0 likes
#training

@ivanfioravanti: One thing's for sure: on Nvidia everything's easier for local AI — inference, training, playing with existing projects.…

X AI KOLs Following ↗ · 2026-05-30 Cached

A developer reflects on the ease of using Nvidia for local AI tasks versus the satisfaction of getting things to work on Apple Silicon, promoting a 'hungry and foolish' mindset.

0 favorites 0 likes
#training

Bias Compounds, Variance Washes Out

Hacker News Top ↗ · 2026-05-29 Cached

This article demonstrates that using stochastic rounding for BF16 optimizer state can match FP32 performance because unbiased errors cancel over time, whereas round-to-nearest stalls due to compounding bias. An experiment with an MLP shows BF16+SR achieves similar loss to FP32 while using less memory.

0 favorites 0 likes
#training

Me train LLM on 8GB from Scratch. Me happy

Reddit r/LocalLLaMA ↗ · 2026-05-29

Built a repository to train a tiny language model (25M parameters) from scratch on 8GB VRAM, with support for MTP but noting limitations of mHC and BitNet.

0 favorites 0 likes
#training

For over a decade, we've accepted that end-to-end backprop is the only way to train deep networks (1 minute read)

TLDR AI ↗ · 2026-05-29 Cached

Sakana AI presents DiffusionBlocks, a method that trains neural networks block-wise by interpreting forward passes as diffusion denoising, significantly reducing memory requirements compared to traditional end-to-end backpropagation.

0 favorites 0 likes
#training

@FrancoisChauba1: If you train on (unsorted list, bubble sort procedure, sorted list) traces, you will never test time compute (TTC) your…

X AI KOLs Following ↗ · 2026-05-26 Cached

A critique arguing that training LLMs on human-generated data limits their ability to discover novel solutions via test-time compute, and that true AGI requires models that can explore hypothesis spaces more broadly, similar to AlphaZero.

0 favorites 0 likes
#training

@ShaokunZhang1: Want to train your own Claude Code/Codex agent with your own model? We are excited to roll out ProRL Agent V2: Polar. A…

X AI KOLs Timeline ↗ · 2026-05-26 Cached

NVIDIA releases Polar, an open-source infrastructure for black-box agentic reinforcement learning, enabling training of coding agents like Claude Code or Codex with any agent harness or framework.

0 favorites 0 likes
#training

Everyone is selling AI agents, but almost nobody is selling the workflows to make them useful.

Reddit r/AI_Agents ↗ · 2026-05-26

The article argues that while many are building and selling AI agents, the real value lies in the workflows and training that make them useful, not the underlying technology.

0 favorites 0 likes
#training

Found in Conversation: LLMs Teach Themselves to Close the Multi-Turn Gap

arXiv cs.CL ↗ · 2026-05-26 Cached

This paper introduces Found in Conversation (FiC), a training framework using View-Asymmetric Self-Distillation to close the multi-turn performance gap in LLMs. The method teaches models to recover single-turn competence from underspecified multi-turn prompts, achieving 92-100% recovery across model families and sizes.

0 favorites 0 likes
#training

A lift for input-convex neural network training

arXiv cs.LG ↗ · 2026-05-26 Cached

Proposes a 'lift' method for training input-convex neural networks (ICNNs) that uses an unconstrained hypernetwork to emit non-negative inter-layer weights, softening the loss landscape and escaping gradient attenuation, achieving lower test loss than projected gradient descent and softplus reparametrization.

0 favorites 0 likes
#training

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models

Hugging Face Daily Papers ↗ · 2026-05-26 Cached

This paper systematically studies scale vectors in LLM normalization layers, showing they optimize training through a self-amplifying preconditioning effect, and proposes three lightweight improvements that enhance performance and scaling behavior with negligible overhead.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback