training-pipeline

Tag

Cards List
#training-pipeline

Scaling Inherently Interpretable Language Models

Hugging Face Daily Papers · 2026-08-06 Cached

This paper introduces Steerling-8B, a diffusion language model trained with interpretability as a constraint, showing that interpretability improves with scale and enabling concept steering without retraining.

0 favorites 0 likes
#training-pipeline

My learnings from optimizing training pipeline to go from 36 steps/minute to 47 steps/minute

Reddit r/LocalLLaMA · 2026-07-21

The author shares techniques that improved training pipeline performance from 36 to 47 steps per minute.

0 favorites 0 likes
#training-pipeline

I trained a vision-language model to play Snake, and so can you. [P]

Reddit r/MachineLearning · 2026-07-14

A demonstration of training a vision-language model to play Snake using the FeynRL framework, illustrating the full training pipeline in an accessible manner.

0 favorites 0 likes
#training-pipeline

@seclink: MiMo-V2.5-Pro-UltraSpeed Ultra Fast, Large Model Training Pipeline.

X AI KOLs Following · 2026-07-03 Cached

MiMo-V2.5-Pro-UltraSpeed is an ultra-fast large model training pipeline.

0 favorites 0 likes
#training-pipeline

@Ryrenz: Explosive! Hugging Face fully open-sourced and reproduced the entire training pipeline of DeepSeek-R1. GitHub 26.4k stars, from the official Hugging Face team, the most authoritative one. DeepSeek-R1 amazed everyone, but the real training recipe…

X AI KOLs Timeline · 2026-07-03 Cached

The official Hugging Face team fully open-sourced and reproduced the entire training pipeline of DeepSeek-R1 (Open-R1 project), including data, training, and evaluation. It has received 26.4k stars on GitHub, providing a reproducible textbook for training reasoning models for the industry.

1 favorites 1 likes
#training-pipeline

@0xkozue: https://x.com/0xkozue/status/2072607035624247732

X AI KOLs Timeline · 2026-07-02 Cached

A concise one-page cheat sheet covering PyTorch tensors, models, matrix multiplication, and autograd for deep learning training pipelines.

0 favorites 0 likes
#training-pipeline

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning

arXiv cs.CL · 2026-06-17 Cached

This paper proposes the LLM-as-Environment-Engineer framework, where a policy model analyzes failures to automatically redesign the training environment for reinforcement learning, and introduces MAPF-FrozenLake as a controllable testbed. The framework, using Qwen3-4B, outperforms larger models like GPT and Gemini, showing that policy learning improves the model's ability to diagnose weaknesses.

0 favorites 0 likes
#training-pipeline

Training Video Foundation Models with NVIDIA NeMo

Papers with Code Trending · 2025-03-17 Cached

This paper presents a scalable open-source pipeline using NVIDIA NeMo for training and inference of Video Foundation Models, addressing challenges in generating high-quality videos with accelerated dataset curation and parallelized training.

0 favorites 0 likes
← Back to home

Submit Feedback