training

Tag

Cards List
#training

I Trained a 117M parameters Silia model on an H100 in 5 hours.

Reddit r/singularity ↗ · 2026-07-07

A 117M parameter Silia model was trained on an H100 GPU in 5 hours using 82M tokens. The model is severely under-trained, but comparisons with nanoGPT are provided.

0 favorites 0 likes
#training

@YuvrajS9886: After Data Parallelism we move to Model Parallelism starting with Pipeline Parallelism The paper: GPipe: Easy Scaling w…

X AI KOLs Timeline ↗ · 2026-07-06 Cached

This thread explains GPipe, a paper on pipeline parallelism for scaling large models across multiple GPUs using micro-batching and activation checkpointing.

0 favorites 0 likes
#training

@QuixiAI: AUM speaks! I trained my custom architecture from absolute zero. 78M parameters trained with 1B tokens on my 8x3090 rig…

X AI KOLs Timeline ↗ · 2026-07-04 Cached

A user announced they trained a custom 78M parameter AI architecture from scratch using 1B tokens on an 8x3090 rig over 12 hours, claiming the model can now speak.

0 favorites 0 likes
#training

@AnnmariaKAntony: LLMs are good at CUDA because the internet is full of it. But a model that gives you highly optimized CUDA may still st…

X AI KOLs Timeline ↗ · 2026-07-02 Cached

A multi-agent synthetic data pipeline with SFT and GRPO RL post-training improves HIP compilation and correctness on AMD MI350X GPUs for a 14B open-source model.

0 favorites 0 likes
#training

Gradient Smoothing: Coupling Layer-wise Updates for Improved Optimization

arXiv cs.LG ↗ · 2026-07-01 Cached

Introduces Depth-wise Gradient Augmentation, a general optimization paradigm that transforms block-wise optimizer updates along depth dimension. The method, Gradient Smoothing, improves optimization and generalization across diverse architectures including transformers and diffusion models.

0 favorites 0 likes
#training

@AaronWeiHuang: Our new blog looks at how FP4 is moving beyond compression into a practical primitive for training and inference across…

X AI KOLs Following ↗ · 2026-06-30 Cached

NVIDIA's blog details how FP4, with the NVFP4 format and Blackwell hardware, has evolved from a compression trick to a practical primitive for training and inference across LLMs and diffusion models, achieving near 16-bit accuracy.

0 favorites 0 likes
#training

Multi-Block Diffusion Language Models

Hugging Face Daily Papers ↗ · 2026-06-30 Cached

This paper proposes Multi-Block Diffusion Language Models (MBD-LMs), extending single-block diffusion to concurrent multi-block decoding with improved training strategies like Multi-block Teacher Forcing and an optimized Block Buffer decoding algorithm. Experiments show increased tokens per forward pass and improved accuracy on benchmarks.

0 favorites 0 likes
#training

@yibie: Recommend this repo to build a GPT-style transformer from scratch without any advanced libraries. With 13M parameters, it can produce grammatically correct text, trainable in one day on a free Colab T4. Train your own LLM from scratch: 13M parameter GPT implementation - Akshay shares…

X AI KOLs Timeline ↗ · 2026-06-29 Cached

Recommended a GitHub repo for building a GPT-style Transformer from scratch without advanced libraries. With 13M parameters, it can be trained in one day on free Colab to generate grammatically correct text.

0 favorites 0 likes
#training

@che_shr_cat: 1/ We have been treating GPU memory all wrong. What if the GPU didn't need to store your model at all? MegaTrain enable…

X AI KOLs Timeline ↗ · 2026-06-29 Cached

MegaTrain enables full-precision training of 100B+ LLMs on a single GPU by treating VRAM as a transient stateless cache, inverting the memory hierarchy.

0 favorites 0 likes
#training

@Hikari_07_jp: Progress report! Training of the DFlash backbone and markov head is complete, enabling DSpark to be used on 27B. We wil…

X AI KOLs Timeline ↗ · 2026-06-28 Cached

Progress update on DSpark: training of DFlash backbone and markov head is complete, enabling use on 27B. Next is training the confidence head for adaptive drafting, expected 8-14% speed improvement over DFlash.

0 favorites 0 likes
#training

@VukRosic99: When a small model learns from a big one, half the lesson is wasted The setup: a small "student" model writes an answer…

X AI KOLs Timeline ↗ · 2026-06-28 Cached

The paper identifies position bias in on-policy distillation for language models, where later tokens in student-generated answers receive degraded supervision. The proposed Importance-Weighted On-Policy Distillation (IW-OPD) weights corrections based on accumulated drift, improving learning speed and final performance.

0 favorites 0 likes
#training

@ArkadiiBessonov: Three main ways to do FP8 in LLM pretraining — and they differ in mainly one thing: how the scale is attached. per-tens…

X AI KOLs Timeline ↗ · 2026-06-27 Cached

Explains three main approaches to FP8 scaling in LLM pretraining—per-tensor, blockwise, and MXFP8—focusing on how the scale is attached, and derives tile geometries from the constraint that scale must remain constant along the matmul's contracted dimension.

0 favorites 0 likes
#training

South Korea plans to train entire military as "drone warriors"

Ars Technica ↗ · 2026-06-26 Cached

South Korea announces plans to train all 500,000 military personnel to operate drones as a 'universal combat tool,' inspired by drone warfare in Ukraine and the Middle East.

0 favorites 0 likes
#training

A debugger for RL reward functions that detects reward hacking during training [P]

Reddit r/MachineLearning ↗ · 2026-06-26

A debugger that detects reward hacking in reinforcement learning reward functions during training, aiding developers in identifying and fixing issues.

0 favorites 0 likes
#training

@almond_robotics: 1 of ~1000 training episodes.

X AI KOLs Following ↗ · 2026-06-26 Cached

Almond Robotics shares one of approximately 1000 training episodes for their robotic system.

0 favorites 0 likes
#training

@lilianweng: A super long overdue (3+ years?) post on scaling laws. Compute is expensive. Scaling laws are a way to help us reason a…

X AI KOLs Timeline ↗ · 2026-06-25 Cached

Lilian Weng's blog post provides a comprehensive overview of scaling laws in deep learning, covering their derivation, compute-optimal allocation, and the debate between Kaplan et al. and Chinchilla.

0 favorites 0 likes
#training

@SergioPaniego: you can now train @liquidai's LFM2-VL in TRL GRPO and RLOO included, with an example script

X AI KOLs Following ↗ · 2026-06-25 Cached

You can now train Liquid AI's LFM2-VL model using TRL's GRPO and RLOO methods, with an example script provided.

0 favorites 0 likes
#training

@natolambert: Another quick lecture -- I've been asked many times for prereq's to my book and what you should know, so built a little…

X AI KOLs Timeline ↗ · 2026-06-24 Cached

Nathan Lambert shares a video lecture covering prerequisites for his book, including language model basics, probabilities, and training pipelines, using GLM 5.2.

0 favorites 0 likes
#training

Reinforcement Learning Towards Broadly and Persistently Beneficial Models

arXiv cs.AI ↗ · 2026-06-24 Cached

This paper from OpenAI investigates whether reinforcement learning on beneficial behavior can produce broad and persistent alignment generalization beyond the training distribution. Using a dataset of realistic situations, they show that RL training on beneficial traits improves out-of-distribution alignment and persistence against adversarial attacks.

0 favorites 0 likes
#training

@SergioPaniego: we let an agent train a coding agent, live, from one prompt which agent is which, why it makes sense, and every artifac…

X AI KOLs Timeline ↗ · 2026-06-23 Cached

A live demonstration of an AI agent training a coding agent from a single prompt, with all artifacts recapped.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback