training-recipe

Tag

Cards List
#training-recipe

@HarshalsinghCN: yooo guys, the blog is up. i tried to break down the design of BarunLM(35M) in simple language while keeping as much te…

X AI KOLs Timeline · 2026-08-02 Cached

Announcement of a blog post explaining the design of BarunLM, a 35M-parameter language model, covering its architecture, training recipe, and dataset preparation, with a focus on efficiency and performance gains over other sub-100M models.

0 favorites 0 likes
#training-recipe

@tom_doerr: Generates 96,000 high-quality deep research trajectories and provides a fully open-source recipe for training agentic l…

X AI KOLs Timeline · 2026-07-14 Cached

OpenResearcher is an open-source project from TIGER-AI-Lab that provides 96,000 deep research trajectories and a recipe to train agentic language models for long-horizon web research without external APIs. It includes a dataset, model, and demo, and has been adopted by NVIDIA's Nemotron models.

0 favorites 0 likes
#training-recipe

@Ex0byt: A must bookmark.. tiny cracked team, 4 H100 nodes, open source 3 stage recipe, trained on 8k synthetic rubric tasks, fu…

X AI KOLs Timeline · 2026-06-18 Cached

A small team trained a frontier-level Deep Research Agent on an academic budget using only 32 H100s and 8K synthetic samples, releasing fully open weights, code, and paper for models from 2B to 35B that match or beat closed frontier agents on key benchmarks.

0 favorites 0 likes
#training-recipe

Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe

Hugging Face Daily Papers · 2026-06-18 Cached

This paper identifies a fundamental limitation (shrinkage bias) in non-uniform FP4 quantization formats for LLM pretraining and proposes UFP4, a uniform 4-bit training recipe that outperforms existing E2M1-based methods.

0 favorites 0 likes
#training-recipe

AdaMame: A Training Recipe for Adaptive Multilingual Reasoning

arXiv cs.CL · 2026-06-16 Cached

This paper introduces AdaMame, a two-stage training recipe (SFT + GRPO) to adaptively align reasoning language with query language in multilingual mathematical reasoning, mitigating language collapse without sacrificing accuracy.

0 favorites 0 likes
#training-recipe

OdysSim: Building Foundation Models for Human Behavior Simulation

arXiv cs.CL · 2026-06-15 Cached

OdysSim presents a systematic investigation into behavioral foundation models for simulating human behavior, introducing the Soul taxonomy, a corpus of 21.4M interactions, and a training recipe that achieves state-of-the-art on 8 of 23 benchmark tasks while producing more human-like outputs.

0 favorites 0 likes
#training-recipe

@llm_wizard: btw, we publish everything you need to build our Nemotron models including the recipes and pipelines directly. https://…

X AI KOLs Following · 2026-06-09 Cached

NVIDIA released the Nemotron repository with open training recipes, pipelines, and model weights for their Nemotron models, including the new Nemotron 3 Ultra and Nemotron 3 Nano Omni, supporting agentic AI and multimodal capabilities.

0 favorites 0 likes
#training-recipe

Qwen-Image-Flash: Beyond Objective Design

Hugging Face Daily Papers · 2026-06-02 Cached

This paper investigates training recipes for few-step distillation of visual generative models, using Qwen-Image-2.0 as a case study. It reveals non-obvious behaviors and proposes Qwen-Image-Flash.

0 favorites 0 likes
#training-recipe

MiniCPM5-1B Shows Why the Small-Model Race Isn't Over

Reddit r/ArtificialInteligence · 2026-05-31 Cached

MiniCPM5-1B is a 1B parameter model from OpenBMB that achieves impressive scores on AIME 2025 and τ2-Bench Telecom, outperforming larger models. It features both fast and reasoning modes from a single checkpoint, enabled by a three-stage post-training process including supervised fine-tuning, reinforcement learning, and on-policy distillation.

0 favorites 0 likes
← Back to home

Submit Feedback