post-training

Tag

Cards List
#post-training

@bojie_li: Just finished watching Teacher He Jiyan from Zhongguancun College lead 7 PhD students to train a 7B model from scratch …

X AI KOLs Timeline ↗ · yesterday Cached

Teacher He Jiyan and 7 PhD students from Zhongguancun College trained a 7B model from scratch in 3 months, achieving state-of-the-art performance in the 7B model category, with insights on efficient training and data quality.

0 favorites 0 likes
#post-training

AlphaDiverse: Post-Training Local Quantitative Research Agents for Diverse Exploration in Alpha Factor Mining

arXiv cs.AI ↗ · yesterday Cached

AlphaDiverse is a framework for automating alpha factor mining in quantitative research using post-trained local agents to ensure diverse exploration and maintain data confidentiality.

0 favorites 0 likes
#post-training

Thinking Leakage: A Causal Audit of NoThink Post-Training in Hybrid Reasoning Models

arXiv cs.LG ↗ · yesterday Cached

This paper uses a causal mediation framework to audit thinking leakage in hybrid reasoning models, revealing that post-training gains in NoThink mode substantially rely on invoking existing Think behavior, with leakage ratios between 42% and 79%.

0 favorites 0 likes
#post-training

Pistis Technical Report

arXiv cs.AI ↗ · yesterday Cached

The Pistis technical report introduces a family of multimodal large language models (27B and 9B parameters) developed through a novel post-training framework combining distillation and reinforcement learning, with specialized variants for deep reasoning and agentic tasks.

0 favorites 0 likes
#post-training

Post-Training Leaves Behavioral Shadows on Unrelated Decisions

arXiv cs.CL ↗ · yesterday Cached

The paper introduces Active Taskless Distillation (ATD), a method that transfers capabilities from a teacher model to a student model using only single-word responses on task-unrelated prompts, probing the behavioral shadows of post-training.

0 favorites 0 likes
#post-training

Alignment of LRMs via Counter-Aligned Few-Shot Conversation Exposure

arXiv cs.AI ↗ · 2d ago Cached

This paper presents SRCF, an attack that steers Large Reasoning Models (LRMs) via counter-aligned few-shot conversations to cause unsafe or refusal behaviors, and proposes ARCF, a post-training defense that enhances safety and helpfulness without degrading utility.

0 favorites 0 likes
#post-training

When Context Misleads: In-context Learning with Jurisdiction in Large Language Models

arXiv cs.CL ↗ · 2d ago Cached

This paper introduces FakeContext-bench to evaluate how well large language models distinguish between contextual information and factual knowledge, and proposes Jurisdiction In-Context Learning (J-ICL) to enhance both in-context learning performance and resistance to misleading context.

0 favorites 0 likes
#post-training

Gemini 4 Pro nears its preview release. (Yes, another preview)

Reddit r/singularity ↗ · 2d ago

Google's Gemini 4 Pro AI model is nearing a preview release with early post-training versions expected to outperform competitors like Astra, with a possible public release in October.

0 favorites 0 likes
#post-training

ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video Generation

Hugging Face Daily Papers ↗ · 2d ago Cached

ViRDM is a method for few-step causal video generation using representation distribution matching, achieving better results than DMD-based baselines with efficient training on single or multiple GPUs.

0 favorites 0 likes
#post-training

Towards Universal Post-Training for Robotics (18 minute read)

TLDR AI ↗ · 2d ago Cached

The article argues that robotics needs post-training similar to language models to achieve high reliability, discussing challenges and potential approaches for universal post-training in robotic systems.

0 favorites 0 likes
#post-training

Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers

arXiv cs.LG ↗ · 3d ago Cached

This paper introduces a benchmark for evaluating LLM agents as forward-deployed engineers in post-training delivery, highlighting the critical 'trains but does not learn' failure mode where models optimize without actual learning.

0 favorites 0 likes
#post-training

TelecomGPT-R1: Unified Post-Training for Reasoning Across Heterogeneous Telecom Tasks

arXiv cs.CL ↗ · 3d ago Cached

TelecomGPT-R1 is a family of open-source models for unified telecom reasoning, trained with supervised fine-tuning and reinforcement learning, outperforming proprietary models like GPT-5 on benchmarks.

0 favorites 0 likes
#post-training

GRRR: The Geometry of Reshaping, Rotation, and Routing in Decoder LLM post-training

arXiv cs.LG ↗ · 4d ago Cached

This paper analyzes post-training weight updates in LLMs using singular value decomposition, identifying geometric components that drive performance gains, with insights suggesting that reshaping singular values is less critical than rotating and routing changes.

0 favorites 0 likes
#post-training

@cHHillee: One hope for Tinker is that it can handle all the ML infra for posttraining for (approximately) everyone. So, curious, …

X AI KOLs Timeline ↗ · 5d ago

A user questions the barriers to using Tinker for all ML posttraining needs, inquiring about factors like speed, cost, scalability, correctness, and flexibility.

0 favorites 0 likes
#post-training

ACLArena: Agent Continue Learning in Multi-stage Post-training

Hugging Face Daily Papers ↗ · 5d ago Cached

The paper presents ACLArena, a framework for evaluating Agent Continual Learning in multi-stage post-training, analyzing forgetting and generalization mechanisms, and proposing an improved ACL recipe using offline replay and LoRA experts.

0 favorites 0 likes
#post-training

@iamtrask: AI attribution works better than you think. Shapley for inference, LoRA for post-training, and hierarchy for pre-traini…

X AI KOLs Timeline ↗ · 2026-09-19 Cached

The tweet discusses effective methods for AI attribution using Shapley values for inference, LoRA for post-training, and hierarchical approaches for pre-training, proposing a future of routed general intelligence.

0 favorites 0 likes
#post-training

Compositional Reasoning in Language Models under Reinforcement Learning Post-Training

arXiv cs.AI ↗ · 2026-09-18 Cached

This paper proposes a dependency-graph framework to formalize compositional reasoning in language models and evaluates the impact of reinforcement learning post-training, finding an asymmetry where composed-skill training transfers more readily to decomposed tasks than vice versa.

0 favorites 0 likes
#post-training

@HEI: Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning Jiayi Yuan, Hangoo Kang, …

X AI KOLs Following ↗ · 2026-09-18 Cached

The paper introduces MoDA, an online post-training reinforcement learning algorithm that jointly enhances quality and diversity in large language models to mitigate mode collapse, eliminating the need for hand-crafted personas or architectural modifications.

0 favorites 0 likes
#post-training

Turn-level Multiscale Density Ratio Estimation for LLM Agents

arXiv cs.AI ↗ · 2026-09-16 Cached

The paper proposes Turn-level Multiscale Density Ratio Estimation (tlm-DRE), a post-training method that assigns different weights to turns and uses asymmetric token-level training to enhance LLM agents' performance in multi-turn reasoning tasks, showing competitive results on benchmarks.

0 favorites 0 likes
#post-training

ReDraft, Don't Just Distill: Reference-Driven Revision for Continual VLLM Post-Training

arXiv cs.AI ↗ · 2026-09-16 Cached

ReDraft is a reference-driven revision method for continual post-training of large vision-language models that balances learning new tasks and preserving old ones, achieving higher accuracy and less forgetting than standard approaches like SFT.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback