post-training

Tag

Cards List
#post-training

Alignment of LRMs via Counter-Aligned Few-Shot Conversation Exposure

arXiv cs.AI ↗ · 17h ago Cached

This paper presents SRCF, an attack that steers Large Reasoning Models (LRMs) via counter-aligned few-shot conversations to cause unsafe or refusal behaviors, and proposes ARCF, a post-training defense that enhances safety and helpfulness without degrading utility.

0 favorites 0 likes
#post-training

When Context Misleads: In-context Learning with Jurisdiction in Large Language Models

arXiv cs.CL ↗ · 17h ago Cached

This paper introduces FakeContext-bench to evaluate how well large language models distinguish between contextual information and factual knowledge, and proposes Jurisdiction In-Context Learning (J-ICL) to enhance both in-context learning performance and resistance to misleading context.

0 favorites 0 likes
#post-training

Gemini 4 Pro nears its preview release. (Yes, another preview)

Reddit r/singularity ↗ · 17h ago

Google's Gemini 4 Pro AI model is nearing a preview release with early post-training versions expected to outperform competitors like Astra, with a possible public release in October.

0 favorites 0 likes
#post-training

Towards Universal Post-Training for Robotics (18 minute read)

TLDR AI ↗ · 21h ago Cached

The article argues that robotics needs post-training similar to language models to achieve high reliability, discussing challenges and potential approaches for universal post-training in robotic systems.

0 favorites 0 likes
#post-training

Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers

arXiv cs.LG ↗ · yesterday Cached

This paper introduces a benchmark for evaluating LLM agents as forward-deployed engineers in post-training delivery, highlighting the critical 'trains but does not learn' failure mode where models optimize without actual learning.

0 favorites 0 likes
#post-training

TelecomGPT-R1: Unified Post-Training for Reasoning Across Heterogeneous Telecom Tasks

arXiv cs.CL ↗ · yesterday Cached

TelecomGPT-R1 is a family of open-source models for unified telecom reasoning, trained with supervised fine-tuning and reinforcement learning, outperforming proprietary models like GPT-5 on benchmarks.

0 favorites 0 likes
#post-training

GRRR: The Geometry of Reshaping, Rotation, and Routing in Decoder LLM post-training

arXiv cs.LG ↗ · 2d ago Cached

This paper analyzes post-training weight updates in LLMs using singular value decomposition, identifying geometric components that drive performance gains, with insights suggesting that reshaping singular values is less critical than rotating and routing changes.

0 favorites 0 likes
#post-training

@cHHillee: One hope for Tinker is that it can handle all the ML infra for posttraining for (approximately) everyone. So, curious, …

X AI KOLs Timeline ↗ · 3d ago

A user questions the barriers to using Tinker for all ML posttraining needs, inquiring about factors like speed, cost, scalability, correctness, and flexibility.

0 favorites 0 likes
#post-training

ACLArena: Agent Continue Learning in Multi-stage Post-training

Hugging Face Daily Papers ↗ · 3d ago Cached

The paper presents ACLArena, a framework for evaluating Agent Continual Learning in multi-stage post-training, analyzing forgetting and generalization mechanisms, and proposing an improved ACL recipe using offline replay and LoRA experts.

0 favorites 0 likes
#post-training

@iamtrask: AI attribution works better than you think. Shapley for inference, LoRA for post-training, and hierarchy for pre-traini…

X AI KOLs Timeline ↗ · 5d ago Cached

The tweet discusses effective methods for AI attribution using Shapley values for inference, LoRA for post-training, and hierarchical approaches for pre-training, proposing a future of routed general intelligence.

0 favorites 0 likes
#post-training

Compositional Reasoning in Language Models under Reinforcement Learning Post-Training

arXiv cs.AI ↗ · 6d ago Cached

This paper proposes a dependency-graph framework to formalize compositional reasoning in language models and evaluates the impact of reinforcement learning post-training, finding an asymmetry where composed-skill training transfers more readily to decomposed tasks than vice versa.

0 favorites 0 likes
#post-training

@HEI: Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning Jiayi Yuan, Hangoo Kang, …

X AI KOLs Following ↗ · 6d ago Cached

The paper introduces MoDA, an online post-training reinforcement learning algorithm that jointly enhances quality and diversity in large language models to mitigate mode collapse, eliminating the need for hand-crafted personas or architectural modifications.

0 favorites 0 likes
#post-training

Turn-level Multiscale Density Ratio Estimation for LLM Agents

arXiv cs.AI ↗ · 2026-09-16 Cached

The paper proposes Turn-level Multiscale Density Ratio Estimation (tlm-DRE), a post-training method that assigns different weights to turns and uses asymmetric token-level training to enhance LLM agents' performance in multi-turn reasoning tasks, showing competitive results on benchmarks.

0 favorites 0 likes
#post-training

ReDraft, Don't Just Distill: Reference-Driven Revision for Continual VLLM Post-Training

arXiv cs.AI ↗ · 2026-09-16 Cached

ReDraft is a reference-driven revision method for continual post-training of large vision-language models that balances learning new tasks and preserving old ones, achieving higher accuracy and less forgetting than standard approaches like SFT.

0 favorites 0 likes
#post-training

Performance, Efficiency and Collapse -- Advantages and Challenges in Offline Post-training of Code LLMs

arXiv cs.LG ↗ · 2026-09-14 Cached

This paper investigates offline reinforcement learning for post-training code-generating LLMs, showing that it can improve zero-shot code generation performance using existing datasets without online sampling.

0 favorites 0 likes
#post-training

Off-Target Effects of Response-Style Alignment in a Korean 27B Language Model

arXiv cs.AI ↗ · 2026-09-12 Cached

This study investigates off-target effects of response-style alignment in a Korean 27B language model, finding that post-training for style significantly impacts answer propensity and disclosure rates without targeting safety or capability.

0 favorites 0 likes
#post-training

LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation

arXiv cs.CL ↗ · 2026-09-11 Cached

LOCUS is a task-aware low-rank post-training method that reduces output token length in language models while maintaining preference alignment, achieving up to 39.84% reduction on Pythia-2.8B with minimal parameter updates.

0 favorites 0 likes
#post-training

@rohanpaul_ai: This is a seriously strong offer for AI builders from Nebius. For learning inference pipelines, orchestration, retrieva…

X AI KOLs Following ↗ · 2026-09-10 Cached

Nebius has launched the AI Builder Program, offering AI builders resources such as runnable examples, blueprints, courses, and over $400 in credits to facilitate building AI systems.

0 favorites 0 likes
#post-training

Data-Centric Post-Training for Financial Reasoning: Mining, Distillation, and Verifiable Learning

arXiv cs.CL ↗ · 2026-09-10 Cached

This paper introduces a data-centric pipeline for post-training language models to enhance financial reasoning through mining reasoning traces, distilling instruction data, and generating verifiable QA pairs, demonstrating improvements in performance while preventing catastrophic forgetting.

0 favorites 0 likes
#post-training

Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training

arXiv cs.CL ↗ · 2026-09-10 Cached

Direct Diversity Optimization (DDO) is an offline post-training method that improves successful strategy coverage in LLM agents for sequential decision tasks, outperforming other methods in benchmarks like BabyAI, BabaIsAI, and WebShop.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback