post-training

Tag

Cards List
#post-training

Performance, Efficiency and Collapse -- Advantages and Challenges in Offline Post-training of Code LLMs

arXiv cs.LG ↗ · 2026-09-14 Cached

This paper investigates offline reinforcement learning for post-training code-generating LLMs, showing that it can improve zero-shot code generation performance using existing datasets without online sampling.

0 favorites 0 likes
#post-training

Off-Target Effects of Response-Style Alignment in a Korean 27B Language Model

arXiv cs.AI ↗ · 2026-09-12 Cached

This study investigates off-target effects of response-style alignment in a Korean 27B language model, finding that post-training for style significantly impacts answer propensity and disclosure rates without targeting safety or capability.

0 favorites 0 likes
#post-training

LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation

arXiv cs.CL ↗ · 2026-09-11 Cached

LOCUS is a task-aware low-rank post-training method that reduces output token length in language models while maintaining preference alignment, achieving up to 39.84% reduction on Pythia-2.8B with minimal parameter updates.

0 favorites 0 likes
#post-training

@rohanpaul_ai: This is a seriously strong offer for AI builders from Nebius. For learning inference pipelines, orchestration, retrieva…

X AI KOLs Following ↗ · 2026-09-10 Cached

Nebius has launched the AI Builder Program, offering AI builders resources such as runnable examples, blueprints, courses, and over $400 in credits to facilitate building AI systems.

0 favorites 0 likes
#post-training

Data-Centric Post-Training for Financial Reasoning: Mining, Distillation, and Verifiable Learning

arXiv cs.CL ↗ · 2026-09-10 Cached

This paper introduces a data-centric pipeline for post-training language models to enhance financial reasoning through mining reasoning traces, distilling instruction data, and generating verifiable QA pairs, demonstrating improvements in performance while preventing catastrophic forgetting.

0 favorites 0 likes
#post-training

Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training

arXiv cs.CL ↗ · 2026-09-10 Cached

Direct Diversity Optimization (DDO) is an offline post-training method that improves successful strategy coverage in LLM agents for sequential decision tasks, outperforming other methods in benchmarks like BabyAI, BabaIsAI, and WebShop.

0 favorites 0 likes
#post-training

Improving Cross-Lingual Token Representations by Adding a Pinch of SALT

arXiv cs.CL ↗ · 2026-09-10 Cached

The paper proposes SALT, a lightweight post-training method that injects span-level supervision into cross-lingual sentence encoders to improve token representations, achieving top results on multilingual token-level benchmarks and enhancing sentence-level performance.

0 favorites 0 likes
#post-training

ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation

Hugging Face Daily Papers ↗ · 2026-09-08 Cached

ActReview is a rebuttal-guided post-training framework that generates diagnostic claims and revision suggestions for peer reviews by leveraging author responses as supervision, along with a human-curated benchmark for evaluation.

0 favorites 0 likes
#post-training

Miles v0.1: Production-Level Post-Training

Hugging Face Daily Papers ↗ · 2026-09-08 Cached

Miles v0.1 is an open-source, production-ready system for large-scale reinforcement learning and post-training, supporting diverse backends and models like GLM-5.2, with a focus on scalability and accessibility.

0 favorites 0 likes
#post-training

ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF

Hugging Face Models Trending ↗ · 2026-09-07 Cached

This repository provides GGUF quantizations of the Qwen3.8-Flash-Next model using gradient-based methods GSQ and RCO for optimized low-bit representation, enabling efficient deployment in standard tools.

0 favorites 0 likes
#post-training

Revisiting Complete Reasoning Traces for Post-Training

Hugging Face Daily Papers ↗ · 2026-09-07 Cached

This paper finds that large language models can gain reasoning improvements from truncated reasoning trajectories rather than full ones during post-training, reducing redundancy while benefiting methods like supervised fine-tuning and reinforcement learning.

0 favorites 0 likes
#post-training

Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation

Hugging Face Daily Papers ↗ · 2026-09-06 Cached

MovieGrid is a multi-grid post-training paradigm that decomposes long videos into spatially arranged chunks to improve multi-shot coherence and efficiency, achieving state-of-the-art intra-shot and inter-shot consistency in video generation.

0 favorites 0 likes
#post-training

@radixark: The real world is multimodal. For AI to understand and recreate it, models need to learn across modalities. In our late…

X AI KOLs Timeline ↗ · 2026-09-03 Cached

Radixark shares a blog post about how Miles supports multimodal learning for AI models with a shared post-training design for vision-language models and diffusion models.

0 favorites 0 likes
#post-training

@ArizePhoenix: Arize Phoenix now supports https://Z.ai GLM-5.3 uses the same base as GLM-5.2. Every gain came from post-training. http…

X AI KOLs Following ↗ · 2026-09-03 Cached

Arize Phoenix now supports GLM-5.3, which builds on GLM-5.2 with improvements from post-training and training on scaled long-horizon environments using an open-source reinforcement learning framework.

0 favorites 0 likes
#post-training

XMerge: Cross-Axis Selection and Reconstructive Layer Merging for LLM Depth Compression

arXiv cs.LG ↗ · 2026-09-03 Cached

XMerge introduces a post-training method for compressing transformer-based LLMs by merging layers through cross-axis selection and reconstructive reconstruction, outperforming existing methods at aggressive depth reduction while maintaining performance.

0 favorites 0 likes
#post-training

@baseten: Today, we're releasing GLM-5.3 Fast: one of the most intelligent open-weight models ever at an even higher TPS. Designe…

X AI KOLs Timeline ↗ · 2026-09-03 Cached

Z.AI releases GLM-5.3 Fast, an advanced open-weight AI model optimized for agentic coding and cybersecurity, featuring a 744B-A40B MoE architecture with substantial benchmark improvements.

0 favorites 0 likes
#post-training

H3 Max by fal

Product Hunt ↗ · 2026-09-02 Cached

Fal.ai launched H3 Max, a post-trained variant of MiniMax H3 optimized for high-quality video generation. It claims top rankings in prompt adherence, aesthetics, and speed—generating 5-second videos in ~3 seconds with 35× the throughput of the official H3 model.

0 favorites 0 likes
#post-training

Qwen will be the king?

Reddit r/LocalLLaMA ↗ · 2026-09-02

Extended reasoning and post-training are key techniques for enhancing AI model performance, with speculation that future models like Qwen 4 could match or surpass large parameter models on specific tasks.

0 favorites 0 likes
#post-training

Base Models Stopped Being the Bottleneck (15 minute read)

TLDR AI ↗ · 2026-08-31 Cached

The article argues that base AI models are no longer the primary bottleneck, with improvements now driven by post-training enhancements as seen in recent releases like GLM5.3 and Qwen3.6.

0 favorites 0 likes
#post-training

Scaling Automatic Research Agents via World Models

Hugging Face Daily Papers ↗ · 2026-08-29 Cached

This paper introduces World Model RL to scale automatic research agents by replacing environment execution with a learned world model, thereby accelerating post-training by 3-4x and enabling smaller agents to outperform larger ones on benchmarks.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback