Tag
This paper presents SRCF, an attack that steers Large Reasoning Models (LRMs) via counter-aligned few-shot conversations to cause unsafe or refusal behaviors, and proposes ARCF, a post-training defense that enhances safety and helpfulness without degrading utility.
This paper introduces FakeContext-bench to evaluate how well large language models distinguish between contextual information and factual knowledge, and proposes Jurisdiction In-Context Learning (J-ICL) to enhance both in-context learning performance and resistance to misleading context.
Google's Gemini 4 Pro AI model is nearing a preview release with early post-training versions expected to outperform competitors like Astra, with a possible public release in October.
The article argues that robotics needs post-training similar to language models to achieve high reliability, discussing challenges and potential approaches for universal post-training in robotic systems.
This paper introduces a benchmark for evaluating LLM agents as forward-deployed engineers in post-training delivery, highlighting the critical 'trains but does not learn' failure mode where models optimize without actual learning.
TelecomGPT-R1 is a family of open-source models for unified telecom reasoning, trained with supervised fine-tuning and reinforcement learning, outperforming proprietary models like GPT-5 on benchmarks.
This paper analyzes post-training weight updates in LLMs using singular value decomposition, identifying geometric components that drive performance gains, with insights suggesting that reshaping singular values is less critical than rotating and routing changes.
A user questions the barriers to using Tinker for all ML posttraining needs, inquiring about factors like speed, cost, scalability, correctness, and flexibility.
The paper presents ACLArena, a framework for evaluating Agent Continual Learning in multi-stage post-training, analyzing forgetting and generalization mechanisms, and proposing an improved ACL recipe using offline replay and LoRA experts.
The tweet discusses effective methods for AI attribution using Shapley values for inference, LoRA for post-training, and hierarchical approaches for pre-training, proposing a future of routed general intelligence.
This paper proposes a dependency-graph framework to formalize compositional reasoning in language models and evaluates the impact of reinforcement learning post-training, finding an asymmetry where composed-skill training transfers more readily to decomposed tasks than vice versa.
The paper introduces MoDA, an online post-training reinforcement learning algorithm that jointly enhances quality and diversity in large language models to mitigate mode collapse, eliminating the need for hand-crafted personas or architectural modifications.
The paper proposes Turn-level Multiscale Density Ratio Estimation (tlm-DRE), a post-training method that assigns different weights to turns and uses asymmetric token-level training to enhance LLM agents' performance in multi-turn reasoning tasks, showing competitive results on benchmarks.
ReDraft is a reference-driven revision method for continual post-training of large vision-language models that balances learning new tasks and preserving old ones, achieving higher accuracy and less forgetting than standard approaches like SFT.
This paper investigates offline reinforcement learning for post-training code-generating LLMs, showing that it can improve zero-shot code generation performance using existing datasets without online sampling.
This study investigates off-target effects of response-style alignment in a Korean 27B language model, finding that post-training for style significantly impacts answer propensity and disclosure rates without targeting safety or capability.
LOCUS is a task-aware low-rank post-training method that reduces output token length in language models while maintaining preference alignment, achieving up to 39.84% reduction on Pythia-2.8B with minimal parameter updates.
Nebius has launched the AI Builder Program, offering AI builders resources such as runnable examples, blueprints, courses, and over $400 in credits to facilitate building AI systems.
This paper introduces a data-centric pipeline for post-training language models to enhance financial reasoning through mining reasoning traces, distilling instruction data, and generating verifiable QA pairs, demonstrating improvements in performance while preventing catastrophic forgetting.
Direct Diversity Optimization (DDO) is an offline post-training method that improves successful strategy coverage in LLM agents for sequential decision tasks, outperforming other methods in benchmarks like BabyAI, BabaIsAI, and WebShop.