instruction-tuning

Tag

Cards List
#instruction-tuning

I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens

Reddit r/LocalLLaMA ↗ · 2h ago Cached

YOON1v released Apex-2, a from-scratch decoder-only Mixture-of-Experts LLM with 3.87B total / 1.45B active parameters, pretrained on 86.5B tokens and SFT-tuned, with full code, architecture writeup, and benchmark results (HumanEval 43.9) released on Hugging Face and GitHub.

0 favorites 0 likes
#instruction-tuning

On-Policy Self-Distillation for Multi-Turn Image Editing

Hugging Face Daily Papers ↗ · 6d ago Cached

This paper proposes MT-OPSD, a non-policy self-distillation framework to address degradation in multi-turn image editing and introduces LME-Bench for evaluating long-horizon robustness.

0 favorites 0 likes
#instruction-tuning

Rewired or Gated? How Instruction Tuning Shapes Knowledge-Conflict Circuits in LLMs

arXiv cs.LG ↗ · 2026-09-23 Cached

This paper provides a mechanistic comparison of knowledge-conflict circuits in LLMs under instruction tuning, finding that tuning gates rather than rewires these circuits across multiple model families, with implications for interpretability.

0 favorites 0 likes
#instruction-tuning

"As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It

arXiv cs.LG ↗ · 2026-09-23 Cached

The study reveals that chat templates control whether language models adopt a disclaimer voice (e.g., 'I'm just an AI') or an experiential voice (e.g., 'I feel'), and identifies an activation direction that can steer this behavior, impacting AI safety research.

0 favorites 0 likes
#instruction-tuning

ImIR: Image-Instruction Tuning for All-in-One Image Restoration

Hugging Face Daily Papers ↗ · 2026-09-21 Cached

ImIR adapts a pretrained image-editing model for six image restoration tasks using image-derived instructions, enabling efficient and task-agnostic restoration.

0 favorites 0 likes
#instruction-tuning

@AtriaASI: We’re grateful to OpenBMB for supporting Atria Dawn Preview with UltraData-SFT-Agent-2609. The dataset’s high-quality a…

X AI KOLs Timeline ↗ · 2026-09-16 Cached

OpenBMB releases UltraData-SFT-Agent-2609, a dataset of 500,000 samples for agent instruction-tuning, used in the post-training of the MiniCPM5-2B model.

0 favorites 0 likes
#instruction-tuning

SalamandraTA at WMT 2026 Terminology Shared Task: Hard Examples Are Better Teachers

arXiv cs.CL ↗ · 2026-09-10 Cached

The paper presents a data selection method for terminology-aware translation that trains only on hard examples where the model's output contradicts the glossary, achieving improved term accuracy, and describes the BSC system submission to the WMT26 Terminology Shared Task.

0 favorites 0 likes
#instruction-tuning

LLMs Learn Better In-Context from Rules than from Examples

arXiv cs.CL ↗ · 2026-09-04 Cached

This paper compares rule-based and example-based in-context learning in LLMs, finding that rules generally lead to more reliable learning across diverse tasks, with implications for model and task properties.

0 favorites 0 likes
#instruction-tuning

EmoStance: Response-Side Affective-Orientation Control for Empathetic Response Generation via Emoji Weak Supervision

arXiv cs.AI ↗ · 2026-09-03 Cached

EmoStance is a method for empathetic response generation that uses emoji weak supervision to model response-side affective orientation, improving contextual specificity and perceived responsiveness in dialogues.

0 favorites 0 likes
#instruction-tuning

How Output Format Confounds Data Quality and Capability in Instruction Tuning

arXiv cs.CL ↗ · 2026-09-03 Cached

This paper demonstrates that output formats confound data quality metrics and model capability assessments in instruction tuning, causing significant accuracy shifts and rendering current practices ineffective without interface-aware adjustments.

0 favorites 0 likes
#instruction-tuning

Incremental Risk Assessment of Progressive Elder Financial Scams via Instruction-Tuned Small Language Models

arXiv cs.AI ↗ · 2026-09-02 Cached

The paper proposes a cumulative turn-based risk assessment framework using fine-tuned small language models to incrementally detect financial scams in multi-turn conversations targeting older adults.

0 favorites 0 likes
#instruction-tuning

From Specialization to Generalization: Instruction-tuned LLMs for Robust Harmful Content Mitigation

arXiv cs.CL ↗ · 2026-08-27 Cached

This paper investigates instruction-tuning general-purpose LLMs for robust harmful content mitigation, specifically hate speech detection, using a unified corpus of 36 datasets, achieving state-of-the-art performance and enhanced cross-domain and cross-lingual generalization.

0 favorites 0 likes
#instruction-tuning

Domain Agnostic Text Redaction from Natural Language Rules using Instruction Tuning

arXiv cs.CL ↗ · 2026-08-18 Cached

The paper introduces an explainable, domain-agnostic text redaction method using instruction tuning of language models to identify and redact sensitive information based on natural language rules.

0 favorites 0 likes
#instruction-tuning

Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity

Hugging Face Daily Papers ↗ · 2026-08-13 Cached

This paper studies how instruction tuning affects model confidence and lexical diversity in question answering, finding that it alters confidence and reduces rationale diversity without improving calibration.

0 favorites 0 likes
#instruction-tuning

NeuPAT: Neuron-aware Plasticity Allocation Tuning for Language-Preserving MLLMs

arXiv cs.CL ↗ · 2026-08-11 Cached

NeuPAT is a lightweight, architecture-agnostic framework that allocates neuron-wise update constraints during multimodal instruction tuning to preserve language capabilities in MLLMs, recovering 94.5% of language degradation from vanilla tuning.

0 favorites 0 likes
#instruction-tuning

SemiAdapt-Instruct: Extensible Instruction Tuning via Latent Domain-Specialised Adapters

arXiv cs.CL ↗ · 2026-08-07 Cached

SemiAdapt-Instruct proposes a modular framework that discovers latent instruction domains, trains per-domain LoRA adapters in parallel, and routes among them without extra parameters, enabling extensible instruction tuning where new domains require only single-adapter updates instead of full retraining.

0 favorites 0 likes
#instruction-tuning

The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models

arXiv cs.CL ↗ · 2026-08-03 Cached

This paper shows that knowledge distillation has asymmetric effects on bias in small language models: it improves context-following on unambiguous tasks but harms refusal calibration on ambiguous ones, and proposes PCCD, a protocol to diagnose such per-item harms that aggregate metrics miss.

0 favorites 0 likes
#instruction-tuning

HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs

arXiv cs.CL ↗ · 2026-07-31 Cached

HSS-Synth introduces the first data synthesis pipeline for humanities and social sciences, producing 237k high-quality instruction-tuning samples that outperform baselines and set new SOTA on Qwen3-8B.

0 favorites 0 likes
#instruction-tuning

Instruction-Tuned Models Locally Reuse Human Syntax More Than Humans Do

arXiv cs.CL ↗ · 2026-07-29 Cached

This paper investigates syntactic convergence in instruction-tuned large language models, finding that they reuse human syntax more than humans themselves do in dialogue contexts, with instruction-tuning increasing reliance on syntactic patterns.

0 favorites 0 likes
#instruction-tuning

IKS-Instruct: A 24,000-Example Multilingual Dataset for Teaching Language Models Indian Knowledge Systems

arXiv cs.CL ↗ · 2026-07-28 Cached

IKS-Instruct is a multilingual dataset of 24,795 instruction-response pairs for teaching language models Indian Knowledge Systems, spanning seven Indian languages and covering 41 pedagogical techniques. Evaluation shows a fine-tuned 7B model performs competitively with larger general-purpose models on IKS-specific tasks.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback