llm-finetuning

Tag

Cards List
#llm-finetuning

Matryoshka attribution: Learning to attribute language model outputs to representations and weights

arXiv cs.CL ↗ · 5d ago Cached

The paper introduces Matryoshka Attribution (MAttr), a mask-learning method for attributing language model outputs to internal components, achieving top performance on the Mechanistic Interpretability Benchmark and demonstrating practical use in modifying LLM behaviors by adjusting weights.

0 favorites 0 likes
#llm-finetuning

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling

arXiv cs.LG ↗ · 2026-08-05 Cached

Introduces SP3O, a novel reward-model-free, critic-free, gradient-based preference-based RL algorithm that leverages segment-level preferences, demonstrating improved performance in robotic control and LLM fine-tuning, especially for long-horizon tasks.

0 favorites 0 likes
#llm-finetuning

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning

arXiv cs.LG ↗ · 2026-07-31 Cached

This paper introduces KGPS, a Kalman-guided prompt selection method for adaptive RL finetuning of LLMs, which models prompt difficulty as a dynamic state to improve accuracy and rollout efficiency.

0 favorites 0 likes
#llm-finetuning

FoRA: Fisher-orthogonal Rank Adaptation for Parameter-Efficient Fine-Tuning

arXiv cs.CL ↗ · 2026-05-29 Cached

FoRA introduces a parameter-efficient fine-tuning method that selects task-informative layers via Fisher scores and trains LoRA down-projections on the Stiefel manifold, reducing parameters while preserving accuracy.

0 favorites 0 likes
#llm-finetuning

@Suryanshti777: NVIDIA just revealed the hidden tricks they’re using to make LLM fine-tuning dramatically faster. Not new GPUs. Not big…

X AI KOLs Timeline ↗ · 2026-05-07

NVIDIA and Unsloth have published a technical guide detailing three low-level optimizations that can accelerate LLM fine-tuning by up to 25%, including packed-sequence caching, double-buffered checkpointing, and optimized MoE routing. The guide provides deep systems-level explanations and benchmarks aimed at ML engineers and developers.

0 favorites 0 likes
#llm-finetuning

CBRS: Cognitive Blood Request System with Bilingual Dataset and Dual-Layer Filtering for Multi-Platform Social Streams

arXiv cs.CL ↗ · 2026-04-21 Cached

Researchers from Bangladesh University of Engineering and Technology present CBRS, a multi-platform framework that filters and parses blood donation requests from social media using a dual-layer architecture and a novel 11K bilingual dataset in Bengali and English. Their LoRA fine-tuned Llama-3.2-3B model achieves 99% filtering accuracy and 92% zero-shot parsing accuracy, outperforming GPT-4o-mini and other LLMs with 35× reduced token usage.

0 favorites 0 likes
#llm-finetuning

Reinforcement Learning via Value Gradient Flow

Hugging Face Daily Papers ↗ · 2026-04-15 Cached

Value Gradient Flow (VGF) presents a scalable approach to behavior-regularized reinforcement learning by formulating it as an optimal transport problem solved through discrete gradient flow, achieving state-of-the-art results on offline RL and LLM RL benchmarks. The method eliminates explicit policy parameterization while enabling adaptive test-time scaling by controlling transport budget.

0 favorites 0 likes
← Back to home

Submit Feedback