training-framework

Tag

Cards List
#training-framework

Counterfactual Benchmarking and Training for Factuality Consistency and Order-Robust Grounded Reasoning in LLMs over Heterogeneous Knowledge

arXiv cs.AI · 2026-08-11 Cached

This paper introduces TKFQA, a counterfactual benchmark of 10,130 QA pairs over tables, texts, and knowledge graphs for evaluating LLM factuality consistency and order-robust reasoning, and proposes ORLF, a training framework that improves reasoning-chain accuracy and reduces input-order sensitivity.

0 favorites 0 likes
#training-framework

iFAN: Inference-Aware Learning for Plain Mask Transformers

Hugging Face Daily Papers · 2026-08-07 Cached

iFAN is a training framework that improves mask transformers for segmentation by aligning query ranking with mask quality and distilling intermediate predictions to the final layer, yielding consistent gains across benchmarks.

0 favorites 0 likes
#training-framework

MiniWorld: Democratizing the Training of Video World Models from Scratch

Hugging Face Daily Papers · 2026-08-02 Cached

MiniWorld is a reproducible framework for training video world models from scratch using a block-causal Video Diffusion Transformer with Flow Matching, enabling efficient streaming generation and trainable in days on a single 8-GPU server.

0 favorites 0 likes
#training-framework

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning

Hugging Face Daily Papers · 2026-07-22 Cached

Molt is a PyTorch-native training framework for agentic reinforcement learning designed to be compact and clean for easy modification, while achieving performance comparable to Megatron-based stacks.

0 favorites 0 likes
#training-framework

Training a harness for model-agnostic and task-environment-agnostic capability improvements with PyTorch-like framework [P]

Reddit r/MachineLearning · 2026-07-20

Introduces a framework for training a harness to improve task LLM capabilities in a model-agnostic and task-environment-agnostic way, with results on Terminal Bench and SWE-Bench.

0 favorites 0 likes
#training-framework

@samsja19: https://x.com/samsja19/status/2076846033922035818

X AI KOLs Following · 2026-07-14 Cached

PRIME-RL is a framework for large-scale asynchronous reinforcement learning, designed to be hackable and scale to 1000+ GPUs with support for various models and environments.

0 favorites 0 likes
#training-framework

Search, Fail, Recover: A Training Framework for Correction-Aware Reasoning

arXiv cs.AI · 2026-07-09 Cached

Introduces Pyligent, a training framework that uses task validators to label failures and teaches LLMs to backtrack during reasoning, improving solve rates on hidden graphs, Sudoku, and Blocksworld.

0 favorites 0 likes
#training-framework

DeepSeek open-sources inference optimizations with 60–85% faster generation [pdf]

Hacker News Top · 2026-06-27 Cached

DeepSeek open-sourced DeepSpec, a full-stack codebase for training and evaluating draft models for speculative decoding, enabling 60-85% faster generation. It includes data preparation, training, and evaluation scripts with support for multiple draft model algorithms (DSpark, DFlash, Eagle3).

0 favorites 0 likes
#training-framework

Learning Robust Pair Confidence for Multimodal Emotion-Cause Pair Extraction

arXiv cs.CL · 2026-06-18 Cached

This paper introduces RPCL, a training-only framework for robust pair confidence learning in multimodal emotion-cause pair extraction, which improves discriminative separation of gold pairs from hard negatives and achieves significant gains in Pair F1 and AUPRC on three datasets.

0 favorites 0 likes
#training-framework

Open weights are not enough: we need open training frameworks for research and better algorithms [P]

Reddit r/MachineLearning · 2026-06-15

A call for open training frameworks in AI research, introducing FeynRL, a modular and explicit framework for RL post-training of LLMs, VLMs, and agents, designed to make training processes visible and modifiable.

0 favorites 0 likes
#training-framework

Retrospective Progress-Aware Self-Refinement for LLM Agent Training

arXiv cs.CL · 2026-06-15 Cached

This paper introduces RePro, a framework that trains LLM agents to self-generate progress signals through a forward-then-reflect rollout paradigm, achieving up to 12% absolute success rate gains on WebShop, ALFWorld, and Sokoban benchmarks.

0 favorites 0 likes
#training-framework

Orchard: An Open-Source Agentic Modeling Framework

Hugging Face Daily Papers · 2026-05-14 Cached

Orchard is an open-source framework for scalable agentic modeling that enables training diverse autonomous agents, achieving state-of-the-art results on coding, GUI navigation, and personal assistance tasks.

0 favorites 0 likes
#training-framework

EasyVideoR1: Easier RL for Video Understanding

Hugging Face Daily Papers · 2026-04-18 Cached

EasyVideoR1 is an efficient reinforcement learning framework for training large vision-language models on video understanding tasks, featuring offline preprocessing with tensor caching for 1.47x throughput improvement, a task-aware reward system covering 11 problem types, and evaluation across 22 video benchmarks. It also supports joint image-video training and a mixed offline-online data training paradigm.

0 favorites 0 likes
#training-framework

Agent Lightning: Train ANY AI Agents with Reinforcement Learning

Papers with Code Trending · 2025-08-05 Cached

Agent Lightning introduces a flexible reinforcement learning framework for training large language models in AI agents, achieving decoupling between agent execution and training to handle complex interactions.

0 favorites 0 likes
#training-framework

Gathering human feedback

OpenAI Blog · 2017-08-03 Cached

OpenAI releases RL-Teacher, an open-source tool for training AI systems through human feedback instead of hand-crafted reward functions, with applications to safe AI development and complex reinforcement learning problems.

0 favorites 0 likes
← Back to home

Submit Feedback