supervised-learning

Tag

Cards List
#supervised-learning

Improving Cross-Format Robustness in Language Models with Multi-Format Training

arXiv cs.CL ↗ · 2026-06-11 Cached

This paper introduces FormatMix, a multi-format training approach that improves LLM consistency across different answer formats by expanding a subset of training items into multiple equivalent formats, showing that format diversity is key to robustness.

0 favorites 0 likes
#supervised-learning

Rich Sutton on AI creativity and discovery

Hacker News Top ↗ · 2026-06-10 Cached

Rich Sutton argues that generative AI trained by supervised learning cannot achieve genuine novelty and quality simultaneously, and that true discovery requires a 'vary, evaluate, select' mechanism found in reinforcement learning rather than pure imitation.

0 favorites 0 likes
#supervised-learning

Bayes-Sufficient Representations in Supervised Learning

arXiv cs.LG ↗ · 2026-06-04 Cached

This paper formalizes the concept of Bayes-sufficient representations in supervised learning, defining when a representation retains exactly the information needed for Bayes-optimal prediction under a given loss function. It introduces the Bayes quotient as a canonical loss-dependent object and connects the framework to property elicitation, illustrating distinctions between sufficiency, minimality, and excess retained information through experiments.

0 favorites 0 likes
#supervised-learning

Return-to-Go Is More Than a Number: Q-Guided Alignment for Return-Conditioned Supervised Learning

arXiv cs.LG ↗ · 2026-05-29 Cached

This paper proposes Q-align DT, a framework that aligns return-to-go with Q-values to improve controllability and performance in offline reinforcement learning, achieving superior results on D4RL benchmarks.

0 favorites 0 likes
#supervised-learning

Assessing the Operational Viability of Foundation Models for Time Series Forecasting

arXiv cs.LG ↗ · 2026-05-26 Cached

This paper presents an applied evaluation of foundation models for time series forecasting compared to supervised approaches across four operational domains, and proposes a Complexity Router to selectively assign series to the optimal model class for balancing accuracy and inference cost.

0 favorites 0 likes
#supervised-learning

Goal-Conditioned Supervised Learning for LLM Fine-Tuning

arXiv cs.LG ↗ · 2026-05-19 Cached

This paper proposes goal-conditioned supervised learning (GCSL) as an offline fine-tuning framework for LLMs, which treats feedback as an explicit goal and trains models via supervised learning with a novel goal formulation and natural-language goal representations. Evaluated on non-toxic generation, code generation, and recommendation, it outperforms standard offline baselines.

0 favorites 0 likes
#supervised-learning

From Imitation to Interaction: Mastering Game of Schnapsen with Shallow Reinforcement Learning

arXiv cs.AI ↗ · 2026-05-19 Cached

This paper investigates whether shallow neural network agents can master the card game Schnapsen using reinforcement learning, outperforming a supervised imitation baseline and achieving competitive results against a strong search-based opponent.

0 favorites 0 likes
#supervised-learning

@jennyzhangzt: general Intelligence requires rethinking exploration

X AI KOLs Timeline ↗ · 2026-05-16 Cached

This paper argues that exploration is essential for all learning systems, including supervised learning, and proposes a framework for generalized exploration to drive open-ended learning towards general intelligence.

0 favorites 0 likes
#supervised-learning

A Unified Geometric Framework for Weighted Contrastive Learning

arXiv cs.LG ↗ · 2026-05-15 Cached

This paper introduces a unified geometric framework showing that weighted InfoNCE objectives can be interpreted as Distance Geometry Problems, providing exact characterizations of optimal embeddings for supervised and weakly supervised contrastive learning methods and revealing when such embeddings are geometrically realizable, degenerate, or inconsistent.

0 favorites 0 likes
#supervised-learning

UniSD: Towards a Unified Self-Distillation Framework for Large Language Models

Hugging Face Daily Papers ↗ · 2026-05-07 Cached

This paper introduces UniSD, a unified self-distillation framework for adapting large language models that integrates mechanisms for supervision reliability, representation alignment, and training stability. Experimental results show that UniSD improves performance over base models and existing baselines across multiple benchmarks.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback