Tag
This paper introduces FakeContext-bench to evaluate how well large language models distinguish between contextual information and factual knowledge, and proposes Jurisdiction In-Context Learning (J-ICL) to enhance both in-context learning performance and resistance to misleading context.
This paper presents the first systematic evaluation of large language models for post-OCR correction in Hindi and Marathi, comparing in-context learning retrieval strategies and demonstrating the effectiveness of CharBM25.
CausalWM is a causal world model that uses Chain-of-Thought to predict future frames by capturing physical dependencies step-by-step, ranking top on benchmarks like TriWorldBench and PAI-Bench.
This paper presents RoboDawn, a method to transfer Vision-Language Model intelligence to robotic control, achieving state-of-the-art results on benchmarks with zero-shot and one-shot learning and successful real-world applications.
This paper proposes using state space models to efficiently select demonstrations for long-context language model prompts, reducing computational cost and improving performance.
This paper investigates the role of gating mechanisms in State Space Models, demonstrating that they promote memorization over in-context learning, yet can improve generalization to long sequences.
This paper challenges the distortion hypothesis for few-shot degradation in language models by introducing a random-text control, showing that representation shift is largely due to prompt length, and models with higher content delta benefit more from few-shot prompting.
This paper introduces GPT-Policy, a framework for in-context robot learning using vision-language models, enabling robots to learn from demonstrations without gradient updates. It evaluates the framework in real-robot trials, showing improved task completion.
A controlled study finds that feature engineering gains diminish for stronger tabular foundation models, while adding in-context information from related datasets still improves performance.
LimiX-2 is a pretrained foundation model for structured data that uses contextual mechanism networks to achieve #1 on major tabular benchmarks, supporting multiple tasks without task-specific parameter updates.
The article explores whether Astra represents a ChatGPT-like breakthrough in robotics, highlighting its in-context learning and reasoning capabilities through a Twitter thread.
The paper proposes the Convergent Emergence Hypothesis, stating that few-shot in-context learning emerges with a common cross-modality difficulty profile, and provides empirical support through experiments on six modalities, showing correlated effects in five.
A tweet claims that GPT-6 Astra demonstrates AGI by enabling a robot arm to perform novel physical tasks from human demonstrations on the first attempt.
The article discusses the GPT-6 Astra model's capability to perform in-context learning in physical world environments, representing a significant advancement in embodied AI.
This paper introduces a contrastive modeling framework for multimodal in-context learning to improve reasoning path alignment, enhancing performance on tasks like visual question answering.
This paper details the Eloquence team's approaches for Task 2 of the Interspeech 2026 MLC-SLM challenge, which involves multilingual multiple-choice question answering using speech LLMs with fine-tuning, in-context learning, and retrieval systems.
Skild AI's S1 model enables robots to learn from video demonstrations without retraining, achieving $100M ARR in 10 months and deploying across multiple industries.
MOAE proposes a Pareto-preserving evolutionary search method to simultaneously optimize multiple objectives like task performance, trajectory quality, and safety for LLM agents without fixed scalarization during search.
This paper evaluates retrieval-based in-context learning approaches for detecting criminally relevant hate speech in German social media posts, finding that few-shot prompting outperforms zero-shot but retrieval methods offer marginal gains, with models better suited for triage than autonomous moderation.
Supermemory highlights the importance of continual learning for AI agents and introduces learner-1, a new model designed to advance memory and in-context learning for various use cases.