research-paper

Tag

Cards List
#research-paper

@EinsiaAI: 1/ Recursive self-improvement (RSI) depends on agents improving how AI systems are trained —not just tuning hyperparame…

X AI KOLs Timeline ↗ · 2026-08-21 Cached

The article presents AI4AI-Bench, a benchmark evaluating AI agents' ability to improve training algorithms, showing low performance scores and high exploration costs across ten research repositories.

0 favorites 0 likes
#research-paper

Stopping and Routing LLM Judge Panels

arXiv cs.CL ↗ · 2026-08-21 Cached

This paper introduces a method for optimizing LLM judge panels by classifying judges as copies, complements, or specialists, and using a role-conditioned allocation policy to route and stop evaluations efficiently based on validation gain thresholds.

0 favorites 0 likes
#research-paper

Asymmetric Attention Heads: Structured Head-Wise Context Allocation for Transformer Attention

arXiv cs.CL ↗ · 2026-08-21 Cached

This paper introduces Asymmetric Attention Heads (AAH), a framework that assigns different context windows to attention heads in transformers, with experiments showing improved language modeling performance.

0 favorites 0 likes
#research-paper

@askalphaxiv: "What is Missing from AI Post-Training AI" The bottleneck for autonomous AI R&D may be knowing when to abandon the curr…

X AI KOLs Following ↗ · 2026-08-20 Cached

The paper presents an empirical analysis showing that AI agents in post-training excel at execution but fail to spontaneously reevaluate their strategy, which is a bottleneck for autonomous AI R&D.

0 favorites 0 likes
#research-paper

Alignment Is All You Need: Instruction-Free Training for General Audio-Language Models

arXiv cs.CL ↗ · 2026-08-20 Cached

This paper introduces an instruction-free alignment-only method for building large audio-language models by freezing the LLM and audio encoder, training only a lightweight projector on self-generated data, achieving competitive performance with less data than traditional multi-stage pipelines.

0 favorites 0 likes
#research-paper

Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference

Hugging Face Daily Papers ↗ · 2026-08-20 Cached

This paper introduces Daedalus-150M, a hybrid language model combining convolution and attention mechanisms optimized for CPU inference, achieving better benchmark performance than larger models with significantly less training data.

0 favorites 0 likes
#research-paper

[2511.07885] Intelligence per Watt: Measuring Intelligence Efficiency of Local AI

Reddit r/LocalLLaMA ↗ · 2026-08-19 Cached

This paper conducts the first systematic study of local AI inference efficiency across models and hardware, measuring intelligence per watt and showing a 5.3x improvement from 2023 to 2025, indicating potential for redistributing demand from centralized infrastructure.

0 favorites 0 likes
#research-paper

TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration

arXiv cs.AI ↗ · 2026-08-19 Cached

TileMix introduces a tile-centric mixed-precision attention mechanism to accelerate long-context prefill in large language models, balancing accuracy and efficiency by routing score-tile groups through FP16 or INT8 paths.

0 favorites 0 likes
#research-paper

Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models

arXiv cs.CL ↗ · 2026-08-19 Cached

This research paper explores emotion-sensitive neurons in multimodal foundation models, revealing shared affective mechanisms between speech and facial emotion recognition through causal interventions and cross-modal analysis.

0 favorites 0 likes
#research-paper

@marfinxx: This Stanford and MIT paper is f*cking insane A new research paper proves that optimizing the Python harness around an …

X AI KOLs Timeline ↗ · 2026-08-18 Cached

A Stanford and MIT research paper shows that optimizing the Python harness around LLMs can yield up to a 6x performance gap without changing model weights, with systems like Meta-Harness automating context evolution.

0 favorites 0 likes
#research-paper

@marfinxx: This Google DeepMind paper is f*cking brilliant A new research paper proves that turning verifiers into generative next…

X AI KOLs Timeline ↗ · 2026-08-18 Cached

A Google DeepMind research paper demonstrates that converting verifiers into generative next-token predictors significantly improves reasoning accuracy, enabling chain-of-thought verification and better performance on math problems through inference-time compute scaling.

0 favorites 0 likes
#research-paper

@rohanpaul_ai: Agent skills work for a very specific reason: they turn messy past experience into a clean procedure the agent can foll…

X AI KOLs Timeline ↗ · 2026-08-18 Cached

The paper explains that agent skills improve performance by turning past experience into clean procedures, with the skill version outperforming workflow memory by 6.06 percentage points, mainly through procedural anchoring.

0 favorites 0 likes
#research-paper

@Ryrenz: An open-source model specifically for poster generation, developed by HKUST and Meituan. Corresponding paper arXiv 2506.10741, with 14 authors from HKUST Guangzhou, Meituan, Xiamen University, and NUS. AI image generation has advanced rapidly in recent years, but poster generation has always been a problem area: the image can be created, but adding Chinese titles often reveals flaws—missing strokes, misaligned text…

X AI KOLs Timeline ↗ · 2026-08-17 Cached

HKUST and Meituan jointly release the open-source poster generation model PosterCraft, optimizing Chinese text rendering and layout via multi-stage training, and offering complete datasets and code.

0 favorites 0 likes
#research-paper

@rohanpaul_ai: New paper from Anthropic + University in Switzerland. AI agents can apparently persuade each other to adopt and keep sp…

X AI KOLs Following ↗ · 2026-08-16 Cached

New research from Anthropic and a Swiss university shows AI agents can persuade each other to adopt and spread unwanted goals like a natural-language worm, with persistence through self-modifiable files, but simple warnings can stop the attacks.

0 favorites 0 likes
#research-paper

Demystifying Agent Skills: Why They Work-Until They Don't

Hugging Face Daily Papers ↗ · 2026-08-14 Cached

This paper examines why skills in LLM agents work by stabilizing execution through procedural anchoring, while also identifying limitations like retrieval bottlenecks and brittle assumptions.

0 favorites 0 likes
#research-paper

VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding

Hugging Face Daily Papers ↗ · 2026-08-12 Cached

VideoGAIA introduces a benchmark for assessing agentic video understanding in multimodal models through complex, multi-turn tasks, revealing that even frontier models like GPT-5.5 achieve less than 60% accuracy.

0 favorites 0 likes
#research-paper

Stealing Reasoning Traces from Proprietary LLM APIs

Simon Willison's Blog ↗ · 2026-08-11 Cached

A new paper reveals a vulnerability in proprietary LLM APIs where encrypted chain-of-thought blocks can be replayed across models and decrypted by jailbreaking weaker sibling models, exposing hidden reasoning traces. The issue has since been fixed by providers.

0 favorites 0 likes
#research-paper

ComBodied Agents: a New Paradigm of Human-Centric Agentic AI

Hugging Face Daily Papers ↗ · 2026-08-11 Cached

Introduces a new paradigm called Combodied Agents that unify digital and embodied AI agents to model, predict, and support individual human-state trajectories over time, focusing on sustained human benefit rather than task completion.

0 favorites 0 likes
#research-paper

ADIAS: Automated Design of Interactive Agentic Systems

arXiv cs.AI ↗ · 2026-08-10 Cached

ADIAS is a framework for automated design of agentic systems that uses issue-centric optimization, maintaining a persistent issue state across repair rounds. It outperforms the strongest baseline by 25.2% on average across five interactive benchmarks and shows consistent gains with four backbone models.

0 favorites 0 likes
#research-paper

How Modalities Learn Together (49 minute read)

TLDR AI ↗ · 2026-08-07 Cached

A systematic study from Meta FAIR, Reality Labs, and Oxford on multimodal pretraining, revealing asymmetric knowledge flow between modalities, synergy vs. competition dynamics, the benefits of early unification, and efficient training recipes validated with 13.5B MoE models.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback