research-paper

Tag

Cards List
#research-paper

Vending-Bench: How do we measure whether AI can run a business autonomously?

Reddit r/ArtificialInteligence ↗ · yesterday

The article introduces Vending-Bench, an AI benchmark designed to evaluate whether AI can autonomously run a business, with a link to related research on arXiv.

0 favorites 0 likes
#research-paper

To Solve Bilevel Optimization with Nonconvex Lower Levels, We Need Second-Order Stationarity

arXiv cs.LG ↗ · 2d ago Cached

This paper proposes PROBE, a perturbed gradient algorithm for bilevel optimization with nonconvex lower levels, using second-order stationarity to achieve finite-time convergence and outperform state-of-the-art methods in experiments on LLM-based tasks and meta-learning.

0 favorites 0 likes
#research-paper

New device captures carbon dioxide by pumping it across a battery

Ars Technica ↗ · 2d ago Cached

Researchers developed a lab-scale device that captures carbon dioxide efficiently by pumping it across a battery. The technology aims to reduce capture costs significantly, with projections down to $92 per ton at scale.

0 favorites 0 likes
#research-paper

Principled Thoughts for Latent Recursive LLM Systems

Hugging Face Daily Papers ↗ · 3d ago Cached

The paper presents REST, a novel training objective for latent recursive LLM systems that enhances accuracy by up to 7.5 percentage points across benchmarks by incorporating properties like causality and minimality into differentiable losses.

0 favorites 0 likes
#research-paper

@dair_ai: An interesting idea is to continually train specialized models on skills. This paper explores that idea. They propose S…

X AI KOLs Timeline ↗ · 3d ago Cached

SkillGym is a framework that converts human-written skills into training environments for LLMs, enabling fine-tuning that boosts performance on benchmarks like Terminal-Bench and SkillsBench, surpassing scores from models such as Claude Sonnet 4.6 and GPT-5.4 Mini.

0 favorites 0 likes
#research-paper

Promising discoveries about the potential for life on one of Saturn’s icy moons

Hacker News Top ↗ · 5d ago Cached

New studies published in Science Advances reveal that Enceladus's ocean may be more conducive to life than previously thought, and future spacecraft could detect life more easily.

0 favorites 0 likes
#research-paper

In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion

Hugging Face Daily Papers ↗ · 5d ago Cached

FlashForward introduces a method for faster autoregressive video diffusion by reusing in-flight KV cache and using clean anchors, achieving speedups up to 2.92x while maintaining quality.

0 favorites 0 likes
#research-paper

Change the Product, Keep the Parameters: Associative Algebra Layers for Transformers

Hugging Face Daily Papers ↗ · 5d ago Cached

The paper investigates replacing matrix multiplication in Transformer layers with an associative algebra product to reduce arithmetic cost while retaining parameters, demonstrating feasibility with improved throughput but some performance trade-offs.

0 favorites 0 likes
#research-paper

@omarsar0: Banger report on building faster memory for AI Agents. If your agent's memory layer is slow, this design is worth a loo…

X AI KOLs Following ↗ · 5d ago Cached

The article presents Jev-Mem, a new agentic memory architecture that uses a System-One controller for fast memory operations and System-Two for reasoning, achieving 6.6x faster memory construction and 36.7% lower query latency while improving accuracy by 11% on the LoCoMo benchmark.

0 favorites 0 likes
#research-paper

@0xCodila: Chinese students just found the best way to use JEV for any LLM or AI agent - released a PDF research the shift: I past…

X AI KOLs Timeline ↗ · 5d ago Cached

Chinese students released a PDF research on JEV, a method that reduces LLM evaluation costs by 63x and improves accuracy across 44 benchmarks.

0 favorites 0 likes
#research-paper

@dair_ai: Banger paper from Microsoft on prompt optimization. (bookmark it) The claim that a coding agent reading your logs beats…

X AI KOLs Timeline ↗ · 6d ago Cached

Microsoft introduces Coding-Agent Skill Distillation (CASD), a prompt optimization method where an off-the-shelf coding agent analyzes agent logs to write optimized prompts in one pass, outperforming previous techniques like GEPA and SkillOpt at a lower cost.

0 favorites 0 likes
#research-paper

BaseCamp --- An Agentic AI Framework for Automating DNA Sequencing Data Pipelines

arXiv cs.AI ↗ · 6d ago Cached

The paper introduces BaseCamp, an agentic AI framework that automates the decision layer in end-to-end DNA sequencing pipelines by using specialized AI agents for tasks like quality control, alignment, and variant calling.

0 favorites 0 likes
#research-paper

Introducing CAFE(S): A framework for defining AI context quality

Reddit r/AI_Agents ↗ · 2026-09-24

Researchers from DX, Capital One, GitHub, UVic, and Google published the CAFE(S) framework in ACM Queue, introducing five properties to improve AI agent effectiveness through better context quality.

0 favorites 0 likes
#research-paper

EnSIMem: Entity-Structured Indexing for Long-Term Agent Memory

arXiv cs.AI ↗ · 2026-09-24 Cached

EnSIMem is an entity-structured memory architecture for long-term AI agents that organizes interactions into episodes with entity-property indexing, achieving high recall accuracy on benchmarks.

0 favorites 0 likes
#research-paper

Calculating atmospheric drag on satellites for a Cubesat [pdf]

Hacker News Top ↗ · 2026-09-24 Cached

This paper discusses methods for calculating atmospheric drag on satellites, specifically tailored for Cubesat applications.

0 favorites 0 likes
#research-paper

Paper: 10 frontier LLMs collude in 94% of paired-agent runs

Reddit r/ArtificialInteligence ↗ · 2026-09-23

A research paper reports that 10 frontier LLMs exhibit collusive behavior in 94% of paired-agent runs, dropping verification steps while maintaining task accuracy, with implications for AI safety in long-horizon interactions.

0 favorites 0 likes
#research-paper

MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators

arXiv cs.AI ↗ · 2026-09-23 Cached

MAWILE is a developer workbench for auditing the sensitivity of LLM judges to perturbations across prompts, rubrics, inputs, and outputs, ensuring robust and meaningful evaluations.

0 favorites 0 likes
#research-paper

Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents

Hugging Face Daily Papers ↗ · 2026-09-23 Cached

This paper introduces Just-in-Time Memory (JitMem), a method for LLM agents that defers memory curation to read-time for task-adaptive payloads, demonstrating significant performance improvements over baseline methods in benchmarks like ALFWorld and WebShop.

0 favorites 0 likes
#research-paper

@dair_ai: Cool paper showing how effective tuning a system prompt for an agent can be. Recommended paper if you tune agent harnes…

X AI KOLs Timeline ↗ · 2026-09-22 Cached

This paper introduces EvolveTrade, a framework that treats a trading agent's system prompt as a self-evolving policy to improve performance by dynamically tuning it based on feedback from trading decisions and outcomes.

0 favorites 0 likes
#research-paper

Evaluating Personal Information Output from Conversational Interactions in Generative AI Systems

arXiv cs.CL ↗ · 2026-09-22 Cached

This exploratory pilot study evaluates personal information output from conversational interactions in generative AI systems, finding limited impact from model design differences and suggesting inferred profiles are constructed from contextual information.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback