deep-research

Tag

Cards List
#deep-research

@CycleDecoded: WebRover is an open-source AI agent. Give it a sentence, and it opens the browser, identifies webpage elements, clicks, turns pages, scrapes data, gets the job done, and finally organizes the results you want clearly. License: MIT License. Positioning: Autonomous web automation AI agent…

X AI KOLs Timeline · 5d ago Cached

WebRover is an MIT-licensed open-source AI agent that uses natural language to drive the browser for web automation, cross-site data scraping, and deep research, with support for local deployment.

0 favorites 0 likes
#deep-research

Deep Research Pretraining via Predictive Navigation

arXiv cs.CL · 5d ago Cached

Introduces Deep Research Pretraining (DRP), an offline framework that generates search-open-write trajectories from citation and hyperlink evidence structures. Qwen3-14B models pretrained on 1B tokens with DRP outperform matched no-DRP baselines on deep research benchmarks, even with less supervised fine-tuning data.

0 favorites 0 likes
#deep-research

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

Hugging Face Daily Papers · 5d ago Cached

Video-DeepResearch (Video-DR) extends multimodal agents from static images to continuous video streams, introducing a decoupled perception-exploration pipeline and a new benchmark Video-DR-Bench. Their Video-DeepResearch-35B-A3B model achieves 64.0% accuracy, surpassing Claude-4.5-Sonnet, GPT-5, and Gemini 2.5 Pro.

0 favorites 0 likes
#deep-research

Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents

Hugging Face Daily Papers · 6d ago Cached

This paper introduces SIEVE, a search-inspect-fetch strategy that uses Boolean Query Language to make deep-research agents retrieve only relevant document sections, achieving higher accuracy with 20.7–50.6% fewer tokens across multiple benchmark datasets and agent backbones.

0 favorites 0 likes
#deep-research

FinanceHarness: Autonomous Financial Deep Research Framework

arXiv cs.CL · 2026-07-31 Cached

This paper introduces FinanceHarness, a framework for end-to-end automated financial deep research powered by LLM agents, along with FinanceGym, a verifiable point-in-time benchmark. Expert validation shows an 82% pass rate, while leading models score below 40%, and FinanceHarness improves open-weight backbone performance from 25.3% to 32.4%.

0 favorites 0 likes
#deep-research

@LangChain: How @Similarweb evaluates a Deep Research agent when there's no single right answer: Deterministic checks for tool call…

X AI KOLs Timeline · 2026-07-29 Cached

This article explains how Similarweb evaluates long-form agent research reports using LangSmith, combining deterministic checks for tool calls and LLM-as-judge scoring for quality, with a focus on making regressions inspectable and enabling A/B comparisons.

0 favorites 0 likes
#deep-research

@elonmusk: Grok Build /deep-research

X AI KOLs Following · 2026-07-26 Cached

Elon Musk announces a /deep-research command for Grok Build that performs research with bounded parallel agents, cross-checks evidence, and generates cited reports.

0 favorites 0 likes
#deep-research

Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions

Hugging Face Daily Papers · 2026-07-23 Cached

This paper introduces MisKnow-Agent, a framework for generating misleading knowledge to test DeepResearch agents, showing that limited exposure to credible-looking false information can lead to false conclusions in final reports.

0 favorites 0 likes
#deep-research

AREX: Towards a Recursively Self-Improving Agent for Deep Research

Hugging Face Daily Papers · 2026-07-23 Cached

AREX introduces a family of recursively self-improving agents for deep research, alternating between an inner research loop and an outer self-improvement loop, trained with long-horizon reinforcement learning. It substantially outperforms comparable-scale baselines on benchmarks like BrowseComp and Humanity's Last Exam.

0 favorites 0 likes
#deep-research

@omarsar0: Octen's web search latency is pretty insane. Deep research tools like OpenAI, Gemini, Grok, and Perplexity can take up …

X AI KOLs Following · 2026-07-21 Cached

Octen is a web search tool that delivers full source-backed reports in under 3 minutes, outperforming OpenAI, Gemini, Grok, and Perplexity by 10-17 points on the DeepResearch Bench, combining speed and accuracy.

0 favorites 0 likes
#deep-research

Lattics

Product Hunt · 2026-07-21

Lattics is a brain-like knowledge base that integrates AI writing and deep research capabilities.

0 favorites 0 likes
#deep-research

I burned all my tokens researching how to save tokens

Hacker News Top · 2026-07-19 Cached

The author describes building a custom AI research pipeline using multiple subscriptions and cheaper models to reduce token costs, learning firsthand how to optimize token usage while researching tokenomics.

0 favorites 0 likes
#deep-research

@trendtech33566: 【Breaking News】An OSS that fully reproduces the Deep Research workflow in an open manner has appeared: OpenResearcher, …

X AI KOLs Timeline · 2026-07-15 Cached

OpenResearcher is an open-source project that fully reproduces the Deep Research workflow, providing datasets, models, and demos. It enables researchers and AI agent developers to build, train, and evaluate deep research pipelines in an open manner.

0 favorites 0 likes
#deep-research

@tom_doerr: Generates 96,000 high-quality deep research trajectories and provides a fully open-source recipe for training agentic l…

X AI KOLs Timeline · 2026-07-14 Cached

OpenResearcher is an open-source project from TIGER-AI-Lab that provides 96,000 deep research trajectories and a recipe to train agentic language models for long-horizon web research without external APIs. It includes a dataset, model, and demo, and has been adopted by NVIDIA's Nemotron models.

0 favorites 0 likes
#deep-research

Why has progress on Deep Research products stalled?

Reddit r/singularity · 2026-07-12

An analysis questioning why progress on Deep Research AI products has stalled since their impressive launch in February 2025, noting that known weaknesses like hallucinations and unreliable source verification persist despite incremental improvements.

0 favorites 0 likes
#deep-research

Do You Need a Frontier Model as a Citation Verifier? Benchmarking Rubric LLMs for Deep-Research Source Attribution

arXiv cs.CL · 2026-07-10 Cached

This paper benchmarks 8 LLM judges for citation quality in deep-research systems, finding that cheaper models remain competitive with frontier models on source relevance and factual support, but differ in directional bias which matters for RL training.

0 favorites 0 likes
#deep-research

@karminski3: Finally, a useful local DeepResearch is available! I've packaged a skill for everyone, ready to use out of the box. The driving model is Apodex's newly released open-weight DeepResearch fine-tuned model Apodex-1.0-mini (now available a…

X AI KOLs Timeline · 2026-07-06 Cached

Apodex has released an open-weight DeepResearch fine-tuned model, Apodex-1.0-mini, based on Qwen3.5-35B-A3B. It scores 71.5 on BrowseComp, approaching flagship model performance, and can run efficiently locally. The author packaged an out-of-the-box skill.

0 favorites 0 likes
#deep-research

@_akhaliq: LiteResearcher A Scalable Agentic RL Training Framework for Deep Research Agent

X AI KOLs Following · 2026-07-01 Cached

LiteResearcher is a scalable reinforcement learning training framework designed for deep research agents.

0 favorites 0 likes
#deep-research

Reliability is becoming the actual axis the serious AI releases compete on, not how smart they sound

Reddit r/artificial · 2026-07-01

The article argues that the next major competitive axis for serious AI systems is reliability and trustworthiness, not just capability or fluency. It highlights emerging verification techniques—such as independent checks and rubric-based grading—that aim to catch confident but false outputs, a failure mode termed 'pseudo-correctness.'

0 favorites 0 likes
#deep-research

@svpino: New paradigm for deep research models! Apodex-1.0-H is a new model that introduces a completely new way of working. Apo…

X AI KOLs Timeline · 2026-06-26 Cached

Apodex-1.0-H is a new deep research model that introduces a multi-agent architecture where the model decomposes tasks, spawns specialist sub-agents, and uses self-verification and iterative improvement to produce answers. Open-weight variants are available on HuggingFace.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback