A curated collection of papers on agentic search — retrieval systems, benchmarks, and agents built around autonomous, multi-step search and research.
-
How We Built Photon, Sep 2026, Perplexity Photon is Perplexity's new Rust-based retrieval and ranking service, built by a small engineering team working with hundreds of coding agents. It indexes over 200B URLs using compact index reads and asynchronous batched disk I/O, and ranks results in multiple stages that combine lexical and semantic retrieval with embedding scorers and cross-encoder rerankers. Internal p99 latency fell from ~800ms to ~65ms on about 20% fewer serving machines while storing 2.5x more data per document, and it powers the new Fast Search API, which returns 95% of results within 230ms.
-
ITER: Interaction-Aware Retrieval for Agentic Search, Aug 2026, arxiv · code A dense retriever for agent-based search that conditions on the main question, the agent's pre-search reasoning, and preceding sub-queries, rather than just the current query and results. Trained with trajectory-based signals where previously seen documents act as negatives, it gives an average relative improvement of 6.9% on InfoSeek-Eval and 15.4% on BrowseComp-Plus across multiple agent models.
-
3x Faster Search: Parallel Test-Time Scaling with Instructed-Retriever-1, Jun 2026, Databricks Introduces Instructed-Retriever-1, a single retrieval-specialized model that runs query generation (broadening search scope) and multi-pivot reranking (improving precision) in parallel instead of sequential agent reasoning, cutting search time by over 3x and halving answer generation time (~2s time-to-first-token) in Databricks' Knowledge Assistant. Trained on synthetic enterprise-style environments and served with FP8 quantization, speculative decoding, and a Mixture-of-Experts architecture, it matches Claude Sonnet 4.5 quality at much lower latency.
-
Beyond Semantic Similarity: Rethinking Retrieval for Agentic Search via Direct Corpus Interaction, May 2026, arxiv · code Challenges the single-shot embedding-and-vector-index retrieval pipeline, proposing direct corpus interaction (DCI) where agents search raw corpora with general-purpose tools like grep and file operations instead of a fixed semantic retriever. Effective across BRIGHT and BEIR benchmarks, showing retrieval quality depends not just on an agent's reasoning ability but on the resolution of the interface it uses to access the corpus.
-
Learning to Retrieve from Agent Trajectories, Mar 2026, arxiv · code Argues that retrieval models for agentic search should be trained directly on agent interaction data rather than human-centric signals. Introduces LRAT, which extracts training signals from agents' browsing actions and reasoning traces, improving evidence recall, task success, and efficiency across different agent architectures.
-
Instructed Retriever: Unlocking System-Level Reasoning in Search Agents, Jan 2026, Databricks Proposes an architecture that propagates system specifications — instructions, examples, and index schema — through every stage of the search pipeline, letting agents follow complex instructions and reason across heterogeneous sources via query decomposition, relevance assessment, and metadata-to-filter translation. On the new StaRK-Instruct benchmark it achieves 35–50% higher recall than basic retrieval, and deployed in Databricks' Agent Bricks Knowledge Assistant it delivers 70%+ gains over a simple RAG baseline, with fine-tuned smaller models matching larger proprietary models.
-
Introducing Contextual Retrieval, Sep 2024, Anthropic · code Fixes the loss of context when documents are split into chunks by having an LLM (Claude 3 Haiku) prepend a short, chunk-specific explanation (50-100 tokens) to each chunk before building both the embedding index (Contextual Embeddings) and the BM25 index (Contextual BM25). Together the two cut the top-20-chunk retrieval failure rate by 49% (5.7% → 2.9%), and adding a reranker brings the reduction to 67% (→ 1.9%). With prompt caching, the one-time cost is about $1.02 per million document tokens.
-
Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses, Jun 2026, arxiv · code Separates state management from policy decisions in search agents: a stateful harness handles bookkeeping (a candidate pool, importance-tagged curated set, evidence links, verification records, deduplicated observations, budget-aware context rendering) while a 20B RL-trained policy focuses on high-level decisions — what to search for, which documents matter, what to verify, when to stop. Across eight retrieval benchmarks it reaches 0.730 average curated recall, beating the next-best open search subagent by +11.4 points, with especially strong generalization on held-out transfer benchmarks.
-
How Search Quality Shapes RL Outcomes, May 2026, Exa Compares RL-trained search agents using Exa's search engine versus a Google SERP baseline with all else held constant, finding Exa-trained agents reach higher pass@k across benchmarks (often beating larger untrained 235B models) while needing 20% fewer tokens and 62% fewer search calls, because Exa surfaces correct answers 10.7% more often per call. The efficiency gains hold even when agents are evaluated with a different search backend at inference time, across MuSiQue, HotpotQA, and out-of-distribution benchmarks like SimpleQA, FRAMES, and 2WikiMultihopQA.
-
Chroma Context-1: Training a Self-Editing Search Agent, Mar 2026, Chroma · code A 20B-parameter model trained as a specialized search subagent for multi-hop retrieval, ranking relevant documents from large corpora for a downstream reasoning model rather than answering questions itself. Its key idea is "self-editing context" — discarding irrelevant retrieved documents mid-search to stay within bounded context windows — trained via RL with synthetic tasks across web, finance, legal, and email domains, matching much larger frontier models while running up to 10x faster; model weights and the data generation pipeline are released publicly.
-
Recursive Language Models, Dec 2025, arxiv · code Proposes an inference-time paradigm where long prompts are treated as an external environment that the LLM programmatically examines, decomposes, and recursively calls itself over, rather than feeding them directly into the context window. RLMs handle inputs up to two orders of magnitude beyond the model's context window and, on GPT-5, beat compaction by a median 26%, CodeAct with sub-calls by 130%, and Claude Code by 13% across four long-context tasks at comparable cost. A post-trained RLM-Qwen3-8B improves on its base model by 28.3% on average, approaching vanilla GPT-5 on three tasks.
-
s3: You Don't Need That Much Data to Train a Search Agent via RL, May 2025, arxiv · code A model-agnostic RAG framework that decouples the searcher from the generator, training only the searcher via RL with a reward measuring improvement over baseline RAG performance rather than fine-tuning the whole LLM or optimizing retrieval metrics directly. Achieves superior results across multiple benchmarks using just 2,400 training samples — about 70x fewer than competing approaches.
-
Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning, Mar 2025, arxiv · code Trains an LLM end-to-end with RL to interleave step-by-step reasoning with autonomously generated search queries against real-time retrieval, using retrieved-token masking for stable training and a simple outcome-based reward rather than process supervision. Supports multiple RL algorithms (PPO, GRPO, REINFORCE), backbone LLMs, and search engines, improving over comparable RAG baselines by 41% on Qwen2.5-7B and 20% on Qwen2.5-3B across seven QA datasets.
- Tongyi DeepResearch Technical Report, Oct 2025, arxiv · code Describes an agentic research model (~30.5B parameters, 3.3B active per inference step) trained via a fully-automated, human-label-free data pipeline with dedicated environments for each training stage. Achieves leading results on deep-research benchmarks including Humanity's Last Exam, BrowseComp, BrowseComp-ZH, and WebWalkerQA, with model and training framework released publicly.
-
Q2D-Web: A Large-Scale Benchmark for Retrieval in Agentic RAG Systems, Sep 2026, arxiv Introduces a 190M-document web corpus with 70k agentic search queries across ten languages, reformulated from real user queries to evaluate machine-written (rather than human-written) query reformulations. Benchmarking 13 retrievers shows model rankings stay consistent across judgment sets but diverge by domain, language, and query type, and the authors show subcorpus sampling via reciprocal rank fusion can approximate full-corpus evaluation while preserving ranking accuracy.
-
NEEDLE: The Benchmark Your Search Engine Can't Memorize, Aug 2026, Keenable · code A live, open-source benchmark comparing search engines on agent-style query patterns rather than human search behavior, covering News, Finance, Scholar, AgenticRare (obscure entities), and Legal categories. Query sets are continuously refreshed from real agent search logs and live sources like RSS feeds and Google Trends (hourly for News, daily for the rest) to prevent overfitting and memorization, with runs executed in public GitHub Actions and results published to a Hugging Face dataset.
Add new papers under the relevant topic section (create a new ## section if none fits) using the format:
- **Title**, Month Year, [arxiv](link) · [code](repo-link)
A 2-3 sentence summary.
If the paper has a public code/implementation repo, include a [code](link) link after the source link (omit it otherwise).
Within each section, entries are sorted by date descending (newest first).
Update the Table of Contents if you add a new section.