Author-Annotated Benchmark for Scientific Paper Retrieval

ScholarCatalyst introduces a benchmark designed to study how researchers identify prior work that inspired their projects. The dataset was built by having 184 lead authors of 207 recent computer science papers label which candidates did or could have advanced their project, each providing a detailed rationale. This author-annotation approach makes scalable evaluation of retrieval systems possible, addressing a gap in how AI progress is measured against human research intuition.

Retrieval Task and Evaluation Setup

The core task: given an initial research question, retrieve the annotated papers from only the literature available when the project began. This constraint—using only pre-project literature—mimics the real-world scenario of building on existing work without knowledge of future developments. The evaluation uses Recall@20 (R@20) as the primary metric, measuring whether relevant papers appear within the top 20 results.

Agentic and Embedding Retrieval Results

Two retrieval approaches were compared against each other and against a strong baseline. Agentic search, which calls the same retriever as a tool during reasoning, achieved 0.42 Recall@20. Plain embedding retrieval achieved 0.48 Recall@20, significantly outperforming the agentic approach despite using the same underlying retriever. An agent built on Claude Fable 5.1, which may have seen the completed papers during training, reached only 0.51 R@20—only marginally better than embedding retrieval alone.

Interpretation of Results

The finding that agentic search does no better than plain embedding retrieval is particularly striking. It suggests that simply calling a retriever as a tool during reasoning does not inherently improve paper retrieval quality. The narrow margin by which Claude Fable 5.1 exceeds embedding retrieval (0.51 vs 0.48) indicates that even state-of-the-art language models struggle to effectively navigate broad research corpora without specialized training signal.

Implications for Scientific AI Agents

These results highlight the need for new training recipes that equip models with expert intuition for searching broad corpora. The benchmark reveals a fundamental difficulty: retrieving inspirational prior work is significantly harder than typical information retrieval tasks, and current agentic or retrieval-based approaches fall far short of human-level performance. ScholarCatalyst is envisioned as a step toward scientific agents that can take a half-formed idea and point to the prior research it needs.

Read the paper on arXiv