AI Papers Reader

Personalized digests of latest AI research

View on GitHub

AI Search Still Misses the Papers That Spark Research

Finding the prior paper that reshapes an early-stage research project is a core scientific skill, but AI search systems remain poor at it. A new benchmark from Stanford, Seoul National University and other institutions asks whether retrieval systems can find “catalyst papers”—prior work whose ideas did or could meaningfully advance a project—rather than merely topically similar papers.

The benchmark, ScholarCatalyst, draws on 184 lead authors of 207 recent computer-science papers. They labeled 894 research questions as they stood before each project’s key findings and identified which papers did or could have advanced their work, including uncited or previously unseen papers. Queries are either core research questions or subfield-specific questions. The search corpus holds 190,896 papers published before each source paper. Systems are scored by Recall@N—the share of author-credited papers in the top N—and nDCG@N, which rewards higher rankings.

Across sparse, dense, multi-vector retrievers and LLM search agents, the strongest system recovered only 48% of gold papers in its top 20. Agentic search did no better than embedding retrieval: in the paper’s summary, agents averaged 0.42 Recall@20 versus 0.48 for the retriever alone, even though agents called that same retriever. General-purpose dense retrievers performed best, with Qwen3-Embedding-8B reaching 0.37 Recall@20 on core queries and 0.51 on subfield queries. A GPT-4.1 tool-calling agent reached 0.37 and 0.43; an o3 deep-research agent reached 0.33 and 0.38. A Claude Fable 5.1 agent, which may have seen the completed source papers during training, reached only 0.51 Recall@20.

One failure illustrates why. For a core query about helping LLM agents use past experience in complex scenarios, the retriever’s top 30 contained three of four author-credited papers. But the GPT-4.1 tool-calling agent’s follow-up searches drifted toward generic transfer learning and never surfaced any of them. The authors attribute such misses to candidate coverage: an agent can only reason over papers its searches return, and repeated querying cannot recover inspirations the retriever never surfaces.

The paper also challenges citation-based proxies. In one spiking-neural-network reinforcement-learning example, authors credited two uncited papers—Jump-Start Reinforcement Learning and Revisiting the Minimalist Approach to Offline RL—for suggesting a secondary controller to bridge a warm-up period, while two cited, topically adjacent papers were labeled hard negatives. Across the benchmark, 43.6% of subfield-specific positives were not cited by the source paper. Similarity measures did not separate positives from hard negatives; hard negatives were at least as similar to queries on BM25 lexical and Qwen3-Embedding-8B semantic measures.

The authors acknowledge limitations. ScholarCatalyst covers computer-science papers from 2025–2026 with uneven area representation, and labels are retrospective author judgments subject to hindsight. Its 190,896-paper corpus is also far smaller than real literature—arXiv exceeds 3 million articles and Semantic Scholar indexes over 225 million—so the task is an easier version of what researchers face. The broader significance is that current systems can retrieve related work but still lack an expert-level sense of which ideas matter. ScholarCatalyst offers a way to measure progress toward that goal, not evidence that it has been achieved.