RATIO: A Benchmark for Retrieval Across Typed Ideation Operations in Scientific Literature
Abstract
RATIO is a large-scale benchmark for retrieving scientific literature via three ideation operations—addressing problems, broadening concepts, and specifying instances—to support literature-grounded ideation.
Retrieved scientific literature can serve as inspiration for both human and AI scientists. Inspiration can take different forms: prior work may directly suggest how to address a problem, or surface directions at different levels of abstraction - zooming out to a more general view or zooming in to a concrete realization. We introduce RATIO (Retrieval Across Typed Ideation Operations), a large-scale benchmark in which relevance is defined by three operations which we name ideation moves: Address retrieves potential approaches for stated problems, Broaden retrieves more general formulations, and Specify retrieves concrete instantiations. RATIO is constructed from millions of full-text scientific papers across CS literature via a general recipe that extends discourse-marker distant supervision - previously used only for classification - to corpus-scale retrieval, combined with extensive LLM and human vetting. Experiments show that operation-specific fine-tuning substantially boosts retrievers but leaves much room for further improvements. RATIO provides a scalable training and evaluation framework for retrieval components that support literature-grounded ideation, opening up new research avenues on scientific inspiration retrieval.
Get this paper in your agent:
hf papers read 2608.27394 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 4
maayans/modernbert-embed-large__ratio-broaden
Datasets citing this paper 1
maayans/RATIO
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper