Scientific research often begins by synthesizing ideas from a set of related papers to identify gaps and formulate new directions. However, training language models to perform this form of literature-grounded ideation remains challenging, as existing approaches based on prompting or feedback lack structured supervision for how papers should be synthesized. We introduce IdeaAnchor, a paradigm for training LLMs to perform research ideation using structured specifications as privileged signals. Each IdeaAnchor instance encodes how each input paper should be synthesized into a successful idea, including their functional roles, relationships, and target synthesis criteria. We build this paradigm by mining instances from published papers, capturing how real ideas emerge from prior literature. We then train models via demonstration, self-distillation, and reinforcement learning, and further enhance generation with retrieval at inference time. Experiments show consistent improvements in ideation quality. Our analysis reveals a functional decomposition: anchor-based training strengthens creative synthesis, retrieval enhances detail elaboration, and combining both yields the best performance.
Why LLM-Mediated Ideation Lacks Structured Supervision
LLMs have demonstrated impressive downstream scientific assistance, from summarizing and reasoning over papers to executing well-defined experiments. Even so, generating novel and actionable research ideas remains challenging. Given some related papers, LLMs tend to produce fluent overviews or shallow concept blends, which often struggle to identify meaningful research gaps and develop concrete methods grounded in existing scholarly evidence. Comparing ideas generated from the same prior works, other work finds a consistent distributional gap from human research taste: LLM ideas concentrate on bridge-like opportunities and synthesis methods, whereas human papers frame gaps and construct contributions in far more diverse ways. Recent efforts address this through goal-conditional plan generation with rubric rewards, execution feedback in sandboxed environments, or single-paper hypothesis inversion. However, these approaches either assume the research goal is already given, require domain-specific execution environments, or do not support multi-paper synthesis.
The authors argue that the bottleneck of ideation is the absence of training signals that capture what makes a good research idea emerge from prior work. To address this, they introduce IdeaAnchor, per-instance and literature-derived specifications that encode the functional roles each input paper should play, the gaps that should be identified, and the patterns that should be leveraged in the synthesis. IdeaAnchor serves as a unified data interface for improving research ideation via three distinct training paradigms: Anchor-Guided Demonstration, Anchor-Based Self-Distillation, and Anchor-Privileged Reinforcement Learning. At inference time, the framework uses either related papers or a topic query as input and can incorporate retrieval augmentation, which complements the learned synthesis capability by supplying factual and methodological details for elaborating generated research ideas.
IdeaAnchor: Mining and Teaching Paper-to-Idea Synthesis
This section first defines the task of literature-conditional research ideation and then introduces IdeaAnchor, the structured specifications that form the backbone of the training framework. Let P={p₁,…,p_N} denote a set of N related papers within a research topic, where each paper p_i consists of a title t_i and a brief summary c_i. A model π_θ is required to produce a structured output y=(T,M,R) consisting of:
- Thinking Trace T: a structured analysis of each input paper's functional role and inter-paper relationships, articulating the synthesis logic that leads to a new research direction.
- Research Motivation M: a novel problem definition derived from identifying gaps, limitations, or unexplored combinations across the input papers.
- Research Plan R: a detailed methodology that synthesizes techniques from the input papers to address the identified problem.
Here, the research motivation M must be discovered from the literature rather than given as input, and the plan R must be grounded in the specific papers provided.
An IdeaAnchor A_D for a training instance is a structured specification derived from a ground-truth paper D and its input literature P. It encodes three types of information:
- Functional Role Assignments {ρ_i}_{i=1}^N: Each prior work p_i is assigned one of four roles that characterize its relationship to D: Direct Predecessor (the method most immediately extended or improved upon), Inspiration Source (a technique or idea from a different context that sparked the approach), Gap & Motivation (work whose limitations define the research problem), and Methodological Ingredient (a specific technical component incorporated into the proposed method). Each role assignment is accompanied by a learned-insight summary: what the authors of D learned from p_i and how it influenced the contribution.
- Relationship Analysis S_D: A structured analysis that articulates the synthesis logic connecting the prior works P to D's contribution, including per-paper analysis (what each work achieved, what limitation remains, why that limitation matters) and cross-paper relationship analysis (how the works relate to each other, whose ideas address whose limitations, what combinations open new possibilities, and how the logical chain across papers points toward an unexplored direction).
- Checkable Criteria U_D={u₁,…,u_K}: A set of verifiable items that a successful idea should satisfy, spanning motivation criteria (does the output identify the key gap or limitation that D addresses?), method criteria (does the proposed approach incorporate details from the prior works?), and overall criteria (is the synthesis coherent, and does the proposed direction logically follow from the cross-paper analysis?).
The value of IdeaAnchor is that they are instance-specific and literature-grounded: rather than generic quality criteria, each anchor is derived from the particular way a real set of papers led to a real published contribution. This makes anchors a form of privileged information available at training time (from the ground-truth paper) but not at inference, when the model must synthesize without knowing the answer. The anchor concept unifies the structured reasoning process, the rewards signal for RL, and the specification for high-quality demonstrations. By framing all three as aspects of a single underlying specification, the authors enable a comparison of different training strategies that exploit the same information source.
Building IdeaAnchor Instances from Published Literature
The authors construct IdeaAnchor at scale by reverse-engineering the intellectual genealogy of existing research papers. The pipeline is automated, using a sample creator model to progressively build the structured specification from a ground-truth paper in three stages:
- Prior work extraction. Given the full text of the target paper, the model extracts the core idea and identifies 5–7 prior works that substantively shaped the target's contribution. Each prior work is assigned a functional role and a learned-insight annotation.
- Literature enrichment. The extracted prior works are queried against scholarly search engines to retrieve their abstracts, providing factual grounding beyond how the target paper describes them.
- Criteria generation. Given the prior works, corresponding analysis, and reference proposal, the model generates several checkable criteria spanning motivation, method, and overall dimensions, ensuring that evaluation criteria are anchored in the specific intellectual genealogy of the target paper.
Beyond being topically related, the mined prior works are expected to be the ones that actually shaped the target's core idea. To verify this, the authors sent authors of benchmark papers the extracted prior works together with the reconstructed idea-formation path and asked whether these reflect how their idea formed, which important works are missing, and which included works were not important. Responses from 20 authors found 16 satisfied with the mined set; five named one to three missing works (9 in total), so the mined sets cover 93.7% of the works the authors credit; two authors flagged included works as unimportant, amounting to only 3 of the 134 extracted works (2.2%). The mined sets therefore closely reflect the prior works that authors themselves credit for their ideas.
The pipeline was applied to two domains: (i) Machine Learning, collecting 7,494 accepted papers from ICLR, ICML, and NeurIPS spanning 2023–2025; and (ii) Natural Science, 6,689 papers published from 2023–2025 on Nature Communications spanning 71 major scientific subjects. An evaluation benchmark was constructed from 924 papers accepted at ICLR 2026, including all 224 oral presentations. For each paper, the same pipeline generated candidate prior works and rubric criteria, which were then reviewed and refined by expert annotators who verified correctness of extracted prior works, adjusted role assignments where necessary, and edited criteria items to ensure they are unambiguous and faithfully reflect the paper's contribution.
Training LLMs via Demonstration Self-Distillation and RL
All three training strategies treat IdeaAnchor as privileged information, which is available during training but absent at deployment, and differ in how they convert it into learning signal.
Anchor-Guided Demonstration
Direct use of privileged information is to steer a strong teacher model toward higher-quality demonstrations. Without structured guidance, teacher outputs might be fluent but often fail to deeply engage with the input literature. Providing a specification alongside the input paper list yields demonstrations that are more grounded. For each instance (P,A_D), an external LLM is prompted with both inputs to produce a structured proposal. The target model is trained on the expert output with the anchor removed via supervised fine-tuning:
L_demo = -E_(P,y_demo)[log π_θ(y_demo | P)]
The anchor's influence is thus embedded in the demonstration: the target model learns to satisfy anchor criteria without ever observing A_D at inference.
Anchor-Based Self-Distillation
Self-distillation removes the external-teacher requirement by exploiting the asymmetry between a model's output quality with versus without privileged context. The privileged-conditioned model serves as its own teacher while the unprivileged model is the student. A system persona is used to prevent the model from revealing privileged information during thinking, avoiding training illusions. We sample y_self = π_θ(P,A_D) and train π_θ on its own anchor-conditioned outputs with A_D removed:
L_self = -E_(P,y_self)[log π_θ(y_self | P)]
Training is initialized from the base model, while the judge is flexible to be either frozen or evolving during training.
Anchor-Privileged Reinforcement Learning
Demonstration-based methods optimize a proxy (matching teacher output) rather than directly optimizing what defines a good idea. RL closes this gap: instance-specific privileged knowledge supplies a reward that is both more informative and harder to game than generic criteria. A judge receives (P,y,A_D) and scores each checkable item u ∈ U_D:
r_anchor = (1/|U_D|) ∑_{u∈U_D} I[θ_r(P,y,u) = satisfied]
Each u_i is deemed satisfied only if no general quality guideline (specificity, soundness, feasibility) is violated. The authors optimize with GRPO: for each P the policy samples G candidates, scored by reward(y) = r_anchor(y) - λ ⋅ I{format violation}, and updates toward higher-scoring ones. Training is initialized from the base model, while the judge is flexible to be either frozen or evolving during training.
Retrieval-Augmented Idea Generation at Inference
Training internalizes creative synthesis, gap recognition, and direction formulation. These capabilities let the model effectively conceptualize promising research directions, but lack specific details to translate ideas into actionable plans. Producing actionable proposals requires detail elaboration, which the authors address with two inference-time mechanisms.
Role-Aware Retrieval Augmentation
When only abstracts are available, the model lacks methodological depth for tangible plans. The model first generates a thinking trace that assigns each prior work a functional role. The system then fetches targeted full-text sections, including methods for Key Methodology, results for Primary Baseline, and problem statements for Gap & Motivation, and produces the final proposal from the augmented set:
(M,R) = π_θ({(p_i,d_i)}_{i=1}^N, T)
where d_i = Retrieve(p_i,ρ_i) denotes role-guided passages.
Automated Topic-Driven Ideation
The model also accepts a high-level topic as an alternative to specific papers. It queries a search API to retrieve candidate papers, then applies the same synthesis pipeline. This pipeline can be optionally combined with full-text retrieval to generate a proposal in an end-to-end manner.
Experimental Validation Across Ideation Benchmarks
The authors evaluate whether IdeaAnchor-driven training improves literature-conditional research ideation beyond prompting alone. Experiments ask three questions: (i) whether specialized training improves over the Qwen3-8B base model and approaches strong proprietary models; (ii) how retrieval-augmented inference complements training; and (iii) whether the learned synthesis behavior transfers across training scale, domains, and open-ended topic-driven ideation.
Experimental Setup
All trained policies are initialized from Qwen3-8B and Qwen3.5-9B. The mined IdeaAnchor corpus, each training instance consists of 5 to 7 prior papers represented by titles and abstracts, together with the privileged IdeaAnchor extracted from the target paper. The three training paradigms differ only in how this privileged information is converted into learning signal. For anchor-guided demonstrations, GPT-4o-mini is prompted to get the structured proposal for fine-tuning. For self-distillation, the trained model itself is prompted with the anchor and then trained on its own anchor-conditioned output under the unprivileged input. For RL, candidate proposals are scored item-by-item against the anchor criteria by GPT-4o-mini judge, and the scores are used as rewards to optimize the policy. All model outputs follow the same structured format: y=(T,M,R).
At inference time, both abstract-only generation and two retrieval variants are evaluated. RAG-Full follows the method: the model first assigns functional roles to input papers during thinking, then retrieves role-relevant passages and generates from the augmented context. RAG-Summary uses the same retrieved passages but compresses each paper into a targeted summary by merging the abstract with the retrieved evidence, keeping input length and style close to the training distribution while injecting more task-directed methodological information.
For automated evaluation, GPT-4o judges the generated proposal against the expert-refined checkable criteria. Each criterion is marked satisfied only if the proposal puts forward a concrete, literature-grounded claim that meets the criterion while maintaining basic soundness and feasibility. The criteria satisfaction rate (CSR) is reported, averaged across all rubric items in the benchmark. For human evaluation, blind pairwise comparison is conducted where experts rank overall strength and assess literature grounding, synthetic novelty, methodological specificity, and feasibility.
Main Results
On the ICLR 2026 benchmark with 924 instances, every training paradigm improves its base model. On Qwen3-8B, self-distillation reaches 16.3% CSR (+5.4), SFT 21.7% (+10.8), and RL 24.6% (+13.7), versus 10.9% for the base. On Qwen3.5-9B, SFT raises CSR from 31.1% to 48.7% and RL to 54.2%, with RL surpassing proprietary model references. RL gives the largest gain at both scales even though its reward contains only instance-specific criteria, suggesting the policy learns a general pattern of gap identification and cross-paper synthesis instead of surface rubric matching. Retrieval adds smaller improvements: 25.6% for RL with RAG-Full and 22.9% for SFT with RAG-Summary, so access to more evidence cannot replace learned synthesis.
Pairwise Ranking Results
Beyond item-wise rubric satisfaction, pairwise preference ranking captures holistic synthesis quality. The judge receives two proposals, one from a trained variant and one from the base model, in randomized order and selects the one that presents better. All three training paradigms achieve more than 75% pairwise win rate against the base model. RL leads, followed closely by SFT. Even self-distillation, which requires no external teacher, attains substantial preference gains, reinforcing that the anchor signal itself is the primary driver of improvement.
Functional Decomposition: Anchor Training versus Retrieval
The authors analyze how much each component contributes to different stages of the ideation pipeline. The key finding is that gap recognition and direction formulation need to be internalized via training, whereas grounding conceived ideas into methodological details can be enhanced by inference-time retrieval.
Table 1 shows that both RAG-Full and RAG-Summary are strongly preferred to abstract-only generation across all training paradigms, despite modest CSR gains. Manual inspection of 100 paired outputs resolves this discrepancy. Retrieved passages add method choices, experimental settings, baselines, and implementation constraints, making proposals more actionable and improving holistic preference. They rarely change the high-level gap, motivation, or direction established earlier during creative synthesis. Because CSR primarily rewards alignment with that direction, extra technical detail cannot recover a proposal whose synthesis is wrong.
This distinction also explains two automated trends. Retrieval slightly hurts the base model because unreliable role assignment can surface evidence for a weak premise and amplify it. Conversely, RAG-Full slightly outperforms the distribution-matched RAG-Summary for the strongest RL policy: once the model chooses a sound direction, richer full-text evidence becomes more useful than avoiding distribution shift. Retrieval is therefore best at elaborating a direction the model has already formulated, not deciding that direction from scratch. Accordingly, CSR and pairwise preference measure complementary stages of the pipeline: the former emphasizes selecting the right research direction, while the latter also rewards how fully that direction is operationalized.
Cross-domain generalization was tested by training on natural science anchors and evaluating on the machine learning benchmark. Dashed curves in Figure 5 show that natural-science-only training transfers reasonably well to machine learning ideation, though machine-learning-specific anchors yield higher absolute performance. This suggests that the core synthesis patterns are partially domain-agnostic, but domain-specific anchors provide an edge.
From-scratch ideation was evaluated by feeding only a topic query to the model, which then searches for relevant papers and generates a proposal. RL with retrieval achieves the highest scores, demonstrating that the learned synthesis capability composes with automated literature discovery. The pipeline handles the full ideation workflow from topic to concrete proposal without manual paper selection.
Constraints and Trade-Offs in Anchor-Based Ideation
The authors acknowledge several limitations. The mining pipeline depends on a sample creator model that may introduce errors in role assignment or criteria generation. While author validation shows high coverage (93.7%) of credited works, the remaining gap could matter for papers with longer intellectual genealogies. The evaluation benchmark, though expert-refined, inherits the ICLR 2026 distribution and may not generalize to all scientific fields. The automated judge (GPT-4o) provides scalable evaluation but may not fully capture the nuance of expert research judgment, though pairwise human evaluation confirms the automated trends. The self-distillation anti-leak strategies required careful design: the naive baseline leaked anchor content in 100% of outputs, while the chosen persona-knowledge approach achieved 0% leakage with perfect rubric compliance. Other strategies either leaked slightly (critique-edit at 4%) or sacrificed compliance (bottleneck PI at 52.2%). Training compute requirements for RL are substantial: GRPO with G candidates per instance and iterative reward optimization demands significant GPU resources.
How Developers Can Use IdeaAnchor
The IdeaAnchor framework offers a practical pathway for teams building AI research assistants. The mined dataset of 14K anchor instances across machine learning and natural science provides a starting point for training custom ideation models. The three training paradigms form a spectrum: self-distillation requires only a base model and the anchor corpus, SFT adds an external teacher, and RL adds a judge model for reward optimization. For inference, the role-aware retrieval module can be integrated with any scholarly search API and full-text access layer. Developers should note that anchor quality depends on the sample creator model; the paper uses GPT-4o for mining, but open-weight alternatives may need prompt adaptation. The evaluation rubrics (CSR, pairwise ranking) are reusable for benchmarking new ideation systems. For deployment, the topic-driven pipeline enables end-to-end ideation from user queries without manual paper curation. The functional decomposition finding suggests a practical design rule: invest training compute in creative synthesis (anchor-driven RL or SFT), then add retrieval at inference for methodological depth. This separation of concerns mirrors how human researchers work: internalize the art of gap-finding, then look up details when drafting proposals.