WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research

[Paper] WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research

WebWeaver, a groundbreaking dual-agent framework led by Zijian Li, revolutionizes AI-driven research by mimicking human processes. It outperforms existing systems on major benchmarks, addressing key limitations in current research agents. The framework features a dynamic planner that continuously refines research outlines based on new evidence, and a writer that employs hierarchical synthesis for efficient information management. WebWeaver's memory bank architecture ensures strong source-groundedness in final reports. Extensive experiments demonstrate its superior performance across challenging open-ended deep research tasks. The approach can be distilled into smaller models, enabling more accessible AI to achieve expert-level performance. WebWeaver represents a paradigm shift in tackling complex, information-intensive tasks, paving the way for more human-like artificial intelligence.

In a significant advancement for artificial intelligence research, a team led by Zijian Li has introduced WebWeaver, a groundbreaking dual-agent framework that reimagines how AI systems conduct comprehensive research. Published on September 16, 2025, this innovative framework tackles the complex challenge of open-ended deep research (OEDR) by mimicking human research processes rather than adhering to rigid, mechanical workflows.

WebWeaver outperforms both proprietary and open-source systems on major benchmarks, establishing a new state-of-the-art in AI-driven research capabilities. This achievement stems from its innovative approach that addresses fundamental limitations in existing research agents, particularly their inability to adapt research outlines based on discovered evidence and their struggle to manage long-context information effectively.

The research team identified two critical flaws in current AI research systems. First, most existing agents either employ a simplistic “search-then-generate” approach that lacks structural coherence or create a static outline before gathering any evidence, preventing adaptation to new discoveries. Second, these systems typically attempt to process all gathered information in a single context window, leading to attentional failures known as “loss in the middle” and increased hallucinations.

WebWeaver’s architecture resolves these issues through a human-inspired dual-agent framework. The first agent, called the planner, operates in a dynamic research cycle that continuously interweaves evidence gathering with outline refinement. Unlike traditional systems that create a fixed outline at the beginning, WebWeaver’s planner treats the outline as a living document that evolves as new information emerges. This dynamic approach enables genuine exploration and allows the research direction to adapt based on discoveries.

“The key lies in abandoning rigid, machine-like pipelines and instead embracing the organic process of human intellect,” explains the research paper. “Our approach teaches the agent to research like a person. A human expert doesn’t finalize their entire plan before starting; they allow their outline to be a living document.”

The second agent, the writer, addresses the challenge of managing extensive information through a hierarchical synthesis process. Rather than attempting to process all gathered evidence simultaneously, it composes the report section by section, performing targeted retrieval of only the most relevant evidence for each part. This focused approach mirrors how human writers work, referencing specific notes for specific chapters, thereby avoiding the attentional failures that plague one-shot generation methods.

The system’s memory bank architecture serves as the bridge between these agents, storing structured evidence with unique identifiers that are explicitly linked to the outline through citations. This connection ensures that the final report maintains strong source-groundedness while enabling more efficient information management. For a deeper understanding of AI memory systems, consider reading Artificial Intelligence and Brain Research: Neural Networks, Deep Learning and the Future of Cognition.

Extensive experiments demonstrate WebWeaver’s exceptional performance across three challenging OEDR benchmarks. On DeepResearch Bench, which features PhD-level research tasks across 22 distinct fields, WebWeaver achieved superior scores in comprehensiveness, insight, instruction-following, and readability compared to systems from major AI companies. On DeepConsult, focused on business and consulting domains, it achieved the highest win rate of 66.86%. And on DeepResearchGym, which evaluates performance on real-world complex queries, it earned top scores for clarity, depth, balance, breadth, and support.

The research team’s analysis reveals that the dynamic outline optimization process is not merely an enhancement but a fundamental requirement for high-quality research. Reports generated after multiple rounds of outline optimization showed significant improvements in comprehensive coverage and logical structure. Additionally, the hierarchical writing approach dramatically outperformed brute-force methods across every metric, confirming that breaking down complex writing tasks into focused subtasks is essential for coherent, well-supported reports.

Perhaps most remarkably, the team demonstrated that WebWeaver’s approach could be distilled into smaller models through supervised fine-tuning. By creating a high-quality dataset called WebWeaver-3k, generated by their framework, they enabled 30-billion-parameter models to achieve expert-level performance previously confined to much larger systems. This achievement suggests that the fundamental skills of thinking, searching, and writing can be effectively taught to more accessible AI models. Those interested in implementing these techniques should explore Mastering Fine-Tuning with LLMs: From Basics to Advanced Techniques.

WebWeaver represents more than just an incremental improvement in AI research capabilities. It establishes a new paradigm for tackling complex, information-intensive tasks by reframing the challenge of long-context reasoning. Rather than relying on brute-force attention mechanisms, it shows how structured, tool-driven approaches can effectively orchestrate complex information workflows, pointing the way toward more human-like artificial intelligence.

The introduction of WebWeaver signifies a crucial step toward autonomous AI systems that can truly assist with the complex, open-ended challenges that define human-level knowledge work—a process driven by curiosity, synthesis, and the discovery of novel insights. For a comprehensive overview of modern AI advancements, A Brief History of Intelligence: Evolution, AI, and the Five Breakthroughs That Made Our Brains provides valuable context on how these developments relate to human cognition.

🚀 Unlock Ads-Free Experience At $5/year

14 days free trial Cancel anytime

1 thought on “[Paper] WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research”

  1. Haha, finally an AI that does research like a human! No more rigid pipelines, just a dynamic outline that evolves like our attention spans. And managing all that info with a hierarchical approach? Now *that*’s organizing! Though, I wonder if the planner gets distracted by shiny new evidence half the time, like me. And the writer better not start making things up just because it can’t find the right citation – we all know how unreliable human memory is (especially mine). Still, turning billion-parameter beasts into capable little models is a win, even if they still might hallucinate a bit. This WebWeaver sounds like it’s finally teaching AI to think outside the (context) window!MIM

Leave a Comment

Your email address will not be published. Required fields are marked *