英文标题:Benchmarking Hybrid Deep Research Across Database Querying and Web Search
作者:Ruofan Wu, Peiran Xu, Xiaolong Li, Fan Shu, Soyoung Yoon, Yite Wang, Xiaodong Yu, Boyi Liu, Feng Yan, Debiao Li, Yuxiong He, Zhewei Yao
arXiv ID:2609.09410 | 分类:cs.CL | 发表:2026-09-08
许可:CC-BY
摘要 尽管自主智能体在“深度研究”方面已取得显著进展,能够迭代地浏览开放网络以综合信息,但现实世界的问题求解很少局限于单一环境。复杂的分析任务本质上要求智能体将来自模糊的非结构化文本(例如开放网络)与高度精确的结构化数据(例如关系数据库)的证据交织在一起。然而,现有基准测试孤立地评估这些模态,未能捕捉关键的“交接”——即在系统之间转移证据时保持约束的能力。我们提出 HybridDeepResearch,据我们所知,这是首个要求同时使用网络搜索和 SQL 才能形成完整、可验证答案的深度研究基准。该基准包含 380 个依赖工具的任务,基于 LiveSQLBench-Base-Lite 数据库和公共
While autonomous agents have made significant strides in "deep research" by iteratively navigating the open web to synthesize information, real-world problem-solving is rarely confined to a single environment. Complex analytical tasks inherently require agents to weave together evidence from both ambiguous unstructured text (e.g., the open web) and highly precise structured data (e.g., relational databases). However, existing benchmarks evaluate these modalities in isolation, failing to capture the critical "handoff" - the ability to preserve constraints when moving evidence between systems. We introduce HybridDeepResearch, to our knowledge the first deep-research benchmark that requires both web search and SQL to form a complete, verifiable answer. The benchmark contains 380 tool-dependent tasks grounded in LiveSQLBench-Base-Lite databases and public web corpora, validated through automated checks and human review, and covering three reasoning patterns: SQL2S, S2SQL, and Parallel. Evaluations across proprietary and open-weight models under various agentic scaffolds reveal that even state-of-the-art models like GLM-5.2, Claude-Sonnet-4.6 and GPT-5 achieve only about 50-54% Pass@8 on the hard subset. Notably, results show that directional reasoning is substantially more difficult than parallel intersection, highlighting that bridging structured and unstructured information spaces without losing constraints remains a major open challenge for agentic systems. Code and datasets are publicly available at GitHub (https://github.com/Snowflake-AI-Research/HybridDeepResearch) and Hugging Face (https://huggingface.co/datasets/Snowflake/HybridDeepResearch).
查看完整双语翻译 →
正在跳转到翻译阅读页… 如果没有自动跳转,请点击这里。