Evaluating deep-research agents from answers to reports
Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports
Paper / source record · 2025-10-02
Research question and approach
Dr. Bench evaluates long research reports with 214 expert-curated tasks across 10 domains. It measures semantic quality, topical focus and retrieval trustworthiness using rubrics, topic keywords and reference-source links.
When this work is relevant
Use this work for report-level evaluation of deep-research agents. Its source-matching metric checks exact links and hostnames against reference sources; it does not establish that every cited sentence is supported. Preserve the task, model and benchmark version when comparing results.
Author-written abstract
As an embodiment of intelligence evolution toward interconnected architectures, Deep Research Agents (DRAs) systematically exhibit the capabilities in task decomposition, cross-source retrieval, multi-stage reasoning, information integration, and structured output, which markedly enhance performance on complex and open-ended tasks. However, existing benchmarks remain deficient in evaluation dimensions, response format, and scoring mechanisms, limiting their effectiveness in assessing such agents. This paper introduces Dr. Bench, a multidimensional evaluation framework tailored to DRAs and long-form report-style responses. The benchmark comprises 214 expert-curated challenging tasks across 10 broad domains, each accompanied by manually constructed reference bundles to support composite evaluation. This framework incorporates metrics for semantic quality, topical focus, and retrieval trustworthiness, enabling a comprehensive evaluation of long reports generated by DRAs. Extensive experimentation confirms the superior performance of mainstream DRAs over web-search-tool-augmented reasoning models, yet reveals considerable scope for further improvement. This study provides a robust foundation for capability assessment, architectural refinement, and paradigm advancement of DRAs.
Abstract source: https://arxiv.org/abs/2510.02190. Checked 2026-09-14. Bibliographic metadata uses the linked paper record or author-maintained catalog. Results, limitations and experimental settings remain defined by the original source.
Citation
BibTeX is preserved from the author-maintained citation repository. The CFF uses the source metadata shown on this page. Version titles or author lists can differ; choose the version you used.
@article{yao2026dr,
title={Dr. bench: A multidimensional evaluation for deep research agents, from answers to reports},
author={Yao, Yang and Wang, Yixu and Zhang, Yuxuan and Lu, Yi and Gu, Tianle and Li, Lingyu and Zhao, Dingyi and Wu, Keming and Wang, Haozhe and Nie, Ping and others},
year={2026}
}