Research arXiv cs.CL

Can Deep Research Agents Retrieve and Organize? Evaluating the Synthesis Gap with Expert Taxonomies

TaxoBenchdeep research agentshierarchical taxonomyevaluation benchmark

TaxoBench is a benchmark built from 72 highly cited LLM surveys, 3,815 cited papers, and their expert-authored taxonomies, designed to jointly test retrieval and organization. It evaluates systems in two settings: Deep Research mode (end-to-end retrieval and organization from a topic) and Bottom-Up mode (given the expert paper set, isolating organization). Metrics include leaf-level ARI and V-Measure, plus hierarchy-level structure metrics US-TED, US-NTED, and Sem-Path. Across 7 Deep Research Agents and 16 LLM configurations, the best agent retrieves only 20.92% of expert-cited papers, and none of 70 standard Bottom-Up runs matches the experts' average taxonomy depth of 4.86. A controlled probe shows that models matching this depth do so by fragmenting the taxonomy, reducing alignment with the expert reference. Additionally, raw Sem-Path remains near a no-organization floor even when a newer model generation gains 3.68 pp ARI; after depth matching, humans outperform on all 10 matched surveys by 13.27 pp. These results identify retrieval and hierarchical organization as separate bottlenecks and underscore the need to calibrate hierarchy metrics before comparing models.

Read original →

← Back to home