Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification
Retrieval-based factuality evaluation, in which LLM-generated claims are verified against evidence from authoritative medical corpora, has become the dominant paradigm for scalable hallucination detection in high-stakes clinical settings. Yet most systems report aggregate metrics such as F1, which obscure where and why failures occur, and existing RAG diagnostics require gold answers or annotated gold evidence that do not exist in this open-ended regime.
The paper introduces two comprehensive taxonomies grounded in a case study on the open-ended MedExpert dataset and three closed-ended datasets. One taxonomy decomposes retrieval-stage errors along five quality dimensions, while the other breaks verifier-reasoning errors into six consecutive steps. The authors adapt an automatic pattern induction pipeline using LLM-as-Judge to label evidence quality and classify verifier reasoning errors at scale, then stress-test their findings across four retrieval methods and six frontier verifier models.
The analysis finds that scaling model size, adding reasoning effort, expanding to authoritative web sources, and applying medical fine-tuning do not resolve these failure modes. This indicates that the failures represent fundamental limitations of the retrieve-then-verify paradigm in open-ended medical settings rather than artifacts of outdated systems. The authors release their code and data at https://anonymous.4open.science/r/Medical_RAG_eval-4AB5 for full reproducibility.