Research arXiv cs.CL

EpiBench: Can LLMs Understand Epitopes for Antibody Drug Discovery?

EpiBenchantibody drug discoveryepitope reasoningLLM evaluation

Epitope understanding is critical for antibody drug discovery because epitopes determine where antibodies bind and influence properties like functional blockade and escape resistance. However, existing resources focus on isolated prediction tasks or require specialized structural data, and general protein benchmarks do not assess epitope-centered reasoning across the antibody development workflow. To fill this gap, the authors introduce EpiBench, a closed-book, sequence-based benchmark for evaluating LLM epitope reasoning.

EpiBench contains 1,609 curated samples grounded in structural antibody-antigen contacts, functional B-cell assays, and deep mutational scanning escape measurements. It spans five connected tasks: targetable region discovery, antibody-conditioned epitope identification, epitope binning, functional epitope assessment, and antibody escape assessment, with controlled sampling to reduce shortcut-based evaluation artifacts. The authors evaluate nine general-purpose LLMs and analyze their performance via task-specific baselines, antigen length stratification, explicit-reasoning comparison, and failure-mode inspection.

Results show that current LLMs capture partial epitope-related signals but are limited in antibody-specific sequence grounding, long-context residue localization, and biologically grounded reasoning. EpiBench thus provides a diagnostic testbed for measuring and improving sequence-aware biomedical LLMs toward reliable LLM-assisted antibody discovery, highlighting specific weaknesses that future models and training strategies should address.

Read original →

← Back to home