Human-Like Anaphor Resolution in Large Language Models
The paper (arXiv:2608.05630v1) investigates anaphor resolution in five open-weight LLMs, testing whether factors identified in cognitive science—such as discourse structure, situation-model properties, and semantic cues—impact model performance. Unlike typical benchmark evaluations, this work draws directly on psycholinguistic theory to design controlled experiments. Findings could inform how LLMs handle coreference in longer contexts and highlight gaps between human and model comprehension. The use of open-weight models allows for reproducibility and further analysis by the research community.