CARAT: Do Materials LLMs Reason or Recite?
The paper asks whether a materials LLM answering a crystal-structure question does so by reasoning from the structure or by copying an answer already printed in its input. Standard accuracy cannot distinguish these, because a structural description often prints the very field it is scored against. CARAT is designed to separate reasoning from recitation.
CARAT holds the question and gold answer fixed across eight matched views, names each structural relation separately in GraphSpace, and combines this with matched fine-tuning, answer masking, evidence injection, paired inference, and a rule that can withhold claims. It also stress-tests the benchmark itself: a shortcut rule that skips the link and reads the list directly answered four of seven hardened families, so the authors rebuilt the benchmark until eleven such shortcuts sat near chance.
On the benchmark's hardest families, the grounded view is worth 17.3 points over formula inputs. GraphSpace beats a plain periodic graph by 19.3 points, but this margin conflates two effects: where the plain rendering already contains everything the question needs, the advantage is only 1.96 points, while where it omits those fields entirely, it is 46.7 points. The headline therefore largely measures what the baseline lacked, not how evidence is presented.
A frozen model quotes the structural link yet answers the same when the link is redirected in 95.6% of paired cases, indicating it repeats the relation without actually using it. After matched supervision that figure reaches 99.8%, and deleting the link drops performance to 23.4%—below the 27.0% achieved by the best shortcut—suggesting both using the relation and withholding claims are learnable.