Recognition, Simulation, and Refusal: A Contamination-Aware Study of Classic Psychological Effects in LLM Agents
An LLM producing a response pattern associated with a human psychological effect is not the same claim as the LLM possessing that bias. The paper presents PsyAgentBench, a benchmark that re-runs classic psychology experiments on LLM agents under a factorial design built to separate these possibilities: each paradigm is run with the paradigm explicitly labeled in the prompt (named) or framed as a routine task (blind), and on the literal textbook version of the task (canonical) or a structurally matched variant written to reduce lexical and scenario overlap with likely training data (counterfactual), crossed with a persona manipulation. It covers five completed paradigms, evaluated on up to three open-weight model families, with 41,904 trials released.
Across those paradigms, apparently human-like effects arise through qualitatively different routes rather than one shared susceptibility. Asch conformity shows paradigm-label gating with explicit override, going from 0 percent blind to 83.3 percent named on gpt-oss-120B. Anchoring shows knowledge-dependent signal reliance, exactly zero on grounded facts versus near total on invented quantities—a pattern equally consistent with rational use of the only available signal. Framing shows amplification on novel content under labeling, sunk cost shows robust absence, and minimal-group allocation shows safety-mediated selection where refusal itself is the primary finding.
A one-sentence persona change (agreeableness, framed as an instruction rather than a verified trait manipulation) eliminates, dampens, or reverses these effects depending on which effect it is, arguing against any single response-bias account. The authors further formalize, and in two cases document empirically, three ways a psychology paradigm can fail to port to LLM agents: persona dominance, population collapse, and safety selection. They argue that scalar bias-susceptibility scores obscure this structure and instead report replication profiles.