The Limits of Simulated Societies: How Post-Training and Survey Fine-Tuning Erase Cross-Cultural Variance
LLMs are increasingly used to simulate diverse human populations for computational social science, but most evaluations score only the average response for a group rather than the spread of opinion within that group. The authors build a diagnostic framework that measures point accuracy alongside dispersion retention — the ratio of predicted to human standard deviation (dr) — tested on 10,000 respondent–question pairs from the World Values Survey covering twelve countries and six continents.
They evaluate eleven zero-shot language models and five variants fine-tuned on WVS data using SFT, DPO, and GRPO. This reveals a failure mode the authors term consensus collapse, in which alignment training compresses model outputs toward a single stereotype per group.
Along the post-training trajectory from the Llama 3.1 70B base model to the Tulu 3 checkpoints, the first stage — supervised instruction tuning — removes half of the response spread with minimal accuracy gain (dr falls from 1.22 to 0.59, while accuracy rises only 0.9 points); later stages do not restore it, and a gap opens between WEIRD and non-WEIRD countries that survey fine-tuning then deepens while chasing higher point accuracy. The most accurate model (Tulu 3 70B-DPO fine-tuned on WVS, 57.9%) retains half the human spread overall (dr = 0.50) and only 11% of it for Nigeria, versus 0.70–0.87 for WEIRD countries. Raising sampling temperature to 1.0 leaves the Wasserstein-1 distance to human distributions unchanged for both fine-tuned DPO models, and GRPO on Qwen 3.5 9B fails to restore spread under either an accuracy reward or a distribution-shaped reward. Mixing the aligned model with an unaligned prior raises dr from 0.51 to 0.62 on a held-out split but leaves Nigeria at 0.36.
The authors conclude that point accuracy alone misjudges these simulators, and that current post-training methods trade diversity for consensus — an important caveat for anyone using LLMs to represent cross-cultural opinion.