Research arXiv cs.CL

Same Chart, Different Story: Bias in Vision-Language Chart Interpretation

vision-language modelschart interpretationbias benchmarkfairness

Vision-language models are increasingly used to interpret charts and generate natural-language explanations for socially consequential data, but they may produce different narratives for the same chart when only the referenced social group changes. This can reinforce stereotypes and mislead decisions, yet no benchmark existed for systematically evaluating bias in chart interpretation across social dimensions.

To address this, the authors introduce ChartBias, described as the first benchmark for auditing bias in VLM-based chart interpretation. ChartBias contains 820 manually curated real-world charts spanning six attributes: race, income, age, religion, immigration status, and gender. These yield 4,319 valid chart-attribute instances and 8,638 paired generations in which the chart is fixed and only the group term is swapped. Across 12 proprietary and open-source VLMs, the benchmark covers 155,484 model responses.

The evaluation finds three widespread failure modes: narrative shift (the same chart producing different narratives), group hallucination (assigning a chart to a group without evidence), and preference polarity (favorable trends often linked to one group). The authors also propose a multi-agent mitigation framework that serves as a strong baseline by separating chart-grounded evidence extraction from group-conditioned generation and using a counterfactual judge to verify that group-driven differences are supported by the chart. This framework substantially reduces narrative shift while preserving chart-grounded reasoning.

The findings show that evaluating chart understanding requires measuring not only accuracy but also fairness and consistency across social groups. ChartBias is released at https://github.com/vis-nlp/ChartBiasBench.

Read original →

← Back to home