Research arXiv cs.CL

What Does 99% Accuracy Measure? A Reproducible Audit of Shortcut Learning in a Widely Used Fake News Corpus

fake news detectionshortcut learningbenchmark auditISOT/Kaggle

Text classifiers trained on the ISOT/Kaggle "Fake and Real News" corpus routinely report accuracy and F1 above 0.98, a result that sits uneasily beside the difficulty of assessing veracity. The authors use a transparent TF-IDF and linear-classifier pipeline as a measurement instrument to audit the corpus along three leakage channels and two distribution-shift protocols, releasing all code and derived numbers.

The audit finds the benchmark is partly degenerate. A classifier given only the subject metadata field, with article text discarded, attains F1 = 1.000 because the two classes have disjoint subjects. Removing all three leakage channels—metadata, a newswire source tag present in 99.2% of real articles, and 6,251 duplicate documents contaminating 19.4% of a naive test split—lowers F1 by only 1.21 points, from 0.9935 to 0.9814. The residual signal is diffuse editorial style rather than a few giveaway tokens: deleting the 1,000 highest-weight unigrams still leaves F1 = 0.926.

That style signal does not transfer. Under a topic-disjoint protocol, average precision falls from 0.9995 to 0.9475 and deployed F1 from 0.9905 to 0.8067, with a prior-matched analysis confirming a genuine 5.2-point loss of discrimination; temporal transfer, by contrast, is nearly lossless. A fine-tuned DistilBERT is stronger in-distribution (F1 = 0.9993) but degrades far more under topic shift, losing 12.9 average-precision points versus the linear model's 5.2.

Transferred to the independent LIAR benchmark, all three models fall to near-chance ranking (ROC-AUC 0.54–0.57), none beating a majority-class baseline. The authors conclude that within-corpus scores on this dataset quantify source and topic separability rather than veracity, that added model capacity exploits the shortcut rather than avoiding it, and they recommend metadata-only, small-sample, and topic-disjoint baselines as inexpensive diagnostics for future work.

Read original →

← Back to home