Can Circuit Alignment Predict OOD Generalization?
Existing representational similarity metrics such as CKA, SVCCA, and RSA compare activations rather than forecast generalization, and the paper shows they are provably insensitive to structural rerouting in the computational graph—the very change distribution shift induces.
To close this gap, the authors propose the Circuit Alignment Score (CAS), which compares class-specific circuits across domains using graph kernels and decomposes the score into same-class coherence and cross-class confusion. They cast CAS as a Lebesgue integral over the domain distribution and prove that its Monte Carlo estimate recovers the ground-truth ranking of learners by OOD accuracy, with pairwise inversion error vanishing at rate O(1/M), where M is the number of sampled domains.
Across 48 learners on PACS, CAS attains 0.88 rank correlation with OOD accuracy, compared with 0.58 for CKA, 0.23 for SVCCA, and 0.14 for RSA, with similar trends on other benchmarks and even against data-dependent methods. The authors describe it as the first provably consistent predictor of distributional robustness that requires neither target-domain data nor labels, and the code is available at https://github.com/ayanban011/ACE.