seq2cause: One Autoregressive Backbone, Four Causal Discovery Tasks in Event Sequences
Complex systems such as vehicles, patients, or genomes emit discrete event sequences, and the operative question about them is causal rather than predictive: which events cause which other events, and which events cause higher-level outcomes such as failures or diseases. The authors decompose this question along two axes — dependency type (event→event versus event→outcome) and causal scope (single sequence versus population) — which yields four structurally distinct regimes, each with its own identifiability conditions.
No existing method addresses more than one of these regimes, because existing approaches all assume multi-stream structure with a low vocabulary and none scales beyond a few hundred event types.
Seq2Cause resolves all four regimes through a single shared primitive: a pretrained autoregressive model repurposed as an amortized conditional independence testing engine, requiring no task-specific retraining. The paper also establishes a prediction–causality duality, in which the model's excess cross-entropy simultaneously bounds causal identification error across all four regimes, so that every improvement in next-token prediction tightens the causal guarantees for free.
Empirically, the method is evaluated on nonlinear SCMs with vocabularies up to 8,000 types and on real-world vehicle diagnostic logs containing 29K event types and 474 failure outcomes. Seq2Cause is reported as the first method to populate all four regimes at scale with a single frozen backbone, while existing methods are either inapplicable, inaccurate, or computationally intractable in this setting.