Research arXiv cs.AI

The Ignition Index: Measuring Global Workspace Dynamics in Language Models

Global Workspace Theoryinterpretabilitylinear probinglanguage models

The Ignition Index (I) operationalizes Global Workspace Theory's all-or-none ignition prediction by fitting a four-parameter sigmoid to per-layer linear probe accuracy as a function of input signal strength. The key extracted parameter, beta-hat, indicates whether a model shows abrupt (ignition-like) or graded transitions. The metric was validated across 11 transformer language models, suggesting variability in how different architectures and scales handle global workspace dynamics. This provides a concrete, quantitative tool for testing cognitive theories in neural networks, potentially guiding interpretability research and model design. Future work might explore correlations with model size, training data, or task performance.

Read original →

← Back to home