PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing
Social media is a major venue for multilingual communication, yet Persian-English code-mixing has been underexplored. Existing Persian resources lack Universal Dependencies (UD) POS annotations for code-mixed words, limiting linguistic analysis and syntax-aware NLP models. PERCEPT addresses this gap as the first publicly available large-scale Persian-English code-mixed corpus with UD POS tags for code-mixed words, comprising 6,800 posts collected from X, Instagram, and Digikala.
The authors introduce an LLM-assisted annotation framework that automatically assigns POS tags and document-level topics. Human evaluation shows high agreement between automated and gold annotations, confirming reliability.
Using PERCEPT, the first comprehensive linguistic analysis of Persian-English code-mixing across multiple platforms reveals that nouns are the predominant category for code-mixed words, while distributions of other POS categories vary across platforms. Positional distribution of code-mixed words is remarkably consistent across platforms, whereas the triggering effect is substantially more pronounced in Digikala. The dataset is publicly available at https://github.com/kalhorghazal/PERCEPT.