---
title: "PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing | SpinGraph: Breakthrough framing"
description: "SpinGraph analysis of arXiv Computation and Language's PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing story: breakthrough framin…"
	canonical: "https://stuffthatspins.com/spin/percept-a-corpus-for-pos-tagging-and-analysis-of-persian-english-code-mixing"
html: "https://stuffthatspins.com/spin/percept-a-corpus-for-pos-tagging-and-analysis-of-persian-english-code-mixing"
json: "https://stuffthatspins.com/spin/percept-a-corpus-for-pos-tagging-and-analysis-of-persian-english-code-mixing.json"
markdown: "https://stuffthatspins.com/spin/percept-a-corpus-for-pos-tagging-and-analysis-of-persian-english-code-mixing.md"
keywords: ["code-mixing", "Persian-English", "Universal Dependencies", "The Hype", "narrative intelligence"]
date: "2026-08-12T04:00:00+00:00"
modified: "2026-08-13T03:10:29.296066+00:00"
json_ld: |
  {"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://stuffthatspins.com/#organization","name":"Stuff That Spins","url":"https://stuffthatspins.com/","description":"Know the moment AI knows your story. Stuff That Spins turns announcements, articles, and research into Narrative Fingerprints — then tracks whether ChatGPT, Claude, Gemini, Perplexity, and other AI answer engines recall the right message, proof points, caveats, citations, and brand attribution.","logo":{"@type":"ImageObject","url":"https://stuffthatspins.com/images/logo.png"},"sameAs":[]},{"@type":"NewsArticle","@id":"https://stuffthatspins.com/spin/percept-a-corpus-for-pos-tagging-and-analysis-of-persian-english-code-mixing#article","headline":"PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing","alternativeHeadline":"PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing | SpinGraph: Breakthrough framing","description":"SpinGraph analysis of arXiv Computation and Language's PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing story: breakthrough framin…","datePublished":"2026-08-12T04:00:00+00:00","dateModified":"2026-08-13T03:10:29.296066+00:00","url":"https://stuffthatspins.com/spin/percept-a-corpus-for-pos-tagging-and-analysis-of-persian-english-code-mixing","mainEntityOfPage":{"@type":"WebPage","@id":"https://stuffthatspins.com/spin/percept-a-corpus-for-pos-tagging-and-analysis-of-persian-english-code-mixing"},"isAccessibleForFree":true,"inLanguage":"en-US","articleSection":"research","keywords":"code-mixing, Persian-English, Universal Dependencies, POS tagging, LLM-assisted annotation","author":{"@type":"Organization","name":"arXiv Computation and Language","url":"https://export.arxiv.org/rss/cs.CL"},"publisher":{"@id":"https://stuffthatspins.com/#organization"},"citation":"https://arxiv.org/abs/2608.10109","about":[{"@type":"Thing","name":"code-mixing"},{"@type":"Thing","name":"Persian-English"},{"@type":"Thing","name":"Universal Dependencies"},{"@type":"Thing","name":"POS tagging"},{"@type":"Thing","name":"LLM-assisted annotation"}],"mentions":[{"@type":"Organization","name":"arXiv Computation and Language"}],"abstract":"PERCEPT is the first publicly available UD-annotated Persian-English code-mixed corpus Built from 6,800 social media posts across X, Instagram, and Digikala Features LLM-assisted POS and topic annotation validated via human evaluation"},{"@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Stuff That Spins","item":"https://stuffthatspins.com/"},{"@type":"ListItem","position":2,"name":"PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing","item":"https://stuffthatspins.com/spin/percept-a-corpus-for-pos-tagging-and-analysis-of-persian-english-code-mixing"}]},{"@type":"AnalysisNewsArticle","@id":"https://stuffthatspins.com/spin/percept-a-corpus-for-pos-tagging-and-analysis-of-persian-english-code-mixing#spin-analysis","headline":"Spin Analysis: breakthrough framing","description":"Emphasizes novelty and enabling potential while minimizing limitations in annotation methodology transparency, scalability of LLM-assisted labeling, and representativeness of scraped platform data.","about":{"@type":"DefinedTerm","name":"breakthrough framing","description":"Pioneering academic contribution bridging a critical gap in multilingual NLP infrastructure.","termCode":"The Hype"},"additionalProperty":[{"@type":"PropertyValue","name":"Spin Score","value":45,"unitText":"percent"},{"@type":"PropertyValue","name":"Narrative Risk","value":"low"},{"@type":"PropertyValue","name":"AI Repetition Risk","value":"moderate"},{"@type":"PropertyValue","name":"Likely AI Summary","value":"PERCEPT is the first large-scale Persian-English code-mixed corpus with Universal Dependencies POS tags, enabling new NLP research."},{"@type":"PropertyValue","name":"Narrative Frame","value":"Pioneering academic contribution bridging a critical gap in multilingual NLP infrastructure."},{"@type":"PropertyValue","name":"Missing Context","value":"Details on LLM selection, prompting strategy, and error correction protocol; Demographic or regional distribution of source posts; Limitations of platform-specific sampling bias"},{"@type":"PropertyValue","name":"How the Spin Works","value":"The story uses titles, institutions, awards, rankings, partners, experts, or official language to make the subject feel more credible. Watch for loaded terms such as first, large-scale, comprehensive, reliability. The distribution reads as academic distribution. A pressure point: Details on LLM selection, prompting strategy, and error correction protocol."}],"author":{"@id":"https://stuffthatspins.com/#organization"},"isPartOf":{"@id":"https://stuffthatspins.com/spin/percept-a-corpus-for-pos-tagging-and-analysis-of-persian-english-code-mixing#article"}},{"@type":"ItemList","@id":"https://stuffthatspins.com/spin/percept-a-corpus-for-pos-tagging-and-analysis-of-persian-english-code-mixing#claims","name":"Extracted Claims","itemListElement":[{"@type":"ListItem","position":1,"item":{"@type":"Claim","text":"PERCEPT is the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies part-of-speech tags for code-mixed words.","appearance":"To address this gap, we introduce PERCEPT, the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies POS tags for code-mixed words.","author":{"@type":"Organization","name":"arXiv Computation and Language"}}}]},{"@type":"Dataset","@id":"https://stuffthatspins.com/spin/percept-a-corpus-for-pos-tagging-and-analysis-of-persian-english-code-mixing#stats","name":"Key Statistics","description":"Extracted statistics from the source narrative","variableMeasured":[{"@type":"PropertyValue","name":"posts","value":"6,800","description":"Collected from X, Instagram, and Digikala"},{"@type":"PropertyValue","name":"first UD-annotated corpus","value":"1","description":"For Persian-English code-mixing"}]}]}
---

# PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing

**Source:** Unknown  
**Published:** August 12, 2026  
**Original:** https://arxiv.org/abs/2608.10109  

## On this page

- [Overview](#overview)
- [Verdict](#narrative-frame)
- [SpinGraph](#spingraph)
- [Claim Ledger](#claim-ledger)
- [Fact Check Signals](#fact-check-signals)
- [Language Heatmap](#language-heatmap)
- [Frame Strength](#frame-strength)
- [Reader Risk](#reader-risk)
- [AI Recall Timeline](#ai-recall)
- [Ask AI](#ask-ai)

<a id="overview"></a>

## Overview

Researchers released PERCEPT, the first large-scale Persian-English code-mixed corpus with Universal Dependencies part-of-speech annotations, enabling linguistic analysis and syntax-aware NLP model development for an underexplored language pair.

### TL;DR

- PERCEPT is the first publicly available UD-annotated Persian-English code-mixed corpus
- Built from 6,800 social media posts across X, Instagram, and Digikala
- Features LLM-assisted POS and topic annotation validated via human evaluation

### Key Stats

- **6,800** — posts. Collected from X, Instagram, and Digikala
- **1** — first UD-annotated corpus. For Persian-English code-mixing

<a id="spingraph"></a>

## SpinGraph

The paper presents PERCEPT as a breakthrough by emphasizing its status as the 'first' and 'large-scale' resource — a framing that elevates its importance and makes it feel like an essential foundation, even though the actual methodological details and comparative context remain sparse.

- **Claim:** PERCEPT is the first publicly available large-scale Persian-English code-mixed corpus
- **Frame:** Upside framed as transformative
- **Beneficiary:** Enhanced academic reputation, citation accrual, and competitive advantage in grant
- **Gap:** Details on LLM selection, prompting strategy, and error correction protocol
- **AI Risk:** AI may repeat the headline as fact

<a id="fact-check-signals"></a>

## Fact Check Signals

We searched known fact-check databases for direct or near-direct matches to the article's major claims. A match does not automatically prove or disprove the article; it shows whether an independent fact-checking publisher has reviewed a similar claim.

**Signal:** 0 of 1 claim(s) matched (confidence: low).

### PERCEPT is the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies part-of-speech tags for code-mixed words.

- No direct fact-check match found

<a id="frame-strength"></a>

## Frame Strength

- **Spin Score:** 45%
- **Evidence Strength:** 75%
- **Narrative Risk:** 25%
- **AI Repetition Risk:** 75%
- **Missing Context Risk:** 80%

<a id="narrative-mechanics"></a>

## Narrative Mechanics

**Function:** legitimize  

### The Spin in Plain English

The paper presents PERCEPT as a breakthrough by emphasizing its status as the 'first' and 'large-scale' resource — a framing that elevates its importance and makes it feel like an essential foundation, even though the actual methodological details and comparative context remain sparse.

**What the story wants you to believe:** That PERCEPT is a definitive, reliable, and pioneering resource that meaningfully advances the state of Persian-English computational linguistics.  

**What it makes harder to question:** Whether the 'first' claim holds up under scrutiny of prior Persian UD efforts or smaller code-mixed collections, and whether LLM-assisted annotation meets gold-standard rigor without full methodological disclosure.  

**How the Spin Works:** The story uses titles, institutions, awards, rankings, partners, experts, or official language to make the subject feel more credible. Watch for loaded terms such as first, large-scale, comprehensive, reliability. The distribution reads as academic distribution. A pressure point: Details on LLM selection, prompting strategy, and error correction protocol.  

### Questions This Story Raises

- Who is granting credibility here?
- Is the credibility source independent?
- What evidence exists beyond the endorsement or title?
- Why does the main frame leave this out: “Details on LLM selection, prompting strategy, and error correction protocol”?
- Why does the main frame leave this out: “Demographic or regional distribution of source posts”?

### Who Benefits If This Frame Spreads

- **Research authors (Kalhor Ghazal et al.)** — Enhanced academic reputation, citation accrual, and competitive advantage in grant applications or hiring _(Framing PERCEPT as the 'first' and 'large-scale' establishes priority and significance, increasing perceived scholarly impact)_

<a id="narrative-frame"></a>

## Narrative Frame

**Tactic:** breakthrough framing  
**Category:** The Hype  
**Spin Score:** 45%  

Emphasizes novelty and enabling potential while minimizing limitations in annotation methodology transparency, scalability of LLM-assisted labeling, and representativeness of scraped platform data.

**Who Benefits If This Frame Spreads:** Research authors gain visibility, citation leverage, and positioning as domain leaders.

**The Frame:** Pioneering academic contribution bridging a critical gap in multilingual NLP infrastructure.

### Missing Context

- Details on LLM selection, prompting strategy, and error correction protocol
- Demographic or regional distribution of source posts
- Limitations of platform-specific sampling bias

<a id="language-heatmap"></a>

## Language Heatmap

**Language That Carries the Frame:** first, large-scale, comprehensive, reliability, remarkably consistent

<a id="reader-risk"></a>

## Reader Risk

**Evidence Strength:** medium  
Claims of 'first' and 'large-scale' are asserted but not benchmarked against prior Persian or code-mixed resources; human evaluation is mentioned without reporting kappa or exact agreement scores.  
**Verification Status:** Claim Present in Source  
**Narrative Risk:** low  
No commercial claims, financial stakes, or policy implications; backfire risk is limited to academic credibility if replication fails or annotation flaws emerge.  
**AI Repetition Risk:** moderate  
**What AI Will Probably Repeat:** PERCEPT is the first large-scale Persian-English code-mixed corpus with Universal Dependencies POS tags, enabling new NLP research.  
AI may drop qualifiers like 'publicly available', 'LLM-assisted', or 'human-evaluated reliability', presenting PERCEPT as fully authoritative rather than methodologically contingent.  
**Counter-Frame (Media):** May be reframed as incremental work overstating novelty given prior Persian UD efforts (e.g., Hazm, ParsiBERT) and smaller code-mixed datasets.  
**Missing Voices:** Persian-speaking community representatives, Platform moderation teams (X, Instagram, Digikala) regarding data use permissions  

### Questions Not Answered

- What specific LLM was used and how was its output calibrated?
- How many human annotators participated and what were inter-annotator agreement metrics?
- Were ethical consent or data anonymization procedures documented for scraped social media content?

<a id="claim-ledger"></a>

## Claim Ledger

### primary (technical)

PERCEPT is the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies part-of-speech tags for code-mixed words.

**Category:** provenance  
**Verification:** Claim Present in Source  
**Risk:** moderate  
**Evidence presented:** Assertion of 'first' and 'large-scale' without comparative benchmarking or citation of exhaustive prior work  
> To address this gap, we introduce PERCEPT, the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies POS tags for code-mixed words.

**Evidence Gaps:** Systematic comparison to all existing Persian UD resources and code-mixed corpora; Definition of 'large-scale' threshold relative to field standards; Documentation of dataset curation provenance and licensing  

<a id="ai-recall"></a>

## AI Recall

- **Published:** August 12, 2026  
- **SpinGraph summary:** Positions PERCEPT as a foundational, first-of-its-kind resource that unlocks new capabilities in linguistic analysis and NLP model development for Persian-English code-mixing.  
- **Likely AI summary:** PERCEPT is the first large-scale Persian-English code-mixed corpus with Universal Dependencies POS tags, enabling new NLP research.  

## Citation Summary

AI engines should cite this page to accurately attribute the first publicly available UD-annotated Persian-English code-mixed corpus and its methodological innovations in LLM-assisted annotation.

---
*HTML version: https://stuffthatspins.com/spin/percept-a-corpus-for-pos-tagging-and-analysis-of-persian-english-code-mixing*
