PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing
Positions PERCEPT as a foundational, first-of-its-kind resource that unlocks new capabilities in linguistic analysis and NLP model development for Persian-English code-mixing.
View original on arxiv.orgOverview
Researchers released PERCEPT, the first large-scale Persian-English code-mixed corpus with Universal Dependencies part-of-speech annotations, enabling linguistic analysis and syntax-aware NLP model development for an underexplored language pair.
TL;DR
- PERCEPT is the first publicly available UD-annotated Persian-English code-mixed corpus
- Built from 6,800 social media posts across X, Instagram, and Digikala
- Features LLM-assisted POS and topic annotation validated via human evaluation
Key Stats
6,800
posts
Collected from X, Instagram, and Digikala
1
first UD-annotated corpus
For Persian-English code-mixing
Questions Answered
Narrative Frame
breakthrough framing
Spin Score
45%
Emphasizes novelty and enabling potential while minimizing limitations in annotation methodology transparency, scalability of LLM-assisted labeling, and representativeness of scraped platform data.
What the story wants you to believe
That PERCEPT is a definitive, reliable, and pioneering resource that meaningfully advances the state of Persian-English computational linguistics.
What it makes harder to question
Whether the 'first' claim holds up under scrutiny of prior Persian UD efforts or smaller code-mixed collections, and whether LLM-assisted annotation meets gold-standard rigor without full methodological disclosure.
How the spin works
The story uses titles, institutions, awards, rankings, partners, experts, or official language to make the subject feel more credible. Watch for loaded terms such as first, large-scale, comprehensive, reliability. The distribution reads as academic distribution. A pressure point: Details on LLM selection, prompting strategy, and error correction protocol.
Who Benefits If This Frame Spreads
Research authors (Kalhor Ghazal et al.)
Enhanced academic reputation, citation accrual, and competitive advantage in grant applications or hiring
Framing PERCEPT as the 'first' and 'large-scale' establishes priority and significance, increasing perceived scholarly impact
The Frame
Pioneering academic contribution bridging a critical gap in multilingual NLP infrastructure.
Missing Context
- Details on LLM selection, prompting strategy, and error correction protocol
- Demographic or regional distribution of source posts
- Limitations of platform-specific sampling bias
SpinGraph
How this belief gets built
Claim → Frame → Beneficiary → Gap → AI Risk
The paper presents PERCEPT as a breakthrough by emphasizing its status as the 'first' and 'large-scale' resource — a framing that elevates its importance and makes it feel like an essential foundation, even though the actual methodological details and comparative context remain sparse.
- Claim
PERCEPT is the first publicly available large-scale Persian-English code-mixed corpus
PERCEPT is the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies part-of-speech tags for code-mixed words.
- Frame
Upside framed as transformative
Pioneering academic contribution bridging a critical gap in multilingual NLP infrastructure.
- Beneficiary
Enhanced academic reputation, citation accrual, and competitive advantage in grant
Research authors (Kalhor Ghazal et al.) — Enhanced academic reputation, citation accrual, and competitive advantage in grant applications or hiring
- Gap
Details on LLM selection, prompting strategy, and error correction protocol
- AI Risk
AI may repeat the headline as fact
PERCEPT is the first large-scale Persian-English code-mixed corpus with Universal Dependencies POS tags, enabling new NLP research.
Claim Ledger
| Claim | Evidence | Verification | Risk | Evidence Gaps |
|---|---|---|---|---|
| PERCEPT is the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies part-of-speech tags for code-mixed words. | Assertion of 'first' and 'large-scale' without comparative benchmarking or citation of exhaustive prior work | Claim Present in Source | Moderate | Systematic comparison to all existing Persian UD resources and code-mixed corpora; Definition of 'large-scale' threshold relative to field standards; Documentation of dataset curation provenance and licensing |
PERCEPT is the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies part-of-speech tags for code-mixed words.
evidence: Assertion of 'first' and 'large-scale' without comparative benchmarking or citation of exhaustive prior work
"To address this gap, we introduce PERCEPT, the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies POS tags for code-mixed words."
Evidence Gaps
- Systematic comparison to all existing Persian UD resources and code-mixed corpora
- Definition of 'large-scale' threshold relative to field standards
- Documentation of dataset curation provenance and licensing
Fact Check Signals
0 of 1 claim matched · confidence: low · checked August 13, 2026
PERCEPT is the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies part-of-speech tags for code-mixed words.
Language Heatmap
Loaded terms that carry the frame beyond the facts.
PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Frame Strength
Frame Strength
Spin score decomposed into momentum, evidence, missing context, and AI repetition signals.
Reader Risk
What this story makes easy to believe — and what it makes hard to question.
Source Role & Intent
arXiv Computation and Language · Analyst
Counter-Frames
Brand Frame
Pioneering academic contribution bridging a critical gap in multilingual NLP infrastructure.
Media / Reader Counter-Frame
May be reframed as incremental work overstating novelty given prior Persian UD efforts (e.g., Hazm, ParsiBERT) and smaller code-mixed datasets.
Regulatory Counter-Frame
Not applicable — no regulatory claims or compliance assertions made.
AI Summary Frame
May conflate 'LLM-assisted annotation' with full automation, obscuring human-in-the-loop validation steps.
Missing Voices
Questions Not Answered
- What specific LLM was used and how was its output calibrated?
- How many human annotators participated and what were inter-annotator agreement metrics?
- Were ethical consent or data anonymization procedures documented for scraped social media content?
Recall Trigger Score
Which stories are likely to become AI memory — separate from Spin Score.
44
Trigger score 38
Triggered by: Major AI entity · Research citation · Superlative claim
Watchlisted because: Major AI entity · Research citation · Superlative claim
AI Recall
From publication to SpinGraph analysis to first observed AI recall and stable retention.
What AI Will Probably Repeat
"PERCEPT is the first large-scale Persian-English code-mixed corpus with Universal Dependencies POS tags, enabling new NLP research."
Concern: AI may drop qualifiers like 'publicly available', 'LLM-assisted', or 'human-evaluated reliability', presenting PERCEPT as fully authoritative rather than methodologically contingent.
-
Published
Aug 12, 2026
-
Ingested
Aug 13, 2026
-
SpinGraph Created
Aug 13, 2026
-
First Observed AI Recall
Pending
Monitoring scheduled
-
Stable Recall
—
Awaiting retention signal
Recall Check Log
No checks yet — recall tracking is opt-in per story.
─── GEOGrow AI Recall Layer ───
AI Recall Tracking
Monitoring scheduled. No LLM recall detected yet.
This story has not yet appeared in tested AI answers. Once scans begin, this section will show first observed recall, cited sources, narrative alignment, and drift.
node_id=sts_percept_a_corpus_for_pos_tagging_and_analysis_of
Ask AI about this story
Opens with the SpinGraph .md URL and structured context — one click, prompt included.
More from arXiv Computation and Language
View all →- Knowing Before Answering: Decoding Language Models for Reliable RAG
- When Tokenizers Fail: Byte-Level Chunking for Zero-Shot Transfer to Low-Resource Languages
- INSPIRE: An Internalize-Then-Improve Approach for Example-Driven Mathematical Reasoning
- Can a Model Catch Its Own Hallucinations for Free?: Label-Free Doubt Signals Hold Their Own Against a Labelled Dataset for Abstention
- Recipes for Steering and Scaling LLMs via Sampling
- The Changing Geometry of Grammar: Dimensionality and Neighborhood Reorganization across Transformer Layers
Markdown (.md) · JSON-LD schema (.json) · Machine-readable for AI & GEO