SPIN Unprocessed July 28, 2026 ai_technology research
BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis
View original on arxiv.orgOverview
arXiv:2607.23319v1 Announce Type: new Abstract: Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages. Sanskrit, Tamil, and other classical Indic languages exhibit agglutinative morphology, productive sandhi (phonological fusion at word boundaries), and domain-specific vocabularies absent from general-purpose training data. Th
SpinGraph analysis pending — check back after processing.
Ask AI about this story
Opens with the SpinGraph .md URL and structured context — one click, prompt included.
More from arXiv Computation and Language
View all →- Co-Evolving Graph and Text Memory for Training-Free Multi-Hop Question Answering
- IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages
- Beyond a Global Norm: Personalizing Toxicity Sensitivity in Language Models Without Retraining
- Interview with Kalle Lyytinen on "Implications of Theories of Language for Information Systems"
- LoRA for Gender-Inclusive Rewriting and Activation Steering for Counter-Narrative Generation
- Attention-Guided Layer Selection for Contrastive Decoding in Large Language Models
Markdown (.md) · JSON-LD schema (.json) · Machine-readable for AI & GEO