SPIN Unprocessed July 28, 2026 ai_technology research
IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages
View original on arxiv.orgOverview
arXiv:2607.23242v1 Announce Type: new Abstract: Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms. We present IndicTalk, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversation
SpinGraph analysis pending — check back after processing.
Ask AI about this story
Opens with the SpinGraph .md URL and structured context — one click, prompt included.
More from arXiv Computation and Language
View all →- BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis
- Co-Evolving Graph and Text Memory for Training-Free Multi-Hop Question Answering
- Beyond a Global Norm: Personalizing Toxicity Sensitivity in Language Models Without Retraining
- Interview with Kalle Lyytinen on "Implications of Theories of Language for Information Systems"
- LoRA for Gender-Inclusive Rewriting and Activation Steering for Counter-Narrative Generation
- Attention-Guided Layer Selection for Contrastive Decoding in Large Language Models
Markdown (.md) · JSON-LD schema (.json) · Machine-readable for AI & GEO