वाणी வாணி India's Open Language Corpus

Ancient texts. Living speech.
Arriving soon.

VĀṆĪ is an open corpus for Indian language research — classical Sanskrit and Tamil literature alongside living, code-switched speech data. We're preparing the first release, drawn from the Vedas to the Sangam poets to real conversations recorded with consent on Vayu.

✓ You're on the list — we'll write when it's ready.

No spam, one email at launch. Or write to us directly at corpus@vanic.org.

What's coming
Living SpeechIn progress

Real-world Indian language audio and transcripts, captured with consent via Vayu. English, Hindi, Tamil, Hinglish.

Ancient TextsIn progress

60+ classical Sanskrit and Tamil texts, annotated in a unified NLP-ready format — from the Rigveda to the Tirukkuṟaḷ.