On 22 July 2026, at the Bible Hall in Aizawl, Chief Minister Lalduhoma used a book launch to make a point that had little to do with books.
The occasion was the release of Zawhzawzo Val, a Festschrift honouring Prof. Laltluangliana Khiangte — the Padma Shri–winning scholar, playwright and poet who retired from Mizoram University's Mizo Department in June — alongside two of Khiangte's own works, Lalzika and Vanthangi and Teaching of Mizo Language in Secondary School. Mizoram University Vice Chancellor Prof. Dibakar Chandra Deka also released Pasaltha Khuangchera, a Bengali retelling of the legendary Mizo warrior's story.
Amid the tributes, the Chief Minister made a technical argument. Mizo literature should be translated into English, Hindi, Bengali and other major languages and put online. Mizo words should be published alongside their English and Hindi equivalents. Do this, he said, and AI-based language tools get better — which matters now that Mizo sits inside Google Translate and, by extension, inside the AI-powered platforms built on top of that infrastructure.
It sounds like cultural policy. It's closer to infrastructure planning. And there is a specific, underappreciated technical reason he's right.
The zero-shot problem
When Google added Mizo to Translate in May 2022 — one of 24 new languages, taking the total to 133 — it was celebrated in Mizoram as arrival. It was. But the way Mizo was added is the part that rarely gets discussed.
Mizo was among the first batch of languages Google onboarded using Zero-Shot Machine Translation. In a conventional translation system, the model learns from parallel text: millions of sentences in one language matched to their equivalents in another. Zero-shot works differently. The model is trained on a large pool of data-rich languages, then shown monolingual text in the new language — Mizo text with no English counterpart attached — and asked to generalise. It learns to translate Mizo without ever seeing an example of a Mizo translation.
It's a genuinely impressive piece of engineering, and Google was upfront that it isn't perfect, promising continued improvement to those models. But understand what it means in practice. Mizo's presence in the world's most-used translation tool rests on inference, not evidence. The system is making educated guesses built on structural patterns borrowed from other languages.
This is why the output behaves the way Mizo speakers describe: broadly comprehensible, frequently stilted, and unreliable precisely where it matters most — idiom, register, tone, and anything domain-specific. Ask it to handle church vocabulary, legal phrasing, agricultural terminology or clinical instructions and the seams show immediately.
Zero-shot got Mizo onto the map. Only parallel data will make it accurate.
Why parallel text is the actual bottleneck
Mizo has somewhere between 800,000 and a million speakers. English has billions of words of indexed, high-quality text online. Hindi has hundreds of millions. Mizo has a fraction of that, and the fraction that exists is overwhelmingly monolingual — Mizo blogs, Mizo Facebook posts, Mizo news sites, Mizo sermons. Valuable, but it feeds the weaker training method.
What's scarce is the aligned pair: a Mizo sentence sitting next to its verified English or Hindi equivalent. This is the single highest-value artefact for machine translation, and it's the one thing that has to be produced deliberately. Nobody generates it by accident.
Three characteristics of Mizo make the shortage bite harder than raw speaker numbers suggest:
Tone. Mizo is tonal — pitch distinguishes meaning. Standard Latin-script orthography under-marks this, so models trained on undifferentiated text conflate words that speakers hear as distinct.
Kuki-Chin structure. Mizo belongs to the Kuki-Chin branch of Tibeto-Burman. Its grammar shares little with Hindi or Bengali, so the cross-lingual transfer that zero-shot depends on has less to borrow from.
Domain gaps. General conversational Mizo is reasonably represented online. Administrative, medical, legal and technical Mizo is not — and those are the registers where mistranslation carries real cost.
Every parallel sentence published closes a bit of this gap, and it compounds: better tools encourage more people to write and publish in Mizo, which produces more data, which improves the tools.
What this actually does to your phone
Most Mizos meet technology through a smartphone, so the payoff is concrete.
Keyboards and predictive text. Swipe typing, autocorrect and next-word prediction are trained on large text corpora. Thin corpora produce the experience Mizo users have now — aggressive autocorrect that mangles correct Mizo into approximate English, missing diacritics, prediction that gives up after two words. Larger corpora fix this at the system level, not app by app.
Voice input. Speech recognition needs audio paired with accurate transcripts. Mizoram produces an enormous volume of recorded speech — sermons, choir performances, YouTube channels, radio, assembly proceedings — and almost none of it is captioned. Captioning existing Mizo video is probably the highest-leverage, lowest-cost intervention available right now, and it improves both speech recognition and the text corpus simultaneously.
Local AI apps. Hriatna, marketed as the first Mizo AI chatbot and available on both Play Store and App Store, has added voice interaction and hooks into the central government's Bhashini initiative. LushAITech built the LushAI Healthy Lunglei app, which handles OPD bookings and answers health queries in Mizo, and has met the Mizo Language Development Board about an AI language partnership. These teams are working against the same data ceiling as everyone else. Raise it and their products improve without a line of new code.
Access. For elderly users, rural residents and anyone not fluent in English or Hindi, this is the difference between using a banking app and asking a relative to do it. Government service portals, health information, school platforms — all of it becomes navigable when the translation layer stops guessing.
The Eighth Schedule angle
This isn't happening in isolation. In March 2026, the Mizoram Legislative Assembly unanimously adopted a resolution — moved by Education Minister Vanlalthlana and initiated by the Mizo Language Development Board after statewide consultation — seeking Mizo's inclusion in the Eighth Schedule of the Constitution.
Mizo has been Mizoram's official language since 1974. The Eighth Schedule currently lists 22 languages, and Mizo is one of roughly 38 competing for entry. Inclusion would let candidates sit central service examinations in Mizo, allow MPs to speak it in Parliament, and unlock central funding for language development — including, relevantly, the corpus-building work that AI systems depend on.
The digital argument and the constitutional argument reinforce each other. A language with demonstrable digital infrastructure makes a stronger case for recognition; recognition brings resources that build more infrastructure.
Who actually has to do the work
Generic calls to "create more content" go nowhere. Specific asks land. In Mizoram's case, the institutions that hold the most valuable untapped material are identifiable:
- The churches. The Synod and other denominations hold decades of Mizo publications, many with existing English versions — hymnals, commentaries, periodicals. Bilingual religious material is a ready-made parallel corpus, and Mizoram's churches are among the best-organised archival institutions in the state.
- Mizo Academy of Letters and Mizoram University. Sixty years of literary output, much of it already translated for academic purposes. Digitising and openly licensing this is a decision, not a research project.
- The state government. Every bilingual gazette notification, circular and scheme document is a parallel dataset the state already owns and already paid for. Publishing these in machine-readable form costs almost nothing.
- Content creators. Captioning your own Mizo YouTube videos in both Mizo and English produces exactly the aligned audio-text data that speech models need.
- YMA and community organisations. The distribution network to run systematic collection at village level already exists.
The licensing detail matters more than it sounds. Content behind copyright, sitting in a PDF, or locked to a platform contributes far less than openly licensed, machine-readable text. "Online" and "available for training" are not the same thing.
The part nobody's discussing
Two risks deserve naming before this becomes policy.
The first is quality control. Models learn whatever they're fed, including errors. If the corpus fills with hasty machine-translated Mizo — output from the current zero-shot system, recycled back in as training data — the result is a feedback loop that entrenches mistakes and calls them standard. Any serious corpus effort needs human verification of the parallel pairs, which means paying qualified Mizo translators. That's a budget line, not a volunteer drive.
The second is standardisation. Building a corpus means making choices about which Mizo gets encoded — which orthographic conventions, which regional forms, whose register. Those choices will shape how a generation of speakers sees their own language reflected back by their devices. They're worth making consciously and transparently rather than by default, according to whoever digitises fastest.
The bottom line
Languages don't die in the digital age because their speakers stop caring. They die because the systems mediating daily life — search, keyboards, voice assistants, chatbots — never learn to handle them, and speakers quietly switch to a language the machine understands.
Mizo got a head start with zero-shot. That head start doesn't renew itself. The work of turning it into genuine fluency is unglamorous, distributed, and mostly consists of publishing things that already exist in a format machines can read.
The Chief Minister asked for more Mizo online. The more precise version of the ask: more Mizo online, paired with its translation, openly licensed, and checked by someone who knows the language.
Sources: PTI/The Print, India Today NE (22 July 2026); Google Translate blog, "Google Translate learns 24 new languages" (May 2022); News on AIR (9 March 2026); Google Play and App Store listings; EastMojo (30 January 2026).