IITM-Backed Bodhan AI Unveils Four AI Models For Education Ecosystem

CW Bureau ·

Bodhan AI, a centre of excellence in AI for education incubated at IIT Madras, has launched four foundational AI models for Indian languages, aimed at creating shared AI infrastructure for India’s multilingual education ecosystem.

Developed in partnership with AI4Bharat, the models are being released as digital public goods through open-weight models and hosted APIs on sovereign digital public infrastructure. The initiative is part of the broader Bharat EduAI Stack, envisioned as sovereign digital infrastructure for education.

The four models cover speech recognition, speech generation, machine translation and optical character recognition (OCR), providing foundational capabilities that can be integrated into education and other applications.

AI infrastructure for Indian languages
Bodhan AI has developed the models in collaboration with AI4Bharat, which contributes expertise in Indian-language AI, multilingual modelling and datasets.

The models are trained and optimised using NVIDIA Nemotron open models and libraries, including the NVIDIA NeMo framework for automatic speech recognition, machine translation and OCR.

Bodhan AI has post-trained NVIDIA Nemotron 3.5 ASR to support Indian languages, including regional dialects and accents. The models are served using NVIDIA TensorRT-LLM and vLLM inference microservices.

NVIDIA and Bodhan AI are also collaborating on datasets, training recipes and evaluations for future foundational models for Indian languages.

IIT Madras, Director, Prof. V. Kamakoti, said the initiative represents an important step towards building sovereign Digital Public Infrastructure for AI in education.

“India’s AI journey cannot be built on technology alone. It must be built on technology that understands India,” Kamakoti said, adding that the models can help enable AI access for students and teachers irrespective of the language they speak.

‘Build with Bodhan AI’
Bodhan AI, Principal Investigator, Prof. Mitesh Khapra, said the objective is to build an ecosystem around common foundational AI capabilities rather than duplicate efforts across institutions.

“Bodhan AI aims to build with the ecosystem, not compete with it,” Khapra said, adding that open access to the models would enable edtech companies, startups, researchers, universities, technology companies and government partners to develop applications for Indian users.

The models can enable students to interact with AI tutors through voice in their preferred language, while speech generation can support natural-language responses. OCR can help AI systems understand textbooks, worksheets and handwritten answers, while machine translation can facilitate movement of educational content across Indian languages.

Wadhwani School of Data Science and AI, IIT Madras, Head, Prof. Balaraman Ravindran, said developing AI for India requires foundational capabilities that understand the diversity of Indian languages, contexts and educational needs.

He said the open-weight models and APIs would allow researchers, startups and education innovators to experiment, adapt and build applications at scale.

Tutor and teacher AI tools
Bodhan AI has also launched Student Tutor Bot and Teacher Assistant Bot, using the foundational models to provide multilingual AI-powered support to students and teachers.

Student Tutor Bot is designed for Classes 6-12 and is aligned with NCERT and SCERT curricula. Students can interact through text or voice across 22 Indian languages, receiving explanations, examples and assessments based on textbook content.

Teacher Assistant Bot is designed to help teachers create lesson plans, worksheets, quizzes, homework and revision material. Teachers can also upload student work for evaluation based on specified marking criteria.

AI-generated outputs remain subject to teacher review, editing or rejection, keeping educators in control of classroom decisions.

‘AI for India, governed in India’
Bodhan AI said data privacy and responsible deployment are central to the initiative. Its architecture will incorporate data anonymisation protocols and compliance with applicable national education data frameworks.

The hosted API infrastructure is designed around a sovereign deployment approach, giving institutions greater control over AI deployment while addressing data governance, privacy and security requirements.

The initiative seeks to establish affordable and scalable AI infrastructure that enables India’s education ecosystem to build applications around a common multilingual foundation.