New multilingual AI stack from IIT Madras aims to personalise education across Indian languages
Developed in partnership with the AI4Bharat research lab at IIT Madras, the models cover speech recognition, speech generation, machine translation and optical character recognition (OCR)
Developed in partnership with the AI4Bharat research lab at IIT Madras, the models cover speech recognition, speech generation, machine translation and optical character recognition (OCR).
Developed in partnership with the AI4Bharat research lab at IIT Madras, the models cover speech recognition, speech generation, machine translation and optical character recognition (OCR).
Developed in partnership with the AI4Bharat research lab at IIT Madras, the models cover speech recognition, speech generation, machine translation and optical character recognition (OCR).
The IIT Madras-incubated Bodhan AI on Friday launched four foundational AI models for Indian languages in partnership with AI4Bharat. Bodhan AI, a centre for excellence in artificial intelligence for education is building a common technology layer for education-focused AI applications in the country.
Developed in partnership with the AI4Bharat research lab at IIT Madras, the models cover speech recognition, speech generation, machine translation and optical character recognition (OCR). Bodhan AI is making them available as open-weight models as well as hosted APIs on sovereign digital infrastructure, allowing startups, edutech companies, researchers, and government institutions to build and customise applications on top of them.
It will be made available to the wider technology and education ecosystem as Digital Public Goods towards building shared AI infrastructure that others can use, adapt and build upon.
Bodhan AI’s model strategy forms part of the broader Bharat EduAI Stack, envisioned as a sovereign digital public infrastructure for education. In partnership with AI4Bharat, which brings expertise in Indian languages AI, multilingual modelling and datasets, Bodhan AI is building the foundational AI infrastructure for India's multilingual education ecosystem.
These voice and vision models are an early flagship output, designed to enable development of AI solutions across Indian languages.
“Bodhan AI aims to build with the ecosystem, not compete with it,” said Prof. Mitesh Khapra, Principal Investigator at Bodhan AI.
Bodhan AI is releasing a suite of AI models spanning four core capabilities: speech recognition, speech generation, machine translation and OCR.
“Bodhan AI aims to build with the ecosystem, not compete with it. We built these voice and vision models, and made them accessible as Digital Public Goods on a Digital Public Infrastructure, so that efforts across the country don't remain fragmented. Instead, there is one common layer that can power all Edu AI in India, without every institution having to duplicate the same work," added Prof. Khapra.
The models are trained and optimised using NVIDIA Nemotron open models, and libraries, including the NVIDIA NeMo framework for automatic speech recognition (ASR), machine translation (MT) and optical character recognition (OCR). Bodhan post-trained NVIDIA Nemotron 3.5 ASR to support Indian languages, including regional dialects and accents. The models are served using NVIDIA® TensorRT™ LLM and vLLM inference micro services.
Both NVIDIA and Bodhan AI are collaborating on datasets, training recipes and evaluations to support future foundational models for Indian languages.
AI4Bharat is developing open-source datasets, tools, models and applications for Indian languages. It is dedicated to advancing AI technology for Indian languages through open-source contributions. AI4Bharat’s work is recognised globally, with publications in top-tier conferences and deployments in real-world use cases, making a significant impact across academia, industry, and government sectors.
For Bodhan AI, the ultimate purpose of these foundational models is to make AI more accessible to teachers and students, particularly those who interact with technology in Indian languages. By opening access to foundational capabilities, it seeks to enable edtech companies, startups, researchers, universities, technology companies and government partners to develop their own applications for Indian users.
Prof. Balaraman Ravindran, Head, Wadhwani School of Data Science and AI (WSAI), IIT Madras, says:
“Building AI for India requires going beyond adapting existing models to Indian languages. It requires developing foundational capabilities that understand the richness and diversity of our languages, contexts and educational needs. The models being released by Bodhan AI represent an important foundation for this effort. By making these capabilities available through open-weight models and APIs, we are enabling researchers, startups, technology companies and education innovators to experiment, adapt and build applications that are relevant to India at scale.”
The launch is designed not simply as a release of four models, but as an invitation to the ecosystem to build with Bodhan AI. The goal is to make AI accessible to at affordable prices mindful of India’s reality. Government partners and public institutions can potentially deploy multilingual AI capabilities for education and other public-service applications using sovereign infrastructure.
Edutech companies can integrate Indian-language AI capabilities into their products through APIs without having to develop foundational models from scratch.
Startups and technology companies can access open-weight models, adapt them and fine-tune them for new applications and user groups.
Researchers and universities can use the models as a foundation for further research, experimentation and development in Indian-language AI.
ASR can allow a student to speak to an AI tutor in their preferred language rather than having to type in English. TTS can enable an AI system to respond naturally in an Indian language. OCR can allow educational AI systems to understand textbooks, worksheets, documents and handwritten answers, while Bodhan-Translate can help educational content move across Indian languages.
Together, these capabilities can form the underlying technology for applications supporting learning, tutoring, practice and assessment across India’s education system. The approach is based on the fundamental premise that AI should understand India’s languages, rather than requiring India to adapt to its language.
Bodhan AI is also launching ‘Student TutorBot’ and ‘Teacher Assistant Bot’, bringing multilingual AI-powered learning and teaching support directly to students and teachers. The models will also form part of the technology foundation powering Bodhan AI’s own education applications, including Student Tutor Bot and Teacher Assistant Bot, while remaining available for others to build upon.
This creates a model in which the same foundational capabilities can be reused across multiple applications and institutions, reducing duplication and enabling a broader ecosystem of innovation.
Student Tutor Bot is an AI-powered learning companion for students in Classes 6-12, designed around NCERT and SCERT curricula to provide structured, personalised learning support beyond the classroom. Students can ask questions through text or voice across 22 Indian languages, with Tutor Bot using their textbook content to provide explanations, examples and assessments.
Student Tutor Bot will guide students step-by-step through problems, using interactive tools and a digital canvas for diagrams, calculations and working. Students can also be assessed on the same curriculum content, with their demonstrated understanding informing subsequent learning support.
The Teacher Assistant Bot is a pan-India AI workspace designed to help teachers plan, create, assign, assess and refine classroom work more efficiently while keeping the teacher firmly in control.
Teachers can generate lesson plans, worksheets, quizzes, homework and revision sheets by specifying parameters such as grade, subject, topic, duration, difficulty and assessment requirements, and can also upload student work for evaluation with specific marking criteria.
Most importantly, every AI-generated output is treated as a starting point for the teacher to review, edit, regenerate or discard, rather than an automatic classroom decision. The Teacher Assistant Bot is therefore positioned not simply as an AI content generator, but as a working environment that helps teachers turn their teaching intent into classroom-ready work, faster.
Data privacy and responsible deployment are foundational to the initiative. Bodhan AI’s architecture will incorporate data anonymisation protocols and comply with applicable national education data frameworks, with the hosted API infrastructure designed around a sovereign deployment approach.
This will provide institutions and ecosystem partners greater control over how AI capabilities are deployed while supporting India’s requirements around data governance, privacy and security.
The objective is to create infrastructure that is not only capable and affordable, but also suited to India's regulatory and institutional environment—AI for India, governed in India.