Main Page
Enhancing Scientific Wiki Articles in Indic Languages
Bridging the knowledge gap for the majority of India's population — making science, biodiversity, and innovation accessible in Hindi and Telugu.
Dr. Radhika Mamidi · Principal Investigator · LTRC, IIIT Hyderabad
The Mission
Democratizing Scientific Knowledge
The Knowledge Gap
India records 117 million daily Wikipedia views, yet the vast majority of scientific content remains locked in English — out of reach for most Indian-language readers. We are fixing this at scale.
Language Empowerment
Inspired by the success of native-language knowledge platforms elsewhere in the world, we are building a lasting Indian-language scientific corpus for generations to come.
Scientific Coverage
From biological sciences and chemistry to biodiversity, species taxonomy, and landmark inventions — every domain covered in both Hindi and Telugu.
AI-Powered Scale
Automated bot pipelines and NLP translation tools allow us to process WikiData and WikiSpecies at a scale impossible with manual effort alone.
Automated Technological Pipeline
01
🗄️ Data Sourcing
Harvesting structured data from WikiData and WikiSpecies — over 800,000 species entries.
02
🤖 Bot Creation
Automated Wikipedia bots programmed to generate, format, and upload article skeletons.
03
🔤 NLP Translation
Translation models fine-tuned on low-resource Indic language corpora.
04
✅ Quality Review
Human-in-the-loop validation by LTRC linguists and domain experts before publication.
05
🌍 Open Release
All tools, datasets, and corpora released open-source for the global research community.
Our Team
Contributing Interns
Recognizing the student researchers and contributors building and reviewing Indic scientific content.
Krupal
Contributor
Abinaya
Contributor
Aditya N Dwivedi
Contributor
Ameya Purohit
Contributor
Saudha
Contributor
Intern 1
Contributor
