Language intelligence for Odia
and low-resource Indian languages
We are building the AI layer for 50 million Odia speakers — tokenizer optimization, language models, speech, and translation for India's underserved languages.
What We Do
Efficient Tokenization
Our Odia-optimized tokenizer is 3x more efficient than generic alternatives. Built for Brahmic scripts with Unicode awareness. This means lower costs and better performance on every API call.
Lekhani Model Family
Odia-exclusive LLMs built on open-weight base models. Named after Odia literary tradition — Pada (verse), Chhanda (meter), Kavya (poetry), Mahakavya (great epic). All released under Apache 2.0 license.
Shruti & Anuvada
Odia speech recognition and synthesis (Shruti = "sound" in Odia) and Odia translation (Anuvada = "translation" in Odia). Built for government, media, and enterprise use.
The Odia Language Gap in AI
- •Global AI models are objectively 3x+ worse at Odia than English
- •No commercial Odia-specific AI API currently exists
- •Existing academic projects use non-commercial licenses
- •50 million speakers remain underserved
How Maelis Solves It
- •Odia-exclusive tokenizer with 3x better efficiency
- •Full commercial license (Apache 2.0) for enterprise use
- •SLA-backed API with dedicated support
- •Built specifically for Odisha government and business needs
Research & Publications
Our research spans tokenizer optimization, model distillation, and dataset construction for low-resource Indian languages. All research models are released under Apache 2.0 license.
An Empirical Evaluation of Efficient Odia Tokenization for Large Language Models
PublishedBenchmarks 16 tokenizer configurations for Odia. Our champion tokenizer achieves industry-leading efficiency for the Odia language.
Distilling Sarvam AI's Odia Generation Capability
Research CompleteAnalysis of open-weight distillation for Odia LLMs. Route B (self-hosted Apache 2.0 weights) recommended over paid API distillation for legal and cost reasons.
odia-pretrain-dataset
Open Source9.75M rows of cleaned Odia text from 25+ sources. CC-BY-4.0 licensed. Monolingual, parallel, instruction, and QA data types with 80.6% mean Odia content ratio.
View on Hugging FaceProduct Architecture
Lekhani
Coming SoonFoundation Model Family
Foundation model family optimized for Odia language with multiple capability tiers.
Shruti
Coming SoonOdia Speech Recognition & Synthesis
Speech recognition and synthesis for Odia, built for enterprise and government use.
Anuvada
Coming SoonOdia Translation (English ↔ Odia)
High-quality translation between Odia and other languages with cultural context.
Khoja
Coming SoonOdia Semantic Search
Semantic search and information retrieval specifically trained for Odia content.
Patra
Coming SoonOdia OCR (Document Understanding)
Optical character recognition and document understanding for Odia text.
Artha
Coming SoonOdia Reasoning (Chain-of-thought)
Advanced reasoning capabilities for complex Odia language tasks and analysis.
About the Founders
Sai Dutta Abhishek Dash
Co-Founder & CEO, Maelis Research
Dhenkanal, Odisha, India
AI Infrastructure & Security Engineer building open-source AI infrastructure, developer tools, and security products. Building language AI for his mother tongue, Odia.
Background
- • AI infrastructure & LLM gateway specialist
- • Code security agents & distributed systems
- • 20+ deployed products, 50+ public repositories
- • Built Epoxy, Vulscany, MarkItDownJS & more
- • NLP research and tokenizer optimization
Connect
Saurav Mahalik
Co-Founder & CTO, Maelis Research
Bhubaneswar, Odisha, India
Engineering scalable language technology to bridge the digital divide for Odia speakers.
Background
- • Machine learning engineering and model deployment
- • Distributed systems and cloud infrastructure
- • Open-source contributor and community builder
- • Experience in building scalable AI pipelines
- • Passionate about Indic language technology
Connect
Mission
Maelis Research is a language infrastructure company. We build the AI layer for the Odia language — tokenization, large language models, speech recognition, translation, semantic search, and document understanding — all optimized for Odia.
We are Odia specialists, not 22-language generalists. Every part of our stack is designed for one language done exceptionally well, not many languages done adequately.
We believe in open-source as a force multiplier for underserved communities. Our core models and tools are released under permissive licenses (Apache 2.0) so developers, researchers, and organizations can build on them freely. At the same time, we support sustainable commercial models — because open-source thrives when the people and companies building it can sustain their work. Founded in 2026 and based in Odisha, we serve government, enterprise, and developer customers who need reliable, commercial-grade Odia language AI.