Maelis Research LogoMaelis Research
ResearchProductsAboutContact

Language intelligence for Odia
and low-resource Indian languages

We are building the AI layer for 50 million Odia speakers — tokenizer optimization, language models, speech, and translation for India's underserved languages.

Get in Touch

What We Do

Efficient Tokenization

Our Odia-optimized tokenizer is 3x more efficient than generic alternatives. Built for Brahmic scripts with Unicode awareness. This means lower costs and better performance on every API call.

Lekhani Model Family

Odia-exclusive LLMs built on open-weight base models. Named after Odia literary tradition — Pada (verse), Chhanda (meter), Kavya (poetry), Mahakavya (great epic). All released under Apache 2.0 license.

Shruti & Anuvada

Odia speech recognition and synthesis (Shruti = "sound" in Odia) and Odia translation (Anuvada = "translation" in Odia). Built for government, media, and enterprise use.

The Odia Language Gap in AI

  • •Global AI models are objectively 3x+ worse at Odia than English
  • •No commercial Odia-specific AI API currently exists
  • •Existing academic projects use non-commercial licenses
  • •50 million speakers remain underserved

How Maelis Solves It

  • •Odia-exclusive tokenizer with 3x better efficiency
  • •Full commercial license (Apache 2.0) for enterprise use
  • •SLA-backed API with dedicated support
  • •Built specifically for Odisha government and business needs

Research & Publications

Our research spans tokenizer optimization, model distillation, and dataset construction for low-resource Indian languages. All research models are released under Apache 2.0 license.

An Empirical Evaluation of Efficient Odia Tokenization for Large Language Models

Published

Benchmarks 16 tokenizer configurations for Odia. Our champion tokenizer achieves industry-leading efficiency for the Odia language.

ArXivGitHub

Distilling Sarvam AI's Odia Generation Capability

Research Complete

Analysis of open-weight distillation for Odia LLMs. Route B (self-hosted Apache 2.0 weights) recommended over paid API distillation for legal and cost reasons.

Internal Research

odia-pretrain-dataset

Open Source

9.75M rows of cleaned Odia text from 25+ sources. CC-BY-4.0 licensed. Monolingual, parallel, instruction, and QA data types with 80.6% mean Odia content ratio.

View on Hugging Face

Product Architecture

MAELIS RESEARCH
├── Lekhani— Foundation Model FamilySoon
│ ├── Lekhani Pada— Lightweight(Efficient for edge deployment)Soon
│ ├── Lekhani Chhanda— Balanced(General purpose)Soon
│ ├── Lekhani Kavya— Large(Deep reasoning)Soon
│ └── Lekhani Mahakavya— Flagship(Maximum capability)Soon
├── Shruti— Odia Speech Recognition & SynthesisSoon
├── Anuvada— Odia Translation (English ↔ Odia)Soon
├── Khoja— Odia Semantic SearchSoon
├── Patra— Odia OCR (Document Understanding)Soon
└── Artha— Odia Reasoning (Chain-of-thought)Soon

Lekhani

Coming Soon

Foundation Model Family

Foundation model family optimized for Odia language with multiple capability tiers.

Shruti

Coming Soon

Odia Speech Recognition & Synthesis

Speech recognition and synthesis for Odia, built for enterprise and government use.

Anuvada

Coming Soon

Odia Translation (English ↔ Odia)

High-quality translation between Odia and other languages with cultural context.

Khoja

Coming Soon

Odia Semantic Search

Semantic search and information retrieval specifically trained for Odia content.

Patra

Coming Soon

Odia OCR (Document Understanding)

Optical character recognition and document understanding for Odia text.

Artha

Coming Soon

Odia Reasoning (Chain-of-thought)

Advanced reasoning capabilities for complex Odia language tasks and analysis.

About the Founders

Sai Dutta Abhishek Dash

Co-Founder & CEO, Maelis Research

Dhenkanal, Odisha, India

AI Infrastructure & Security Engineer building open-source AI infrastructure, developer tools, and security products. Building language AI for his mother tongue, Odia.

Background

  • • AI infrastructure & LLM gateway specialist
  • • Code security agents & distributed systems
  • • 20+ deployed products, 50+ public repositories
  • • Built Epoxy, Vulscany, MarkItDownJS & more
  • • NLP research and tokenizer optimization

Connect

LinkedInWebsiteGitHubTwitterEmail

Saurav Mahalik

Co-Founder & CTO, Maelis Research

Bhubaneswar, Odisha, India

Engineering scalable language technology to bridge the digital divide for Odia speakers.

Background

  • • Machine learning engineering and model deployment
  • • Distributed systems and cloud infrastructure
  • • Open-source contributor and community builder
  • • Experience in building scalable AI pipelines
  • • Passionate about Indic language technology

Connect

LinkedInGitHubEmail

Mission

Maelis Research is a language infrastructure company. We build the AI layer for the Odia language — tokenization, large language models, speech recognition, translation, semantic search, and document understanding — all optimized for Odia.

We are Odia specialists, not 22-language generalists. Every part of our stack is designed for one language done exceptionally well, not many languages done adequately.

We believe in open-source as a force multiplier for underserved communities. Our core models and tools are released under permissive licenses (Apache 2.0) so developers, researchers, and organizations can build on them freely. At the same time, we support sustainable commercial models — because open-source thrives when the people and companies building it can sustain their work. Founded in 2026 and based in Odisha, we serve government, enterprise, and developer customers who need reliable, commercial-grade Odia language AI.

Maelis Research Logo

Maelis Research

Language infrastructure for Odia and low-resource Indian languages.

Dhenkanal, Odisha, India

Quick Links

  • Research
  • Products
  • About
  • Contact

Connect

GitHubLinkedInTwitterHugging Face

Maelis Research © 2026

Dhenkanal, Odisha, India

maelis.research@gmail.com

Language intelligence for Odia and low-resource Indian languages.