· 000· (Araskova.sys)· (Loading Environment)
A

"Where the monsoons are born."

8.5°N · 76.9°E
· SYSTEM_READY: FALSE
000
NLP Research

Malayalam LLM Benchmark: Failure Analysis on Administrative Queries

An internal benchmark of leading LLMs on real-world Malayalam administrative queries reveals a ~73% failure rate — validating our regional data engine thesis.

Overview

We ran an internal benchmark of leading public LLMs — including GPT-4o, Gemini 1.5 Pro, Claude 3 Opus, and Llama 3 70B — against a curated set of 200 real-world administrative queries in Malayalam. These queries were drawn from the kinds of tasks a government office worker, a bank employee, or a primary school teacher in Kerala would actually perform.

The result: approximately 73% failure rate across all models on accurate, contextually-correct Malayalam output.

This is not a failure of intelligence. It is a failure of data. No major model has been trained on sufficient Malayalam at the density required for professional-grade outputs. English translations, transliterations, and approximate answers dominate responses where precise Malayalam is required.

What We Mean by "Failure"

We defined failure as any of the following:

  • Code-switching: the model reverts to English mid-sentence for technical or administrative terms
  • Transliteration errors: phonetically incorrect Malayalam rendering (Romanised Malayalam back-transliterated incorrectly)
  • Semantic drift: the response is grammatically valid Malayalam but loses critical nuance from the original query (e.g., a legal deadline, a form field label, an eligibility condition)
  • Hallucination of terminology: the model invents Malayalam equivalents for administrative terms that have no accepted translation

Why This Matters

Kerala has approximately 38 million Malayalam speakers. Administrative functions — land registry, ration card management, school enrollment, local governance — are conducted in Malayalam. A usable AI assistant for any of these domains requires professional-grade Malayalam output, not approximate translations.

The ~73% failure rate establishes a clear market gap. No existing model closes it. That is the gap Araskova's Malayalam-first data engine is designed to fill.

Our Approach

We are building the data moat, not training the model yet. The first phase is corpus construction: 47.5M tokens of curated Malayalam text spanning administrative documents, educational material, legal text, and conversational data — all sourced, annotated, and quality-checked by native speakers.

Once the corpus reaches sufficient density and domain coverage, fine-tuning and evaluation will follow. We will publish the benchmark methodology and dataset in full at that point.

Status

  • Benchmark dataset: Complete (200 queries, 4 models, scored)
  • Corpus construction: Active — 47.5M tokens and growing
  • Model fine-tuning: Planned for Q4 2026
  • Public paper: In preparation

This benchmark was conducted internally by Araskova Labs in June 2026. Raw data available on request.