hiraia

Decentralized AI tutoring for the global south

Our on-device AI app works on entry-level Android phones worth around $100 and speaks fluent Tagalog and English (with Bisaya coming soon!)

Try the web demo now, or download the APK directly to your mobile phone below.

Learn More

hiraia is a personal AI tutor in your pocket, built with QVAC

Fine-tuned open-weights LLMs

  • Replies in Tagalog, Bisaya, and English
  • Augments in-person classroom learning with review and reinforcement at home on-demand

Offline mode

  • No internet required for chat, entire package is less than 4 GB
  • Internet only needed for the initial download and databank updates

Entry-level phone requirements

Runs on cheap phones from five years ago; no other costs or ongoing fees!

Hiraia App Preview

Free to download

No account, no fees, no Play Store. Just download, install, and start learning.

Hiraia Cat

Most users3B model · GPU

For phones with 8 GB+ RAM.

Larger 3B model on your phone's GPU — higher-quality answers, faster replies.

Download Hiraia Cat

APK ~679 MB · first-launch model ~3.6 GB · Android 12+

Hiraia Kitten

Experimental1B model · CPU

For phones with 8 GB RAM or less.

Smaller 1B model, CPU-only — runs on phones that can't handle the larger one. Slightly less detailed, still grounded in the same curriculum.

Experimental. The smaller 1B model can make mistakes — including on safety and “is this true?” questions. Always double-check important answers with a teacher or trusted adult. For the most reliable answers, use Hiraia Cat if your phone supports it.

Download Hiraia Kitten

APK ~560 MB · first-launch model ~1.3 GB · Android 12+

Not sure which? If your phone has 8 GB+ RAM, get Cat. If it has less than 8 GB, get Kitten. The app is free either way — if Cat won't open on your phone, just install Kitten instead.

How the download works

The app itself is a small download — just Hiraia, its offline science knowledge bank, and the illustrations. The first time you open it, Hiraia downloads the AI model that matches your tier, served from Hiraia's own servers. After that, Hiraia runs fully offline — no internet, no account, and nothing you type ever leaves your phone.

Hiraia performance

hiraiabench measures our fine-tuning honestly. Read the table top to bottom: scores fall as the models get smaller — except our row. Hiraia is a 3B model, but our fine-tune makes it score like one several times its size, beating its own untrained base and matching far larger open models on the dimensions that matter for a Filipino science tutor. Scored 0–5 by a blind LLM judge across 67 probes; higher is better.

ModelScienceTagalogEnglishBisayaPedagogySafetyTaglish
Claude Opus 4.8
cloud frontier (reference ceiling)
5.05.05.05.05.05.05.0
Qwen3.5-27B
recent open model, ~9× our size
4.94.55.03.94.35.04.6
Qwen3.5-9B
recent open model, ~3× our size
3.33.25.01.83.73.43.4
HiraiaHiraiaours
our on-device 3B fine-tune (Sailor2-3B + LoRA)
3.13.93.63.02.73.13.3
Sailor2-3B
our untrained 3B base
2.32.42.02.52.01.52.9
Qwen3-1.7B
untrained small base
0.00.23.90.00.20.40.1

What each category means

Science Accuracy
Is the science correct and pitched for a 5th-grader? A confident wrong answer or an affirmed myth scores near zero — accuracy outranks fluency.
Tagalog Fluency
How natural, warm, and age-appropriate the Tagalog reads for a child — tutor register, not just grammar. A fluent but stiff, essay-style answer scores lower than a simple, friendly one.
English Fluency
The same register and clarity test, for questions a student asks in English.
Bisaya
Cebuano answers, judged on both natural register and factual correctness — the hardest language for most models.
Pedagogy
Does it actually teach? Simple words, builds intuition, and encourages the child — without dumbing the science down into wrongness.
Safety & Honesty
Gently corrects common myths and gives safe everyday answers — and honestly says “I’m not sure” on genuinely unknowable questions (tomorrow’s weather, a lottery number) instead of inventing one.
Code-switching
Filipino kids mix Tagalog and English mid-sentence (“Taglish”). This measures whether the tutor follows that naturally and still answers well.

Single-turn probes from our internal capability suite (no retrieval/RAG), one shared neutral tutor prompt for every model, blind comparative LLM-judging (answers anonymized, scored 0-5 on accuracy/helpfulness/faithfulness/naturalness/pedagogy, then projected to the 7 categories).

  • Qwen3.5 ships in many sizes (0.8B–72B + a 30B-A3B MoE); we show 9B (~3× our size) and 27B (~9×) as open references. No 14B exists in this generation.
  • The 3B/1.7B/9B models and Hiraia run at Q4_K_M (llama.cpp) — the on-device quant a deployer would use; the 27B runs at bf16 (vLLM) as a full-precision larger-model reference.
  • No RAG: every model answers from its own weights under the same tutor prompt. Hiraia’s product adds retrieval grounding on top — not reflected here.
  • Hiraia’s Bisaya here uses the tl/en LoRA (cross-lingual transfer); the product ships a separate Bisaya adapter, not benchmarked.
  • Under the shared Filipino-tutor prompt the 27B answered English questions in Tagalog (a language-matching miss, not an English deficit); its English row was re-measured with an explicit “reply in English” instruction.
  • Claude Opus 4.8 saturates the rubric (5.0) — the frontier reference ceiling, not a same-class model. Qwen3.5-27B was scored against the other five as frozen calibration anchors.

Why we built hiraia

Founder Luis Buenaventura talks about the declining aptitude of Filipino schoolchildren and what we believe AI can do to get us back on track.