Back to Home

EduBench-RU: test of 22 LLMs for schools in Russia

EduBench-RU — benchmark for evaluating LLMs on tasks of Russian schools: FGOS, subjects, pedagogical tools, Chuvash. Tested 22 models, leader — Gemini 3.1 Pro. All failed ChuvashBench.

Test of 22 neural networks: who will cope with FGOS and Chuvash?
Advertisement 728x90

EduBench-RU: Benchmark for Testing LLMs on Russian School Education Tasks

Developers of EduBench-RU tested 22 language models across 50 prompts covering the Federal State Educational Standards (FSES), subject knowledge, pedagogical tools, and the Chuvash language. No model scored above 3 out of 4 on Chuvash tasks. Gemini 3.1 Pro emerged as the leader with an overall score of 3.42.

This benchmark fills a critical gap: existing tests like MERA focus on the Unified State Exam (EGE), while EduBench targets English and Chinese. Here, the emphasis is on real-world scenarios in Russian classrooms.

Test Module Structure

Prompts are divided into four modules for comprehensive evaluation:

Google AdInline article slot
  • Module A (Pedagogy by FSES, 15 prompts): Creating lesson technology maps, student explanations, and analyzing OGE results.
  • Module B (Subject Knowledge, 10 prompts): Math, Russian language, physics, biology, history, and literature problems.
  • Module C (Teacher Copilot, 10 prompts): Developing curriculum plans (KTP), student profiles, parent meeting materials, grading rubrics, and inclusion support.
  • Module D (ChuvashBench, 15 prompts): Translations, grammar exercises in Chuvash, cultural context, and bilingual lessons.

Modules A–C simulate daily teacher workflows. ChuvashBench exposes weaknesses in regional languages.

Tested Models and Methodology

Twenty-two models were selected based on relevance as of March 2026:

  • Frontier: Claude Opus 4.6, Claude Sonnet 4.6, GPT-5.4, Gemini 3.1 Pro, Gemini 2.5 Pro.
  • Mid-tier: GPT-5.4 Mini, Gemini 2.5 Flash, Grok 4.1 Fast, Kimi K2.5, GLM 5, DeepSeek V3.2.
  • Open-source: Qwen3 (8B, 14B, 32B, 235B), Qwen3.5 27B, Mistral Large 3, Llama 4 Maverick, Phi-4.

Testing was conducted via OpenRouter API with max_tokens=8192, temperature=0.7, and a single Russian-language system prompt. Total cost: 1,500 rubles for 2.4 million tokens.

Google AdInline article slot

Evaluation used LLM-as-judge: GPT-5.4 and Claude Sonnet 4.6 served as evaluators. Scoring scale: 1–4 across six criteria—pedagogical quality, Russian language proficiency, factual accuracy, practicality, understanding of Russian context, and Chuvash language use. Final scores are the average of two judges, adjusted for bias (+0.49 points from Sonnet to Claude models).

Overall Model Rankings

| # | Model | Overall | Education | Chuvash | Type |

|---|--------------------|---------|-----------|---------|--------|

Google AdInline article slot

| 1 | Gemini 3.1 Pro | 3.42 | 3.51 | 3.19 | Closed |

| 2 | Claude Opus 4.6 | 3.24 | 3.36 | 2.98 | Closed |

| 3 | Claude Sonnet 4.6 | 3.22 | 3.34 | 2.95 | Closed |

| 4 | Gemini 3.1 Flash Lite | 3.22 | 3.33 | 2.94 | Closed |

| 5 | Gemini 2.5 Pro | 3.21 | 3.31 | 2.98 | Closed |

| 6 | DeepSeek V3.2 | 3.15 | 3.28 | 2.85 | Open |

| 7 | GLM 5 | 3.15 | 3.28 | 2.84 | Closed |

| 8 | Mistral Large 3 | 3.14 | 3.28 | 2.81 | Open |

| 9 | GPT-5.4 | 3.09 | 3.23 | 2.78 | Closed |

|10 | GPT-5.4 Mini | 2.99 | 3.19 | 2.51 | Closed |

Closed models lead (average 3.30), open models trail by 18% (average 2.80). Gemini 3.1 Pro dominates across all metrics; GPT-5.4 ranks 9th.

The Chuvash Language Crisis

ChuvashBench revealed a major failure: no model exceeded 3.0 in accuracy. Distribution (with GPT-5.4 as judge):

  • >3.0 (mostly correct): 0 models.
  • 2.0–3.0 (mixed results): 3 models.
  • 1.0–2.0 (hallucinations): 14 models.
  • =1.0 (complete hallucination): 5 models (Qwen3 variants, Phi-4).

Top performers: Claude Opus 4.6, GPT-5.4, Gemini 3.1 Pro (2.1/4.0). Models generate pseudo-Chuvash—correct words like «Salam» (hello) mixed with fabricated content. Chuvash is spoken by 1.1 million people and official in the Chuvash Republic, currently under UNESCO threat. Data is available (2.9 million sentences, 1.4 million Chuvash-Russian pairs on HuggingFace, CC0 license).

Practical Takeaways for Implementation

For schools, scores of 3.0–3.5 are suitable for draft lesson plans requiring refinement. Under the 152-FZ law (local data), Qwen3.5 27B (3.09, 18 GB VRAM) is viable—12% behind the leader. Regional languages (Chuvashia, Tatarstan, Bashkortostan) remain unsupported. Developers plan ChuvashLM based on Qwen3-32B for local deployment.

Key Takeaways

  • Gemini 3.1 Pro leads with 3.42, outperforming Claude and GPT on FSES and Chuvash tasks.
  • Open models lag by 18%, best being DeepSeek V3.2 (3.15).
  • ChuvashBench: all models hallucinate Chuvash, max score 2.1/4.0.
  • Benchmark is open-source: prompts, results, and code on GitHub.
  • Chuvash training data exists—problem lies in developer priorities.

— Editorial Team

Advertisement 728x90

Read Next