
Papers brief: NOLLI maps where Korean LLM gaps actually break
arXiv NOLLI’s 7,500 procedurally generated English–Korean puzzles show presentation language is cheap; Hangul jamo multi-step work and kinship notation are not.
Source: arXiv
Paper
NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap — Dasol Choi, Joonyong Park, Daegon Yu, Soo Yong Kim, Youngsook Song, Seunghyeok Hong (HAERAE LAB and collaborators; arXiv 2608.04397, Aug 2026).
What it claims
When a frontier model scores lower in Korean than English, the headline is usually “Korean lag.” NOLLI (Korean for “logic”) asks a sharper question: which step breaks — surface presentation, Hangul jamo manipulation, or Korea-specific culture and orthography?
The benchmark is procedurally generated: 15 puzzle types, 25 language-specific tasks, 7,500 items. Every instance is seed-regenerable, verified to have a unique solution, and scored with deterministic exact match. Difficulty is tuned behaviorally on a fixed reference model until accuracy lands in target bands — not by making puzzles structurally bigger.
Three levels organize the design. Level 1 pairs matched direct translations. Level 2 adapts script-intensive puzzles over Hangul jamo at sub-syllabic granularity. Level 3 adds Korean-only tasks grounded in Korean culture or orthography.
The authors evaluate 15 frontier, open-weight, and Korean-developed models. Among 12 above a 3% overall-accuracy floor, matched English–Korean scores on direct translations are statistically equivalent within a ±10 pp TOST margin — little systematic cost from presentation language alone. Writing-system tasks tell a different story: Korean Cipher trails English by up to 68.7 pp, while Cryptarithmetic over the same jamo shows no systematic penalty; Jamo Composition accuracy predicts Korean Cipher accuracy. On Korean-only tasks, rule-application deficits vary in sign, but a Kinship deficit is positive in all 12. Finally, a salient size measure fails to grow from Easy to Hard in 7 of 15 types — structural size is an unreliable hardness proxy.
The breakdown
NOLLI is diagnostic infrastructure, not a trophy board. Level 1 estimates presentation-language effects; Levels 2 and 3 localize script and Korea-specific bottlenecks without treating them as additive causes — a common collapse in multilingual evals.
Behavioral calibration puts heterogeneous puzzle families on one accuracy scale. The jamo findings fit a known Korean NLP tension: models must decompose, transform, and recompose sub-syllabic units across steps, not only read syllable blocks. Kinship tasks probe whether chat fluency extends to formal cultural notation — the kind that shows up in family registers, genealogy forms, and structured Korean labels overseas readers encounter in admin paperwork.
Why readers outside the lab should care
Product and localization teams shipping Korean LLM features should stop treating “translate the prompt” as the main fix. If failure modes are cipher-like multi-step Hangul work or kinship-term chains, you need jamo-aware tests and culture-specific suites — not another English leaderboard screenshot.
Korea-built model buyers can use NOLLI’s three-level split to ask vendors where they fail: matched translation, script adaptation, or Korean-only reasoning. Expats filling Korean family-register forms hit the same kinship notation the benchmark probes — fluency in chat English does not guarantee the model handles those structured labels.
What builders and Korea-touching teams should watch
- Do separate matched-translation evals from jamo-intensive and kinship tasks before declaring “Korean parity.”
- Don’t equate harder with larger item size; NOLLI finds size fails as a difficulty proxy in 7 of 15 types.
- Expect Korean Cipher-style gaps (up to 68.7 pp in the paper) even when direct-translation scores look fine within ±10 pp.
- Re-check any Korea customer-support or genealogy feature against kinship and cultural tasks, not only MT benchmarks.
Context
Read NOLLI as a fault map for English–Korean reasoning, not as proof that Korean is “hard for AI” in the abstract. Korelay’s frame: presentation language is cheap; multi-step jamo execution and kinship literacy are where the tab stays open.
Source
Primary: arXiv:2608.04397 (abstract and framing cited; open the OA PDF for methods and full results). Do not republish the PDF.