Academic papers

Papers brief: K-EXAONE 2.0 scales Korea’s open MoE to 750B

LG AI Research’s arXiv report: upcycled MoE, 10 languages, Apache 2.0 — what overseas buyers should ask beyond the parameter headline.

  • academic papers
  • Korean AI
  • open-weight models

Source: arXiv

Paper

K-EXAONE 2.0 Technical Report: Journey to Global Frontier-Scale Foundation Models — LG AI Research (arXiv 2608.04505, Aug 2026).

What it claims

Korea’s government GPU program funded a second-phase open-weight foundation model that upcycles the earlier K-EXAONE MoE rather than training a larger sibling from scratch. K-EXAONE 2.0 reaches 750B total parameters with about 37B activated per token (from 236B / 23B), keeps a hybrid sliding-window + global attention stack for up to 256K context, and expands language coverage from six to ten languages including Korean and English. The authors say they release the model under Apache 2.0.

The report’s product claim is not only scale: mid- and post-training stress agentic coding, tool use, long-context retrieval, and a Korea-augmented safety taxonomy grown by refining 100+ risk areas and adding 70 new ones. Speculative decoding attaches MTP and a stronger DSpark drafter to the same target weights for serving speedups.

The breakdown

Architecture: 78 layers (two dense + 76 MoE), 256 experts with top-8 routed plus one shared expert activated per token, Clamped SwiGLU on the last 16 MoE layers for activation stability. Depth grows by repeating mid-stack LLLG blocks (48→78 layers); width doubles experts (128→256) with rotation noise so duplicated experts differentiate. Continual pre-training adds a healing stage after upcycling, then roughly 8T further tokens with Active Reading and thinking-augmented synthetic data recipes.

Inference table (same FP8 target, draft budget γ=7): DSpark beats MTP on acceptance length by about 32–66%, with end-to-end speedups reported around 1.81–2.57× versus 1.27–1.77× for MTP across math and code benches. Evaluation spans nine categories — knowledge, math, coding/agentic coding, tool use, instruction following, long context, Korean understanding, multilingual, safety — with the authors highlighting largest gains versus the prior model in agentic coding and long-context work, and advantages versus peer open-weight models on long-context retrieval and safety.

Why readers outside the lab should care

If you procure Korean open weights for enterprise RAG, coding agents, or bilingual KR/EN products, this paper is a vendor questionnaire, not a hype slide. Ask for Apache 2.0 artifact location, activated-parameter cost at your context length, whether DSpark is in the served stack, and how the expanded safety taxonomy maps to your risk register. Overseas teams that only score Korean models on English MMLU will miss the report’s own pitch: long-context retrieval and Korea-grounded safety as the differentiators.

What builders and buyers should watch

  • Do price 37B activated MoE serving, not the 750B billboard — the report’s efficiency claim lives there.
  • Don’t assume “upcycled” means free lunch; healing + 8T tokens still imply serious compute and data ops.
  • Expect ten-language marketing; verify Korean + your target EU/SEA languages on your tasks, not only the paper’s suite.
  • Re-check license and export-control posture before shipping Apache weights into regulated stacks — open-weight ≠ cleared for every jurisdiction.

Context

Read this as Korea industrial-policy AI with a serving story, not as proof that a domestic MoE “beat the frontier.” Korelay’s frame: the keep is upcycling + speculative decoding + safety taxonomy expansion — the parts that change how you evaluate a Korea open model beyond counting parameters.

Source

Primary: arXiv — K-EXAONE 2.0 Technical Report (LG AI Research). Paraphrase of the technical report abstract and body; figures and bench numbers as stated by the authors.