
Papers brief: A.X K2 — SK Telecom’s 688B MoE bets on sparse long context
arXiv A.X K2 technical report: from-scratch agentic MoE, Sparse Gated Attention to 256K, Think-Fusion modes — what Korea open-weight buyers should price before the parameter headline.
Source: arXiv
Paper
A.X K2 Technical Report — SK Telecom (arXiv 2608.30181, Sep 2026).
What it claims
A.X K2 is a 688B-parameter Mixture-of-Experts language model trained from scratch as a high-performance foundation for agentic applications — coding agents, tool use, and long-horizon tasks rather than chat-only serving. SK Telecom positions it as the successor to A.X K1.
The headline efficiency claim is token budget: roughly 8.5T training tokens — fewer than K1 — on a smaller but higher-quality mixture with substantially expanded agentic and software-engineering data, yet still improving over K1 across the board, including gains of over 30 percentage points on some benchmarks. That is a methods story about data curation and MoE scaling, not brute-force token pile growth.
For long context, the authors introduce Sparse Gated Attention (SGA), combining sparse attention with gated attention, plus Gated Norm (GN) to stabilize large-scale training. SGA trains natively at 128K through a sparse indexer warmup that optimizes against its own sparse top-k selection rather than the dense attention distribution. Each query reads only 2,048 positions, yet the authors report unchanged long-context quality and a 94.6 score on RULER out to 256K. GN’s outlier suppression keeps 4-bit NVFP4 serving within one point of FP8 accuracy. A Think-Fusion recipe lets users switch between thinking and non-thinking modes in a single unified model. Evaluations claim competitiveness against strong open-weight baselines, matching or exceeding them on math and Korean-language benchmarks.
The breakdown
Three engineering bets matter if you will never train a 688B MoE yourself. First, sparse long context is the serving story: attending to 2,048 positions instead of the full window is how the model claims 256K-class retrieval without paying dense attention cost on every token. Second, Think-Fusion is an inference-economics knob — reasoning depth becomes a per-request switch, not a separate model SKU. Third, the Korean benchmark hook is explicit in the abstract: this is not marketed as English-only frontier chasing; Korean-language eval is part of the competitive claim alongside math.
Treat the 30+ point jumps as benchmark-specific until you replicate on your tasks. The abstract frames them as token-efficiency gains, not proof that every downstream agent workflow gets cheaper overnight.
Why readers outside the lab should care
Korea’s telecom-backed open-weight stack is now publishing at 688B MoE scale with an agentic and Korean-eval pitch. If you procure models for bilingual KR/EN products, coding agents, or long-document RAG inside Korea, this report is a vendor questionnaire: ask for activated-parameter serving cost at your context length, whether SGA sparse routing is in the deployed stack, and how Think-Fusion maps to your latency SLA. Overseas teams that score Korean models only on English MMLU will miss the abstract’s own differentiator — Korean-language benchmarks alongside math.
Expats and firms routing work through Korea-hosted agents should watch whether “fewer tokens, better benchmarks” survives third-party replication — and whether NVFP4-class quantization holds on your longest contexts.
What builders and buyers should watch
- Do price sparse-attention serving at your real context length — the abstract’s win is 2,048 reads per query, not a free 256K window.
- Don’t treat the 688B billboard as your GPU bill; ask vendors for activated-parameter throughput and NVFP4 vs FP8 accuracy on your stack.
- Expect Think-Fusion to need product wiring — mode switching is a deployment choice, not automatic cost savings.
- Re-check Korean-language claims on your tasks; abstract competitiveness is author-reported until you run your own eval harness.
Context
Read A.X K2 as Korea sovereign-scale open MoE with a serving story, not as proof that a domestic model “won the frontier.” Korelay’s frame: the keep is SGA long context + Think-Fusion economics + Korean eval — the parts that change how you evaluate a Korea open model beyond counting total parameters.
Source
Primary: arXiv — A.X K2 Technical Report (SK Telecom). Paraphrase of the abstract and framing; open the OA HTML/PDF for architecture tables and full benchmark suites. Do not republish the PDF.