Academic papers

Papers brief: Solar Open 2 — 250B MoE built for long agent trajectories

arXiv Solar Open 2 technical report: 250B-A15B MoE, 1M-token hybrid attention, strong Korean benchmark average among compared models.

  • academic papers
  • large language models
  • Korea

Source: arXiv

Paper

Solar Open 2 Technical Report — Park, Kim, Gim, Cho, Ko, et al. (submitted 22 Jul 2026)
ID: arXiv:2607.20062

What it claims

Solar Open 2 is a 250B-A15B Mixture-of-Experts language model aimed at long-horizon agentic tasks, scaled from Solar Open 1 (100B class). It targets a 1M-token context via a hybrid attention stack (one softmax layer among every three linear-attention layers), no positional encoding, and a gated delta rule extended to negative eigenvalues.

Training efficiency under a fixed compute budget comes from initializing from Solar Open 1 (transferring a 5.69B shared skeleton) and curating a higher-value data mixture (refining a 20T pool into a 10T mix). Agent skills are built by training twelve domain specialists, then consolidating with Multi-teacher On-Policy Distillation (MOPD). On English benchmarks, the report claims leadership among comparably sized open-weight models on MMLU-Pro, LiveCodeBench, and APEX-Agents, with competitiveness elsewhere. On Korean benchmarks, it records the highest average among models compared (including fast-tier closed APIs); on in-house Ko-GDPval (Korean officework-agent), it is competitive with a much larger DeepSeek-V4-Pro at under a sixth the size — per the abstract’s framing.

The breakdown

Three engineering bets matter for readers who will never train a 250B MoE. First, the 1M-token hybrid attention story is about keeping whole agent trajectories in one context — tool calls, retries, and Korean/English mixed officework — instead of truncating the audit trail. Second, training under a fixed budget via Solar Open 1 initialization plus a curated 10T mixture (from a 20T pool) is a methods claim about value per token, not magic scale. Third, twelve specialists distilled with MOPD is how agent skills are supposed to consolidate without keeping twelve deployable models. The abstract’s Korean-benchmark average and Ko-GDPval comparison are the local performance hooks; treat them as author-reported until you replicate on your tasks.

Why readers outside the lab should care

If you buy or build Korea-facing agents (officework, Korean + English mixed context, long tool traces), open technical reports like this are how you compare context length and Korean eval without waiting for vendor slide decks alone. Expats and firms routing work through Korean-language agents should watch whether “agentic + Korean average” claims survive third-party replication — and whether a 1M window changes how you log full trajectories for audit.

What travelers and expats should watch

  • Do separate abstract claims (Korean average, Ko-GDPval) from your own task eval before switching production models.
  • Do ask vendors for trajectory-length behavior (1M-class windows) when agents must keep long tool histories.
  • Don’t treat a company-linked technical report as an independent news desk finding — it is an OA methods/results document.
  • Expect benchmark suites and baselines to be detailed in the PDF; cite numbers only as the abstract states them here.

Context

Read this as an open-weight capability map with a Korea eval accent, not as a purchase order. Korelay frame: the overseas-useful question is whether long-horizon agents that score on Korean officework change your verification habits — still verify high-stakes outputs against primary sources. A technical report on arXiv is a readable methods surface; it is not a substitute for your own eval harness.

Source

arXiv:2607.20062 — abstract and framing cited; open the OA PDF for architecture, data mix, and full tables. Do not republish the PDF.