TraviaTechPie Review

Review Tech, Science, Finance

The Story

On August 11, 2026, SK Telecom said its in-house large language model, A.X K2, scored 29 out of 42 on the six problems from this year’s International Mathematical Olympiad. That number matters because 29 is exactly where the gold-medal cutoff sat for IMO 2026. So when SKT calls it “gold-medal level,” it’s not loose marketing — it’s the model landing right on the line.

A quick note on where this claim comes from, because it’s the part most coverage blurs. The IMO figure is SKT’s own evaluation: the company ran the 2026 problems through A.X K2 and graded the results. That’s a company claim, not an independent audit. On the older IMO 2025 set, SKT reports 35 out of 42, with a clean 7-out-of-7 on each of the first five problems. Strong numbers. But again — self-reported.

Where it gets more interesting is the one benchmark SKT didn’t grade itself. On MathArena’s AIME 2026 leaderboard — a public board run by ETH Zurich’s SRI Lab and INSAIT — A.X K2 posted 97.1% accuracy and landed in a tie for first. And the model it tied with is Inkling, the open-weights release from Mira Murati’s Thinking Machines Lab. We wrote about Inkling here a couple of weeks back. The short version was that Thinking Machines built something small and sharp on purpose. Now a Korean telco’s model is sitting next to it at the top of a third-party math board. That’s the data point worth staring at, because nobody at SKT got to pick the grader.

So what actually is A.X K2? It’s a 688-billion-parameter model — but don’t let that number spook you, because it’s a “Mixture-of-Experts” design. In plain terms: the model is huge on paper, but for any single token it only wakes up a small slice of itself, about 3.3 billion parameters’ worth. Think of it as a big building where you only turn on the lights in the rooms you’re using. That’s how you get frontier-scale reasoning without frontier-scale running costs. It’s the same efficiency trick we keep bumping into lately — UNIST’s GMoE work on sharing experts across layers, Thinking Machines going deliberately compact. The whole field is converging on “do more with fewer active parameters.”

Two more facts fill out the picture. First, A.X K2 is genuinely open — the weights are on Hugging Face under Apache 2.0, which means anyone can download it and use it commercially, no permission slip. That puts it in the same lane as the open-weight releases we’ve been tracking, from NVIDIA’s Cosmos to Meta’s Muse Glimmer. Second, the training mix is telling: SKT trained on roughly 8.2 trillion tokens, and the language split is 72.7% English, 15.4% Korean, 8.3% code. Read that again. A model built by a Korean company, pitched as a “sovereign AI” win, is mostly trained on English.

SKT also leaned on one specific framing: among AI models built outside the US and China, A.X K2 is the only one to clear the IMO 2026 gold threshold. On the raw IMO score, several AI models — including Chinese ones — reportedly hit a perfect 42, so this isn’t “best in the world.” It’s “best of the rest” — a very particular slice of the leaderboard.

The Takeaway

Here’s the thing about math-olympiad scores becoming the headline metric for AI. A year ago, the flex was chatbot benchmarks and coding tests. Now it’s IMO problems, AIME accuracy, olympiad medals. That shift isn’t cosmetic. Math olympiad problems are a decent proxy for multi-step reasoning that you can’t fake by memorizing the internet — you either follow the logic or you don’t. When OpenAI buried its next model’s name in a math blog post (we covered that), it was signaling the same thing: math is where labs now go to prove a model can actually think, not just autocomplete. SKT hitting the gold line is a way of saying “our model reasons,” in the language the whole field currently respects.

But I’d separate two questions that the press release wants you to blur together. Question one: is this a real technical result? Mostly yes — the AIME tie with Inkling is third-party, it’s public, and tying a well-funded frontier lab on a live board is not something you fake. Question two: is this “sovereign AI catching up to the frontier”? That’s a bigger claim, and the details push back on it. The IMO gold number is self-graded. The winning IMO score belongs to a Chinese model. And the NYU RITS breakdown of A.X K2 flagged the gap that matters: the model shines on curated math and customer-support simulations but scores badly on open-ended web browsing — around 9 on BrowseComp, versus 98 on a support task. In other words, it’s very good at the narrow, well-defined stuff and still weak at the long, messy agent work that frontier models are racing toward.

That’s actually the honest shape of Korea’s sovereign-AI push right now, and it’s worth naming plainly. The strategy that’s working isn’t “beat OpenAI at everything.” It’s “pick a lane where a focused, efficient, open model can genuinely compete, and win that lane.” Math reasoning is a smart lane to pick — it’s measurable, it’s respected, and it doesn’t require the trillion-token web-scale agent infrastructure the American labs are burning capital on. A 688B MoE model you can download for free that ties a Mira Murati lab on AIME is a real accomplishment. It just isn’t the same thing as closing the frontier gap, and the mostly-English training data is a quiet reminder of how much of the foundation is still borrowed.

My read: watch the third-party boards, not the press releases. The AIME result is the signal here. The IMO gold framing is the story SKT wanted to tell — technically defensible, carefully scoped, and doing a lot of work to make “best outside the US and China” sound like “frontier.” Both things can be true at once. The interesting question for the next year is whether Korea’s open models can carry the math-reasoning strength into the harder, open-ended agent tasks — because that’s where the real gap still is.

This article is for informational purposes only and is not an endorsement to adopt any particular technology or product.


Photo: Growtika / Unsplash

Posted in

댓글 남기기

TraviaTechPie Review에서 더 알아보기

지금 구독하여 계속 읽고 전체 아카이브에 액세스하세요.

계속 읽기