Trackers / Frontier model tracker
China Frontier LLM Release Tracker
Every flagship large language model from China's frontier labs: lab, release date, parameters, context, license, public pricing and a benchmark claim, each row traced to a named source.
As of 28 July 2026. Updated from China AI Dispatch's daily reporting. Machine-readable version:frontier-models.json.
The question behind every Chinese model release is the same: how close is it to the US frontier, and what does it cost to run? This page answers both as a table. It lists the flagship models from the labs that matter, with when each shipped, its size and context where the lab disclosed them, its license, its public price, and one load-bearing benchmark result, each attributed. It is a companion to the domestic chip scoreboard, which tracks the silicon these models increasingly run on.
Two honest caveats sit over the whole table. First, benchmarks are contested ground. Where a number is the lab's own claim, we mark it lab-reported; where it comes from an independent leaderboard (Artificial Analysis, OpenRouter, Chatbot Arena, Code Arena), we mark it independent. A vendor's "we beat everyone" is never presented as neutral fact. Second, several labs do not publish a parameter count or context figure, most notably Alibaba for Qwen; those cells read "n/d" rather than carrying a guess. We track one flagship per lab to keep the table readable; the daily carries every point release.
The tracker
Flagship large language models from Chinese frontier labs, newest first. Cells marked n/d are specs the lab did not disclose. Each benchmark cell is tagged with who reported it. Each row links to its source below.
| Model | Lab | Released | Params | Context | License | Pricing | Benchmark |
|---|---|---|---|---|---|---|---|
| Kimi K3 | Moonshot AI | Week of 17 Jul 2026 | 2.8T per Qbit’s launch-day report (called the largest open-weight model to date); pre-release estimates ranged 2.5T to 3T | n/d | Open-weight | n/d | Took the top spot on the Frontend Code Arena, ahead of the leading US models; Chinese models still trail the top US labs on the hardest independent coding benchmarksindependent Source: Qbit (量子位); Sina Finance; 36Kr · CAD issue |
| LongCat-2.0 | Meituan (LongCat) | 30 Jun 2026 (open-sourced at release) | 1.6T total, about 48B active (MoE) | 1M tokens | Open-weight | n/d | Meituan reports it as the first trillion-parameter model trained and inferenced entirely on a 50,000-card domestic cluster with no foreign accelerators, on over 30 trillion tokenslab-reported Source: SCMP; Reuters; Geek Park; InfoQ · CAD issue |
| GLM-5.2 | Zhipu (Z.ai) | 13 Jun 2026 (API and MIT open-source rolled out the following week) | Not disclosed at release | 1M tokens | Open-weight (MIT) | Roughly one-sixth the API cost of GPT-5.5 (VentureBeat). On the Harvey legal benchmark it averaged about $2.40 per task against about $31 for Anthropic’s Fable, a roughly 13x spread (Applied Compute case study, via CAD, 28 Jul 2026). Absolute per-token price not disclosed | First open-weight model to land at 51 on the Artificial Analysis Intelligence Index, between GPT-5.5 and Opus 4.8; 74.4 on FrontierSWE, within a point of Opus 4.8independent Source: SCMP; VentureBeat; Artificial Analysis; Applied Compute; tmtpost · CAD issue |
| MiniMax M3 | MiniMax | 1 Jun 2026 | Not disclosed | 1M tokens | Open-source (weights released about 10 days after the 1 Jun launch) | ¥119 per month for 18 billion tokens, roughly 15x the token volume of a comparable Claude plan | Global rank #7 on the Artificial Analysis intelligence index; MiniMax reports 93.2% on GPQA Diamond, above Claude Opus 4.8 and 4.7independent Source: Artificial Analysis; MiniMax technical report · CAD issue |
| Qwen3.7-Max | Alibaba (Tongyi/Qwen) | 20 May 2026 | Not disclosed (Alibaba does not publish Qwen parameter counts) | n/d | Proprietary (API only) | Per-token price not disclosed by Alibaba | Highest-ranked Chinese model on the LMArena leaderboard, 13th globally, between GPT-5.5 and Grok 4.2; fifth globally on Artificial Analysisindependent Source: LMArena; Artificial Analysis; Alibaba Qwen team; InfoQ · CAD issue |
| DeepSeek V4-Pro | DeepSeek | 24 Apr 2026 | 1.6T total, 49B active (MoE) | 1M tokens | Open-weight (MIT) | About $3.48 per million output tokens, roughly one-seventh the price of Claude Opus 4.7 | Coding a hair behind the best Western models; the DeepSeek report’s 80.6% on SWE-Bench Verified was independently confirmed, behind Opus 4.7 at 87.6% but leading on LiveCodeBenchindependent Source: DeepSeek technical report; Reuters; independent benchmarks · CAD issue |
| Hunyuan Hy3 | Tencent Hunyuan | Apr 2026 preview (open-sourced Jul 2026) | 295B total, 21B active (MoE) | 256K tokens | Open-weight (open-sourced after an open preview) | About 1.2 yuan ($0.17) per million input tokens on Tencent Cloud, 0.4 yuan on a cache hit | Topped OpenRouter’s weekly token-usage chart at 3.66 trillion tokens (a usage metric, not a quality test); InfoQ rated it top-tier on coding-agent and instruction-following tasks among comparably sized models at launchindependent Source: OpenRouter; tmtpost; InfoQ; Reuters · CAD issue |
| Kimi K2.6 | Moonshot AI | 20 Apr 2026 | 1T (MoE) | n/d | Modified MIT | $0.95 in / $0.16 cached / $4.00 out per million tokens; API prices raised 58% on release | Moonshot claims 58.6% on SWE-Bench Pro, beating every closed-source model, and 54.0% on Humanity’s Last Exam with tools; independently placed at 54 on the Artificial Analysis index, tied with DeepSeek V4 and Qwenlab-reported Source: Moonshot benchmark; The Batch (deeplearning.ai); Artificial Analysis; 36Kr · CAD issue |
| GLM-5.1 | Zhipu (Z.ai) | 8 Apr 2026 | 754B (MoE) per post-launch coverage; a pre-release third-party estimate cited 744B | n/d | Open-weight (MIT) | Priced close to Claude Sonnet after a launch increase; absolute per-token price not disclosed | Zhipu claims #1 globally on SWE-bench Pro, above GPT-5.4 and Opus 4.6; independent breakdowns put it at 94.6% of Opus 4.6lab-reported Source: Zhipu launch docs; WaveSpeed; tmtpost; Reuters · CAD issue |
Scroll the table sideways to see every column.
How to read this
Open weights are the through-line. DeepSeek V4, Zhipu's GLM-5.1 and GLM-5.2, Moonshot's Kimi K2.6 and K3, MiniMax M3 and Meituan's LongCat-2.0 all ship open weights, most under a permissive license. The lone closed flagship in the table is Alibaba's Qwen3.7-Max, which stays proprietary and API-only. When a lab cannot out-spend the US frontier on chips, giving the model away is how it sets the standard.
Price is the other weapon. DeepSeek V4-Pro lists at roughly one-seventh the price of Claude Opus 4.7; MiniMax sells about fifteen times the token volume of a comparable Claude plan for the same money; GLM-5.2 runs at roughly one-sixth the API cost of GPT-5.5. The models are close on capability and far cheaper to run, which is the entire distribution thesis.
The benchmarks are close but not settled. On independent leaderboards the best Chinese models now land between GPT-5.5 and the top Opus builds: on the Artificial Analysis index GLM-5.2 was, by Artificial Analysis's own reading, the first open-weight model to reach 51, and Kimi K3 topped the Frontend Code Arena. But the same independent reporting is clear that Chinese models still trail the top US labs on the hardest coding benchmarks, and the "we beat everyone" figures are the labs' own. For the chips these models run on, see the domestic chip scoreboard.
Who is actually using these models
Increasingly, American enterprises. An IDC survey published in late July 2026 found that 47% of decision-makers at US firms with more than a thousand employees had put a Chinese model into at least one use case, and about one in five said they use one heavily. At a point in mid-July, all five of the top models on OpenRouter, the marketplace that routes enterprise traffic to whichever model wins on price and performance, were Chinese. (IDC and OpenRouter, via The Next Web and China AI Dispatch, 28 Jul 2026.)
The named switchers are not lab experiments. Coinbase moved work onto Moonshot's Kimi and Zhipu's GLM and says it halved its AI spending; DoorDash routes lower-level work to Kimi; Airbnb runs customer service on Alibaba's Qwen; and Cursor, the coding tool, built its Composer 2 model on Kimi foundations. The pattern is consistent: keep an expensive US model for the hardest slice of work and hand the rest to a Chinese model that costs a fraction as much. (China AI Dispatch, "The Trade-Down," 28 Jul 2026.)
The cost gap is structural, not a promotion. UBS estimates the leading Chinese models cost roughly a tenth as much to train as comparable US systems, with API prices at 10 to 20% of the foreign alternative, reached through smaller and more sparsely activated mixture-of-experts designs, GPU utilization above 70% against an industry norm near 40%, and cheaper power. On one legal benchmark the spread was stark: about $2.40 per task for GLM-5.2 against about $31 for Anthropic's Fable, roughly thirteen to one. (UBS via TechNode; Applied Compute's Harvey case study; via China AI Dispatch.)
The caveat cuts both ways. Chinese commentators are not triumphant about it: Caixin argues the cheap-AI story runs into the heavy-industrial cost of actually serving trillion-parameter models at scale, and selling intelligence at cost is not a business. And both Washington and Beijing are now weighing export controls on advanced models and weights, so the open window in which a US company can download a Chinese model and run it in production may not stay open indefinitely. (Caixin and the Financial Times, via China AI Dispatch, 28 Jul 2026.)
New models land in the daily first
This tracker is updated from the newsletter, so the daily is always ahead of the page. When a Chinese lab ships a new flagship or open-sources a frontier model, that is where it breaks first.
Common questions
What are China’s frontier AI models?
The most-watched are DeepSeek V4, Alibaba's Qwen line, Zhipu's GLM, Moonshot's Kimi, and MiniMax, with Tencent's Hunyuan and Meituan's LongCat now shipping trillion-parameter models too. DeepSeek is the most influential; most of these labs release open weights, which is the strategy that sets them apart from the closed US frontier.
What is DeepSeek V4?
DeepSeek V4 is the lab's flagship, open-sourced under the MIT license in April 2026 with 1.6 trillion parameters and a 1-million-token context. The V4-Pro variant activates about 49 billion parameters per token, is priced at roughly $3.48 per million output tokens (about one-seventh of Claude Opus 4.7), and scores 80.6% on SWE-Bench Verified in independent testing, a hair behind the best Western models on coding.
Which Chinese AI models are open-weight?
Most of them. DeepSeek V4 (MIT), Zhipu’s GLM-5.1 and GLM-5.2 (MIT), Moonshot’s Kimi K2.6 (modified MIT) and K3 (open-weight), MiniMax M3, and Meituan’s LongCat-2.0 all ship open weights. Alibaba’s flagship Qwen3.7-Max is the notable exception: it is proprietary and API-only. Open weights are China’s distribution strategy, making the stack the cheapest and most legal one to run anywhere.
How good are Chinese models against US models?
Close, and the gap is a moving target measured release by release. GLM-5.2 was, by Artificial Analysis’s reading, the first open-weight model to land at 51 on its Intelligence Index, between GPT-5.5 and Opus 4.8. DeepSeek V4-Pro’s 80.6% on SWE-Bench Verified, against Opus 4.7’s 87.6%, was independently confirmed. Kimi K3 topped the Frontend Code Arena. But independent reporting notes Chinese models still trail the top US labs on the hardest coding benchmarks, and lab-reported "we beat everyone" numbers should be read as vendor claims.
Are US companies actually using Chinese AI models?
Yes, and increasingly in production rather than in tests. An IDC survey in late July 2026 found 47% of decision-makers at US firms with over a thousand staff had put a Chinese model into at least one use case, with about one in five using one heavily, and at one point in mid-July all five top models on OpenRouter were Chinese. Named adopters include Coinbase (Kimi and GLM, which it credits with halving its AI spend), DoorDash (Kimi for lower-level work), Airbnb (Qwen for customer service), and Cursor, whose Composer 2 model is built on Kimi foundations. The draw is cost: UBS estimates Chinese API prices run at 10 to 20% of the US alternative. Reported by IDC, OpenRouter and UBS, via China AI Dispatch (28 Jul 2026).
Are these models trained on Chinese chips?
Increasingly. Zhipu says GLM-5.1 was trained entirely on about 100,000 Huawei Ascend 910B chips with no Nvidia hardware, the first frontier-scale model to do so. Meituan reports LongCat-2.0 was trained and run entirely on a 50,000-card domestic cluster. DeepSeek V4 ships with Huawei Ascend as a primary deployment path. The training frontier is still harder than inference, but the trend is toward domestic silicon.
Sources and method
Every row traces to a named source carried in China AI Dispatch's reporting: DeepSeek, Zhipu, Moonshot, MiniMax and Alibaba's own technical reports and launch posts for specs and lab-reported benchmarks; Reuters, SCMP, InfoQ, tmtpost, 36Kr, Qbit and Geek Park for release facts; and Artificial Analysis, OpenRouter, Chatbot Arena, LMArena and Code Arena for independent leaderboard placements. Where a row links to a CAD issue, that issue cites the underlying primary source.
Two classes of figure are handled with care. Benchmark results are tagged by who produced them: a lab-reported tag means the number is the vendor's own claim and should be read as such; an independent tag means a third-party leaderboard. Parameter counts and context windows are only listed where the lab disclosed them; Alibaba does not publish Qwen's, so those cells read "n/d" rather than carrying an outside estimate. We track one flagship model per lab to keep the table legible, so a lab's earlier or lighter variants are not all listed here.
This is a living document, updated from the daily newsletter, which always carries the newest releases first.
Get the reporting behind this page
China AI Dispatch is a daily, English-language briefing on Chinese AI, built from Chinese-language sources most of the English press does not read. If this tracker was useful, the daily is where it comes from.
One email each morning. Unsubscribe anytime.