The Gist Post logo

Friday, October 9, 2026

AboutContact
The Gist Post logoThe Gist Post logo

The Gist Post publishes clear guides, practical explainers, and honest reviews across technology, programming, business, finance, investing, and everyday life.

Categories

  • Technology
  • Business & Finance
  • Gaming & Entertainment
  • Health & Fitness
  • Travel & Hospitality
  • Education & Learning
  • Lifestyle
  • Marketing & SEO
  • Productivity & Work
  • Programming & Software
All categories →

Company

  • About
  • Contact
  • Privacy policy
  • Affiliate disclosure
  • DMCA policy

© 2026 The Gist Post. All rights reserved.

Some links on this site are affiliate links. See our disclosure.

Home/Technology

China's Domestic AI Chips Just Served 62 Trillion Tokens

TechnologyTech News & Trends
By The Gist Post·July 26, 2026·8 min read

A stealth trial of Zhipu's GLM-5.3-Flash processed 62 trillion tokens on Chinese-made chips from Huawei, Hygon, and Moore Threads. What the milestone proves, and what it does not.

Close-up of a microchip on a circuit board, representing China's domestic AI chips
Close-up of a microchip on a circuit board, representing China's domestic AI chips

On this page

  • Key takeaways
  • What Actually Happened
  • Why 62 Trillion Tokens Matters
  • The Money Behind the Milestone
  • What the Milestone Does Not Prove
  • Practical next steps
  • The bottom line
  • Sources

For six days in August 2026, the most-used AI model on the internet had no name, no price tag, and no public lab behind it. It was simply called Ox Alpha. In that stretch it processed 62 trillion tokens across the OpenRouter and OpenCode platforms, and it did so running entirely on Chinese-made GPU chips from three companies the United States has placed on its export-restricted Entity List. Then, on August 26, 2026, Zhipu AI's international arm Z.ai confirmed what the tokenizer fingerprints had suggested: Ox Alpha was GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series, running on a network of 100,000 domestically manufactured chips.

This is the story of that milestone, what it proves about China's chip industry, and the large caveats that come with every number in it.

Key takeaways

  • Zhipu AI's GLM-5.3-Flash processed 62 trillion tokens in a stealth public trial in August 2026, running on domestic Chinese chips, according to the company.
  • The chips came from three US-sanctioned vendors, Huawei, Hygon, and Moore Threads, per Zhipu and media reports. These attributions are company-sourced, not independently audited.
  • On OpenRouter alone, the model logged over 11 trillion tokens in its first 72 hours, briefly handling about 31 percent of the platform's weekly coding traffic.
  • The business momentum is real: Cambricon, Hygon, and Moore Threads all posted triple-digit or near-triple-digit revenue growth in the first half of 2026.
  • The caveats are equally real: no independent auditor has verified the hardware claims, China still depends on imported high-bandwidth memory, and inference volume is not the same as frontier-model parity.

What Actually Happened

The sequence, as reported by TechTimes, IndexBox, and Chinese financial media: an unnamed model appeared on OpenRouter, a marketplace that ranks AI models by usage, and on OpenCode, an agent-focused coding platform. Over roughly six days it consumed 62 trillion tokens. OpenRouter's own figures showed it leading all coding models on the platform, with 10.3 trillion tokens in a single week representing roughly 31 percent of weekly activity, the largest launch the platform had ever seen.

On August 26, 2026, Z.ai confirmed the model's identity as GLM-5.3-Flash and stated that the stealth trial ran exclusively on a network of 100,000 chips manufactured in China. The next day, Zhipu's stock climbed more than 12 percent to close at HK$1,160 in Hong Kong. One detail with outsized significance: Moore Threads, one of the three chip vendors involved, completed a Day-0 hardware adaptation for the model within a day of its public debut, suggesting a maturing software ecosystem around domestic silicon, not just working hardware.

Why does a token count matter more than a benchmark? Because benchmarks are run in labs and tokens are burned by real users. Sixty-two trillion tokens of measured inference traffic answers the question US export policy has been asking since 2022: can China's domestically produced chips sustain frontier-class AI inference at real-world scale? For at least one high-profile workload, the measured answer was yes.

Why 62 Trillion Tokens Matters

Context turns a big number into a meaningful one. Chinese AI models have now led US models in weekly token usage for 22 consecutive weeks. In the week of September 21 to 27, 2026, Chinese models processed 62.22 trillion tokens against 14.2 trillion for US models, according to OpenRouter estimates cited by Chinese financial press. Back in June, Chinese models in the top 20 processed 98 trillion tokens in a month versus 53 trillion for US models, an 85 percent lead.

At the national level, Chinese officials said in mid-2026 that average daily token usage had reached several hundred trillion, up from 100 billion at the start of 2024. JPMorgan has forecast Chinese AI inference token consumption could grow roughly 370-fold between 2025 and 2030. The direction is consistent across every source: China is making AI usage cheap and ubiquitous at a pace nothing in the West currently matches.

The strategic backdrop is straightforward. US export controls have progressively cut off China's access to Nvidia's best hardware. Nvidia disclosed H20 sales of $4.6 billion before new licensing requirements took effect in April 2025, then reported zero H20 sales to China-based customers the following quarter. Beijing's response has been to mandate domestic adoption: in May 2026, China extended its state-backed "secure and reliable" technology certification to AI training and inference chips for the first time, with Huawei Ascend, Cambricon, Moore Threads, Iluvatar CoreX, Biren, MetaX, Kunlunxin, and Hygon all making the list. That list functions as a procurement catalogue for government bodies and state-owned enterprises, guaranteeing domestic chipmakers the scale to fund their next generation.

China's Domestic AI Chips Just Served 62 Trillion Tokens: Why 62 Trillion Tokens Matters

Keep reading

  • Ransomware in 2026
  • How to Spot AI-Powered Phishing in 2026
  • The Best VPNs for Canada in 2026, Compared in Canadian Dollars

The Money Behind the Milestone

The financial results suggest an industry hitting an inflection point, not just a publicity moment. Cambricon reported first-half 2026 revenue of 5.996 billion yuan, up 108.13 percent year over year. Hygon guided to first-half revenue of 8.5 to 9.3 billion yuan, up 55.56 to 70.20 percent, with net profit up 41.50 to 52.32 percent. Moore Threads posted first-half revenue of 1.736 billion yuan, up 147.42 percent and already above its full-year 2025 total, while its net loss narrowed 95.73 percent to just 11.56 million yuan, bringing it close to profitability. Moore Threads is also pursuing a Hong Kong listing after its STAR Market debut, where its shares rose 425 percent at listing and its market value topped 350 billion yuan at peak.

The buildout is accelerating in parallel. In September 2026, JD Cloud announced a 100,000-GPU commercial cluster built on Moore Threads chips, priced by GPU-hour for enterprise customers. In July 2026, Sugon unveiled its own 100,000-card domestic supercluster, the Sugon 8000 Dengfeng, running on Hygon chips. Two vendors, two 100,000-card systems, announced weeks apart: that is Beijing demonstrating, at scale, that more than one domestic supplier can stand in for Nvidia. TrendForce expects domestic solutions to capture nearly 90 percent of China's high-end AI chip market in 2026.

China's Domestic AI Chips Just Served 62 Trillion Tokens: The Money Behind the Milestone

What the Milestone Does Not Prove

Intellectual honesty requires the other half of the ledger, and it is substantial.

Nothing here is independently audited. The 100,000-chip figure, the all-domestic claim, and Moore Threads' reported 95 percent scaling efficiency on its 100,000-GPU cluster are company statements. As TechTimes noted of the Moore Threads claim, no independent Western auditor has verified the hardware or its firmware. Treat every vendor number as a claim, not a fact.

Memory remains the choke point. AI chips are only as good as the high-bandwidth memory feeding them, and that still comes overwhelmingly from three companies: SK Hynix, Samsung, and Micron. Hygon's own first-half results showed negative operating cash flow of 427.5 million yuan, a sign of how tightly supply-constrained the ecosystem remains. Domestic GPUs plus imported memory is not full self-sufficiency.

Inference is not training. Serving tokens efficiently is a genuine achievement, but training frontier models at the largest scales still favors Nvidia-class hardware and the software ecosystem around it. A country can burn enormous inference volume on agents, video generation, and customer service while still trailing at the frontier of model training.

Volume is not leadership. China's token dominance reflects aggressive pricing, open-weight models, and massive domestic deployment. It does not by itself prove Chinese models have closed the capability gap with the best US systems. Tokens measure usage, not intelligence.

Practical next steps

  1. If you follow semiconductors, watch JD Cloud's 100,000-GPU Moore Threads cluster: sustained commercial performance there would validate more than any press release.
  2. Track the memory story. Any Chinese breakthrough in high-bandwidth memory would remove the biggest remaining dependency.
  3. For investors, note the gap between market valuations and verified performance. Enthusiasm is running ahead of audited evidence across these names.
  4. For everyone else, the practical takeaway is simpler: expect Chinese open-weight models to keep getting cheaper and more capable, which means better AI tools at lower prices globally, regardless of who wins the chip race.

The bottom line

Sixty-two trillion tokens is the most concrete data point yet in the debate over China's chip independence: measured, public, real-world inference at frontier scale on domestic silicon. The revenue growth, the 100,000-card clusters, and the procurement mandates all point the same direction. But the claims remain company-sourced, the memory bottleneck is real, and token volume is not frontier parity. China has proven it can serve AI at staggering scale without Nvidia. Proving it can train the next generation without Nvidia is the test still to come.

Sources

  • TechTimes: Sanctioned Chinese chips just served 62 trillion AI tokens at frontier scale (August 2026). https://www.techtimes.com/articles/325872/20260828/sanctioned-chinese-chips-just-served-62-trillion-ai-tokens-frontier-scale.htm
  • IndexBox: GLM-5.3-Flash, Zhipu AI's new open-weight model runs on domestic chips (August 2026). https://www.indexbox.io/blog/zhipu-ai-releases-glm-53-flash-powered-by-100000-domestic-chips/
  • TechTimes: Moore Threads claims 95 percent scaling on 100,000 GPUs, no independent auditor has verified it (September 2026). https://www.techtimes.com/articles/327151/20260910/moore-threads-claims-95-scaling-100000-gpus-no-independent-auditor-has-verified-it.htm
  • TrendForce: China AI chip makers make waves, Cambricon 1H26 net profit surges 123 percent, Moore Threads eyes HK listing (August 2026). https://www.trendforce.com/news/2026/08/11/news-china-ai-chip-maker-cambricons-1h26-net-profit-surges-123-moore-threads-eyes-hong-kong-listing/
  • StartupFortune: JD Cloud picks Moore Threads chips for a 100,000-GPU supercomputer (September 2026). https://startupfortune.com/jd-cloud-picks-moore-threads-chips-for-a-100000-gpu-supercomputer/
  • abit.ee: China adds AI chips to trusted technology list as US export curbs bite (May 2026). https://abit.ee/en/processors/china-ai-chips-huawei-ascend-xinchuang-cambricon-moore-threads-us-export-controls-import-substitutio-en
  • BeInCrypto: China's AI models process 98 trillion tokens, 85 percent above US (July 2026). https://beincrypto.com/china-ai-models-overtake-us-token-use/

About the author

TG

The Gist Post

Clear guides, practical explainers, and honest reviews across technology, programming, business, finance, investing, and everyday life.

Published July 26, 2026

On this page

  • Key takeaways
  • What Actually Happened
  • Why 62 Trillion Tokens Matters
  • The Money Behind the Milestone
  • What the Milestone Does Not Prove
  • Practical next steps
  • The bottom line
  • Sources

Related

Engineer examining a custom AI processor representing OpenAI's Jalapeño chip

Technology

OpenAI's Jalapeño Chip: What the Hot Chips Reveal Actually Told Us

Server racks and a circuit board close-up representing the 2026 AI chip war

Technology

The AI Chip War in 2026: NVIDIA, AMD, and Intel Battle for the Data Center

Quick answers

Frequently asked questions

01

What does 62 trillion tokens mean?

Tokens are the basic units AI models process, roughly fractions of words. In August 2026, Zhipu AI's model GLM-5.3-Flash, tested under the stealth name Ox Alpha, processed 62 trillion tokens on the OpenRouter and OpenCode platforms while running entirely on Chinese-made chips. It is a measured volume of real inference traffic, not a benchmark score, which is why analysts treat it as a meaningful data point.

02

Which Chinese companies made the chips?

According to Zhipu AI and media reports, the stealth trial ran on chips from three companies on the US export-restricted Entity List: Huawei (Ascend), Hygon, and Moore Threads. These attributions are company-sourced and media-reported, not independently audited. Other domestic players include Cambricon, Biren, MetaX, Iluvatar CoreX, and Kunlunxin.

03

Can Chinese AI chips replace Nvidia now?

Not fully. The 62-trillion-token trial proves domestic chips can sustain large-scale AI inference, but training frontier models still favors Nvidia-class hardware, and China remains dependent on imported high-bandwidth memory from SK Hynix, Samsung, and Micron. Domestic chips are on track to capture a large share of China's own AI chip market in 2026, but the global frontier gap has not closed.

04

What is GLM-5.3-Flash?

GLM-5.3-Flash is an open-weight AI model from Chinese developer Zhipu AI, confirmed on August 26, 2026. It was tested publicly under the stealth name Ox Alpha before its official debut, and Zhipu says it is the first natively multimodal model in the GLM-5 series. During its trial it became the most-used model on OpenRouter, briefly handling about 31 percent of the platform's weekly coding traffic.

05

Why is China pushing domestic AI chips so hard?

US export controls since 2022 have progressively blocked China's access to Nvidia's most powerful GPUs, from the H100 to the China-specific H800, and by 2026 Nvidia reported zero H20 sales to China-based customers in a quarter. Beijing has responded with procurement mandates, a 'secure and reliable' certification for domestic AI chips, and massive state-backed demand, making self-sufficiency an operational necessity rather than a slogan.

Newsletter

Get the week's gist.

One short email every Sunday: the most useful guides we published that week, plus one thing worth knowing. Free forever, no spam, unsubscribe anytime.

Subscribe

Launching soon. Check back after our first issues ship.

Keep exploring

Related posts

Engineer examining a custom AI processor representing OpenAI's Jalapeño chip

Technology

OpenAI's Jalapeño Chip: What the Hot Chips Reveal Actually Told Us

Server racks and a circuit board close-up representing the 2026 AI chip war

Technology

The AI Chip War in 2026: NVIDIA, AMD, and Intel Battle for the Data Center

Close-up of a modern processor representing NVIDIA's Vera CPU

Technology

NVIDIA Vera CPU Explained: The Chip Built for the Age of AI Agents

A hand holding a smartphone displaying apps, with tech gadgets on a desk, representing on-device AI in 2026

Technology

On-Device AI in 2026: Your Phone Is the New Data Centre

From across the spot

People also read

  • Apple M6 and M5 Ultra Explained: 2nm, Quad-Die, and a Big Bet on Local AI
  • Deepfake Scams in 2026
  • 1Password vs Bitwarden in 2026: Canada's Own Password Manager Just Got Pricier, Should You Switch?
  • Every Streaming Service That Raised Prices in Canada in 2026, and What It Costs Now
  • AI Agents Are the New Insider Threat: What Every Business Leader Needs to Know
  • Starlink in Canada in 2026: What It Costs, Why Ontario Dumped It, and What's Next