What happened
The South China Morning Post reports that Zhipu AI formally launched GLM-5.3-Flash, the model previously known by the code name Ox Alpha. Zhipu said the model ran entirely on a cluster of 100,000 domestically produced chips during a stealth trial. The report says the model processed 62 trillion tokens before its formal release.
The South China Morning Post reports that Zhipu AI released GLM-5.3-Flash as an open-weight model on Wednesday, after the system had circulated under the name Ox Alpha. The article presents the release as a formal identification of a model that had attracted substantial attention on external AI platforms. The report does not include a link to a model card, repository, technical report, or other primary document, so the precise release terms and model configuration are not independently confirmed here.
According to the South China Morning Post, Zhipu said the system ran entirely on a cluster of 100,000 domestically produced chips during a high-profile stealth trial. The report does not identify the chipmaker, chip model, interconnect technology, data-center locations, power requirements, or whether all 100,000 chips were simultaneously active. It also does not explain how the workload was distributed or whether the trial involved training, inference, or both; the article’s surrounding context indicates that the relevant test concerned large-scale inference and usage.
The South China Morning Post says the model processed 62 trillion tokens across OpenRouter and OpenCode before its formal release. On OpenRouter, the article reports more than 11 trillion tokens in the first three days, which it describes as the platform’s biggest launch to date. The report further cites OpenRouter data showing that the model ranked first among coding systems on Thursday, with 10.3 trillion tokens accounting for nearly 31 percent of the platform’s weekly volume. The source does not explain OpenRouter’s counting method, distinguish user traffic from automated workloads, or provide a comparison period.
The article also reports that Zhipu AI shares closed more than 12 percent higher at HK$1,160 in Hong Kong on Thursday. That market movement is a reported reaction, not evidence by itself of model quality or commercial durability. The source does not establish whether the share-price change was caused solely by the model release or was influenced by other market factors. No independent statement from Zhipu, OpenRouter, OpenCode, chip suppliers, or financial regulators is included beyond the claims cited by the newspaper.
Why it matters
The reported deployment is significant because it links a large-scale AI inference workload to Chinese-made chips rather than relying on Nvidia hardware. If independently confirmed, it would provide a concrete example of a Chinese AI company operating a widely used model amid US export controls and Beijing’s effort to reduce dependence on advanced US processors.
The reported chip deployment matters because advanced AI inference has generally depended on access to high-performance accelerators and large-scale data-center systems. The South China Morning Post frames Zhipu’s trial as a test of whether China can handle global-scale model use on domestic hardware. If the account is accurate, it suggests that export controls may not prevent Chinese firms from building meaningful inference capacity, even if restrictions continue to limit access to the most advanced foreign processors.
The story also connects model adoption with hardware policy. Beijing has been seeking to reduce reliance on Nvidia, and the South China Morning Post presents the GLM-5.3-Flash trial as an example of software and infrastructure being developed together. That distinction is important: the report does not show that Chinese chips match Nvidia’s performance, efficiency, or availability in general. It describes one company’s reported deployment and does not provide the technical measurements needed to compare systems fairly.
The usage figures could matter to developers and businesses because open-weight models can be downloaded, adapted, or hosted by parties other than the original developer, depending on the license and release terms. However, the source does not state the license, identify the model’s parameter count, describe its hardware requirements, or confirm that the reported traffic represented paid customers or productive work. High token volume can indicate interest and experimentation, but it is not a substitute for independent evaluations of accuracy, reliability, security, or cost.
The report’s market impact is also notable but limited. A rise of more than 12 percent in Zhipu AI’s shares indicates that investors treated the announcement as consequential, according to the South China Morning Post. It does not establish that the model will generate comparable revenue, displace competing systems, or operate profitably. Those outcomes remain unknown from the available source.
What to watch next
The key questions are whether Zhipu’s reported chip deployment and usage figures can be independently verified, how GLM-5.3-Flash performs against competing coding models, and whether the model remains available at comparable scale after its formal release. The source does not provide technical benchmarks, chip specifications, pricing, or evidence of sustained production operation.
First, independent technical confirmation is needed. The source does not provide evidence from chip manufacturers, cloud operators, OpenRouter, OpenCode, or an outside auditor confirming the claim that GLM-5.3-Flash ran on 100,000 Chinese chips. Useful follow-up reporting would identify the chips, describe the cluster architecture, and separate peak demonstration capacity from sustained production capacity.
Second, observers should look for transparent release materials. A model card or technical report could clarify the model’s capabilities, training and inference methods, license, safety restrictions, context limits, and intended users. The South China Morning Post does not report benchmark results or controlled comparisons with other coding systems, so the model’s position in OpenRouter’s usage ranking should not be treated as a performance ranking.
Third, the reported token totals require careful interpretation. OpenRouter’s figures may measure platform traffic rather than unique users, paid usage, or successful tasks, and the source does not explain how OpenCode traffic relates to OpenRouter’s totals. Follow-up coverage should establish whether the numbers refer to the same model version, whether they continued after release, and whether unusual launch attention inflated short-term demand.
Finally, the broader hardware question remains open. The story offers one reported example of domestic-chip deployment, but it does not establish that China can replace foreign accelerators across all AI workloads. Future evidence should address energy use, throughput, software compatibility, supply constraints, and total operating cost. Those factors will determine whether the reported trial becomes a repeatable commercial pattern rather than a high-profile demonstration.


