Moonshot has halted new subscriptions for Kimi K3, its latest AI model, after overwhelming demand consumed nearly all available GPU capacity within two days. The Chinese AI startup faced infrastructure strain so severe that it had to close the product to new customers.

The pause reflects a common problem facing AI companies launching capable models. Moonshot invested heavily in GPU infrastructure to support Kimi K3, but demand outpaced supply faster than anticipated. The 48-hour window before maxing out capacity suggests either extremely strong market interest or insufficient hardware provisioning.

The company plans to restructure its subscription model to distribute computational resources more evenly across users. This likely means tiering access levels or implementing usage caps rather than unlimited compute access. Such approaches let companies serve more subscribers on existing hardware, though users often experience slower response times or reduced usage allowances in return.

Moonshot competes in a crowded market where Chinese AI companies like Alibaba, Baidu, and Zhipu have already deployed capable models. Kimi K3 apparently offers features or performance compelling enough to attract thousands of customers immediately. The rapid uptake validates market demand for the model but exposes the capital intensity of AI deployment.

The subscription pause also signals a temporary revenue ceiling. Moonshot cannot onboard new paying customers until either GPU capacity expands or the revised subscription model reduces per-user resource requirements. Neither solution is quick.

This situation affects the broader AI infrastructure market. GPU shortages have become recurring bottlenecks as models improve and adoption accelerates. Nvidia's H100 and H200 chips remain expensive and limited. Companies without access to cutting-edge silicon must either wait for supply, pay premium prices, or redesign products to run on less powerful hardware.

Moonshot's move highlights an uncomfortable truth about AI products. Building a capable model solves only half the problem. Deploying it