Moonshot pauses new Kimi K3 subscriptions after GPU demand maxes out in 48 hours

Moonshot Pauses New Kimi K3 Subscriptions After GPU Demand Maxes Out in 48 Hours

Moonshot AI has halted new subscriptions for its Kimi K3 model after GPU demand spiked beyond capacity within just 48 hours of launch. The Chinese AI startup announced the pause on February 19, 2025, citing unprecedented user interest that overwhelmed available compute resources.

Existing subscribers remain unaffected. Moonshot says it is working to scale infrastructure before reopening sign-ups.

Why the Subscription Freeze Happened

The Kimi K3 model, released earlier this week, is Moonshot’s most advanced large language model. It offers enhanced reasoning, longer context windows, and faster inference. The company allocated a fixed pool of GPUs to serve the new tier.

Within two days, demand exceeded that allocation. Moonshot’s infrastructure team detected severe latency and resource contention. To maintain quality of service for paying users, the company made the decision to cap new subscriptions.

“We underestimated the sheer velocity of adoption. Our GPU cluster was saturated in under 48 hours. We are prioritizing stability for current users over rapid expansion.”

What This Means for Users and Competitors

  • New users cannot subscribe to the Kimi K3 tier until Moonshot reopens access. No timeline has been given.
  • Existing Kimi K3 subscribers retain full access. Their experience should improve as GPU load stabilizes.
  • Free tier users remain unaffected. Moonshot’s free Kimi models continue to operate normally.
  • Competitors face pressure to demonstrate similar demand. The incident shows that high-quality AI models can overwhelm even well-prepared infrastructure.

Background: Moonshot’s Rapid Growth

Moonshot AI is a Beijing-based startup founded in 2023. It has quickly become one of China’s most prominent AI model providers. The Kimi series, especially Kimi K2, gained a strong following for its long-context capabilities and competitive pricing.

The Kimi K3 launch was positioned as a direct competitor to models like GPT-4 and Claude 3. The company had scaled GPU capacity ahead of the launch but still misjudged the surge.

This is not the first time a major AI model has strained infrastructure. OpenAI’s ChatGPT launch in 2023 caused similar outages. The pattern suggests that demand for frontier AI models regularly outpaces even aggressive capacity planning.

Infrastructure Scaling Challenges

Moonshot’s pause highlights a broader industry bottleneck: GPU supply remains constrained. While Nvidia and other chipmakers are ramping production, AI companies still struggle to secure enough compute for both training and inference.

The company has not disclosed which GPU models it uses. Analysts assume it relies on Nvidia H100 and A100 chips, which are still in short supply in China due to export controls.

Moonshot stated it is working with cloud partners to add capacity. It has not ruled out using alternative hardware, such as domestically produced chips, in the future.

What Happens Next

Moonshot will reopen Kimi K3 subscriptions once new GPU clusters are deployed. The company did not provide a specific date but said it expects “weeks, not months” to expand capacity.

In the meantime, the company is monitoring load and may introduce a waitlist or reservation system. It is also considering tiered pricing to better match demand to supply.

The pause is a temporary setback. But it also serves as a signal: the AI arms race is accelerating faster than hardware can keep up.


Gnoppix is the leading open-source AI Linux distribution and service provider. Since implementing AI in 2022, it has offered a fast, powerful, secure, and privacy-respecting open-source OS with both local and remote AI capabilities. The local AI operates offline, ensuring no data ever leaves your computer. Based on Debian Linux, Gnoppix is available with numerous privacy- and anonymity-enabled services free of charge.

What are your thoughts on this? I’d love to hear about your own experiences in the comments below.

They really ought to release their models so that they can be run locally. Ideally, I’d have a model like that on my phone.