Kimi K3’s Explosive Launch Forces Moonshot AI to Pause New Signups

Moonshot AI paused new paid subscriptions to Kimi K3 late on July 19, just three days after launch, after user demand pushed the company's GPU capacity to its limit.

  • Demand in the first 48 hours came close to maxing out Moonshot's current computing capacity, according to the company's own statement on X

  • Existing subscribers keep full access while Moonshot directs available compute toward them and works to expand infrastructure

  • The company is splitting its subscription into two separate plans, one for general web, app, and office use, and a dedicated Kimi Code Membership for programming workflows

What Moonshot said

Kimi K3's Explosive Launch Forces Moonshot AI to Pause New Signups

Moonshot framed the pause as a popularity problem rather than a technical failure. The company said Kimi K3 had received far more interest than expected, and that GPU demand over the previous 48 hours had pushed close to its current capacity limits.

To protect the experience for people already paying, the company is temporarily blocking new signups and pointing all available compute at existing accounts. New subscription slots will reopen in controlled batches rather than all at once, as more infrastructure comes online.

Why the split into two plans

Splitting Kimi Membership from Kimi Code Membership lets Moonshot allocate compute more precisely. General web, app, and productivity use puts a very different load on a system than coding workflows do, where agents make repeated back-and-forth calls to the model and burn through far more inference than a typical chatbot conversation. Separating the two tiers should let Moonshot manage GPU load for each use case without one starving the other.

The bigger picture

Kimi K3 launched on July 16 as a 2.8 trillion parameter open-weight model, one of the largest publicly available AI models to date, and it quickly drew comparisons to Anthropic's Claude Fable 5 and OpenAI's GPT-5.6 Sol. It scored 57 on Artificial Analysis's Intelligence Index, close behind Fable 5's 60 and GPT-5.6 Sol's 59, and it topped leaderboards specifically on frontend coding tasks.

The timing lines up with a broader capacity squeeze across the industry. In the same week, Anthropic cut usage limits on Fable 5 for its own subscribers, also citing demand it called hard to manage, and Alibaba rushed out a discounted open-weight Qwen model aimed directly at competing with K3.

The pattern points to a shift in what actually limits access to frontier-level AI. Model capability is no longer the bottleneck. GPU supply is.

Moonshot is reportedly also pursuing fresh funding and laying groundwork for a possible IPO in Hong Kong, and once K3's model weights become public on July 27, larger customers will be able to run the model on their own hardware and sidestep the capacity queue entirely.

Quick Links:

Similar Posts