Kimi K3 has received far more love than expected (1 minute read)
Summary
Kimi K3's popularity has strained GPU capacity, leading to a temporary pause on new subscriptions and a planned split into two membership plans: Kimi Membership for general use and Kimi Code Membership for coding workflows.
View Cached Full Text
Cached at: 07/20/26, 09:26 PM
Kimi K3 has received far more love than we expected, and our GPUs are feeling it.
Over the past 48 hours, demand has pushed close to the limits of our current capacity. To protect the experience of existing subscribers, we’re temporarily pausing new subscriptions and prioritizing compute for current members. Existing subscribed users are not affected.
We’re adding capacity as fast as we can and will reopen new subscription spots in batches.
Going forward, we’ll also split membership into two more focused plans: Kimi Membership for Kimi Web, App, and Work; and Kimi Code Membership for coding workflows. This will help us match compute more precisely and keep the experience stable.
Thank you for your patience and understanding!
Similar Articles
Kimi is temporarily pausing new subscriptions and prioritizing compute for current members due to surging demand.
Kimi is pausing new subscriptions to prioritize compute resources for existing members due to surging demand.
For those who didn't get to join Kimi, there is a waitlist now.
Kimi Code is an AI-powered coding assistant and CLI tool powered by Kimi K3 model. A waitlist is now available for those who missed initial access.
Kimi K3 API Pricing
Kimi K3 API pricing details announced.
Kimi K3 is now live
Kimi AI has launched K3, a new model built for agentic coding and knowledge work, now live on their platform.
@thealexker: underrated gems in Kimi-K3 release: > an early K3 wrote the majority of the kernels in the late development stages > it…
Kimi.ai released Kimi K3, a 2.8 trillion parameter multimodal model with 1 million context, featuring novel Delta Attention and Attention Residuals, and a self-optimizing stack including MiniTriton compiler. The model achieves up to 6.3x faster decoding and ~25% higher training efficiency.