TL;DR: Kimi is the model family from Moonshot AI, a Beijing lab, released with downloadable weights under the lab's own license and known for very long context windows and long-running agentic and coding work. You can download the weights and serve them yourself, call Moonshot's hosted API through OpenAI-compatible or Anthropic-compatible endpoints, or reach the same models through third-party inference providers.
What it is and how you use it
Moonshot AI was founded in Beijing in 2023 and made context length its identity before most of the field treated it as a headline feature: the Kimi chat assistant launched on long-document handling, and the models have kept that trait ever since. Under the hood the flagships are very large mixture-of-experts systems, where only a small fraction of the parameters on disk activate for each token - which is how a model measured in trillions of parameters stays affordable to serve. The K2 line carried the family through 2025 and the first half of 2026 in instruct, thinking and coding variants; the K3 generation that followed in mid-2026 is natively multimodal and ships with thinking on by default, with an effort setting to dial it down.
Three things define the family. The first is the context window: the vendor lists a million tokens for the current flagship, enough to hold a large codebase or a long document set in a single request. The second is reasoning - the thinking variants work through problems before answering, and the lab quotes the same agentic benchmark suites its Western peers lead their launches with, such as SWE-bench Verified, Terminal-Bench and BrowseComp, as vendor claims. The third is agentic work: Moonshot positions the models as engines for long-horizon tool use and ships its own open-source terminal coding agent, Kimi Code, alongside a desktop agent app.
The weights are on Hugging Face, but the license is Moonshot's own rather than Apache 2.0 or MIT. The K2 line shipped under a "Modified MIT" license that adds an attribution clause for very large deployments; K3 moved to a bespoke "Kimi K3 License" that requires products above a user or revenue threshold to display the model's name and requires a separate agreement for large model-as-a-service businesses. Smaller research releases - the Kimi-Linear, Kimi-VL and Kimi-Audio models - are plain MIT. Same rule as the other open families: read the license for the checkpoint you deploy, not the family's reputation.
Access comes in the usual three forms. Moonshot's platform API exposes an OpenAI-compatible Chat Completions endpoint, an Anthropic-compatible Messages endpoint and a Responses-style endpoint, so most existing SDKs and coding agents work against it with a base-URL change. Western inference providers such as Together and Fireworks host the open weights on their own hardware, and OpenRouter lists the flagship; hyperscaler coverage lags, with Microsoft Foundry carrying it through a partner provider while the other clouds stopped at earlier generations. Self-hosting means vLLM or SGLang on a multi-GPU node - the flagships are far too large for a workstation, and Ollama's Kimi entry routes to a cloud endpoint rather than running local weights. Where it fits: long-document and long-session agent workloads, and teams that want frontier-tier open weights with a choice of who serves them. The counterweights are the ones every China-based first-party API carries - jurisdiction and data governance - which is exactly why the provider path matters.
Where it sits in the AI stack
Kimi sits at the model layer as an open family: the weights are the product, and whether Moonshot, a provider or your own GPUs serve them is your decision:
Key tools and implementations
-
Open-weight downloads
The K-series flagships on Hugging Face under Moonshot's own license, with attribution and revenue thresholds; the smaller research models under MIT.
-
Moonshot platform API
The first-party endpoint, with OpenAI-compatible, Anthropic-compatible and Responses-style request shapes so existing tooling works with a base-URL change.
-
Kimi Code and the Kimi apps
An MIT-licensed terminal coding agent with MCP support and editor integration, plus the Kimi chat assistant and a desktop agent app.
-
Providers and serving runtimes
Together, Fireworks and OpenRouter serve the weights on Western infrastructure; vLLM and SGLang run them on self-managed multi-GPU nodes.
Related entries
- DeepSeek A Chinese AI lab known for open-weight models with strong reasoning and unusually low training costs.
- Qwen Alibaba's family of open-weight models spanning many sizes, strong in multilingual and coding tasks.
- GLM Z.ai's model family (formerly Zhipu AI), released mostly under MIT and positioned as a flat-rate engine behind popular coding agents.
- Context window The maximum amount of text, measured in tokens, that a model can consider in a single request.
- Open-weights models Models whose trained weights are published for anyone to download, run locally, and fine-tune.