Kimi K2 Thinking

Kimi K2 Thinking is the latest, most capable version of open-source thinking model. Starting with Kimi K2, we built it as a thinking agent that reasons step-by-step while dynamically invoking tools. It sets a new state-of-the-art on Humanity's Last Exam (HLE), BrowseComp, and other benchmarks by dramatically scaling multi-step reasoning depth and maintaining stable tool-use across 200–300 sequential calls. At the same time, K2 Thinking is a native INT4 quantization model with 256k context window, achieving lossless reductions in inference latency and GPU memory usage.

Context
262K
Max output
100K
Input
$0.60
Output
$2.5
DECISION SUMMARY

Recommended use cases

Strengths in this dataset

  • 262,144-token context window
  • text input
  • 17 supported API parameters listed

Limits and caveats

  • Provider behavior and pricing can change; verify the linked sources before production use.
CAPABILITIES

Capability

Model-native facts

Model
Reasoning
Supported
Open weights
Unknown

Provider endpoint facts

Provider endpoint
Tool calling
Supported
Structured output
Supported
Streaming
Unknown
Prompt cache
Supported
Batch
Unknown
Fine-tuning
Unknown
PROVIDER PRICING

Kimi K2 Thinking Provider pricing

Provider endpoint: moonshotai/kimi-k2-thinking

Input
$0.60
per 1M tokens
Output
$2.5
per 1M tokens
Cached input
$0.15
per 1M tokens
Image output
Unknown
per 1M tokens
SOURCE RECORDS

Sources and verification

Hugging Face model card

Fields: tags, gated, license, summary, library name, pipeline tag

OpenRouter Models API

Fields: identity, description, modalities, context window, maximum output, pricing, supported parameters

MODEL FAQ

Frequently asked questions

Answers are generated from the same sourced model and provider facts shown above.

Kimi K2 Thinking is the latest, most capable version of open-source thinking model. Starting with Kimi K2, we built it as a thinking agent that reasons step-by-step while dynamically invoking tools. It sets a new state-of-the-art on Humanity's Last Exam (HLE), BrowseComp, and other benchmarks by dramatically scaling multi-step reasoning depth and maintaining stable tool-use across 200–300 sequential calls. At the same time, K2 Thinking is a native INT4 quantization model with 256k context window, achieving lossless reductions in inference latency and GPU memory usage.

Model specifications and prices may vary by provider and change over time. AIToolly displays sources and verification dates so users can confirm critical details before production use.