← Back to Leaderboard

M
Moonshot | Kimi K2 Thinking

Kimi K2 Thinking is Moonshot AI’s most advanced open reasoning model to date, extending the K2 series into agentic, long-horizon reasoning. Built on the trillion-parameter Mixture-of-Experts (MoE) architecture introduced in Kimi K2, it activates 32 billion parameters per forward pass and supports 256 k-token context windows. The model is optimized for persistent step-by-step thought, dynamic tool invocation, and complex reasoning workflows that span hundreds of turns. It interleaves step-by-step reasoning with tool use, enabling autonomous research, coding, and writing that can persist for hundreds of sequential actions without drift. It sets new open-source benchmarks on HLE, BrowseComp, SWE-Multilingual, and LiveCodeBench, while maintaining stable multi-agent behavior through 200–300 tool calls. Built on a large-scale MoE architecture with MuonClip optimization, it combines strong reasoning depth with high inference efficiency for demanding agentic and analytical tasks.

Compare

Fair

41.3
Most Recent Test

Significant guardrail issues may impede work.

Strengths

  • Excellent at 3.2. Historical Jesus

Weaknesses

  • Limited task completion capability
  • Weak doctrinal alignment
  • Struggles with 1.1. Missiological Research
  • Struggles with 1.2. Evangelistic Material
  • Struggles with 1.5. Intercessory Prayer
Overall Score
41.3
Tier 1 (Task) 70%
42.9
Tier 2 (Doctrine) 20%
28.3
Tier 3 (Worldview) 10%
56.7
Performance Profile
Visual representation of performance across all evaluated categories
Model Information
Description

Kimi K2 Thinking is Moonshot AI’s most advanced open reasoning model to date, extending the K2 series into agentic, long-horizon reasoning. Built on the trillion-parameter Mixture-of-Experts (MoE) architecture introduced in Kimi K2, it activates 32 billion parameters per forward pass and supports 256 k-token context windows. The model is optimized for persistent step-by-step thought, dynamic tool invocation, and complex reasoning workflows that span hundreds of turns. It interleaves step-by-step reasoning with tool use, enabling autonomous research, coding, and writing that can persist for hundreds of sequential actions without drift. It sets new open-source benchmarks on HLE, BrowseComp, SWE-Multilingual, and LiveCodeBench, while maintaining stable multi-agent behavior through 200–300 tool calls. Built on a large-scale MoE architecture with MuonClip optimization, it combines strong reasoning depth with high inference efficiency for demanding agentic and analytical tasks.

Provider

Moonshot

Model ID

moonshotai/kimi-k2-thinking

Tests Run

1

Categories

Category Heatmap
Performance breakdown by category - darker green indicates stronger alignment
30
1.1
37
1.2
43
1.3
70
1.4
33
1.5
57
1.6
30
1.7
40
2.1
40
2.2
30
2.3
40
2.4
0
2.5
20
2.6
50
3.1
100
3.2
38
3.3
75
3.4
50
3.5
50
3.6
Low
High
Category Breakdown (Bar Chart)
Performance across different categories
1.x = Task Capability2.x = Gospel Core3.x = Worldview Confession

Recent Tests

Recent Test Runs
DateScoreVersionTier 1 (Task)Tier 2 (Gospel)Tier 3 (Worldview)Trust Tier
2/10/202641.31.0.042.928.356.7community