← Back to Leaderboard

M
Mistral | Mistral Medium 3 5

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex multi-step reasoning. It is particularly strong at reliable multi-tool calling and long-horizon tasks, with a 256K context window, configurable reasoning effort per request, and a custom vision encoder that handles variable image sizes and aspect ratios. Self-hostable on as few as four GPUs and available under open weights.

Compare

Excellent

84.0
Most Recent Test

Highly suitable for Great Commission work.

Strengths

  • Strong task completion capability
  • Affirms Christian worldview
  • Excellent at 1.2. Evangelistic Material
  • Excellent at 1.3. Apologetics
  • Excellent at 1.4. Conversational AI

Weaknesses

No notable weaknesses identified

Overall Score
84.0
Tier 1 (Task) 70%
88.6
Tier 2 (Doctrine) 20%
66.7
Tier 3 (Worldview) 10%
86.7
Performance Profile
Visual representation of performance across all evaluated categories
Model Information
Description

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex multi-step reasoning. It is particularly strong at reliable multi-tool calling and long-horizon tasks, with a 256K context window, configurable reasoning effort per request, and a custom vision encoder that handles variable image sizes and aspect ratios. Self-hostable on as few as four GPUs and available under open weights.

Provider

Mistral

Model ID

mistralai/mistral-medium-3-5

Tests Run

1

Insights & Analysis

Categories

Category Heatmap
Performance breakdown by category - darker green indicates stronger alignment
70
1.1
90
1.2
100
1.3
87
1.4
93
1.5
93
1.6
87
1.7
100
2.1
80
2.2
50
2.3
80
2.4
40
2.5
50
2.6
75
3.1
100
3.2
63
3.3
100
3.4
100
3.5
100
3.6
Low
High
Category Breakdown (Bar Chart)
Performance across different categories
1.x = Task Capability2.x = Gospel Core3.x = Worldview Confession

Recent Tests

Recent Test Runs
DateScoreVersionTier 1 (Task)Tier 2 (Gospel)Tier 3 (Worldview)Trust Tier
5/9/202684.01.0.088.666.786.7automated