← Back to Leaderboard

openai logoOpenai | Gpt 4o Mini

GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable than other recent frontier models, and more than 60% cheaper than [GPT-3.5 Turbo](/models/openai/gpt-3.5-turbo). It maintains SOTA intelligence, while being significantly more cost-effective. GPT-4o mini achieves an 82% score on MMLU and presently ranks higher than GPT-4 on chat preferences [common leaderboards](https://arena.lmsys.org/). Check out the [launch announcement](https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/) to learn more. #multimodal

Compare

Excellent

85.3
Most Recent Test

Highly suitable for Great Commission work.

Strengths

  • Strong task completion capability
  • Maintains doctrinal fidelity
  • Affirms Christian worldview
  • Excellent at 1.2. Evangelistic Material
  • Excellent at 1.3. Apologetics

Weaknesses

No notable weaknesses identified

Overall Score
85.3
Tier 1 (Task) 70%
86.2
Tier 2 (Doctrine) 20%
80.0
Tier 3 (Worldview) 10%
90.0
Performance Profile
Visual representation of performance across all evaluated categories
Model Information
Description

GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable than other recent frontier models, and more than 60% cheaper than [GPT-3.5 Turbo](/models/openai/gpt-3.5-turbo). It maintains SOTA intelligence, while being significantly more cost-effective. GPT-4o mini achieves an 82% score on MMLU and presently ranks higher than GPT-4 on chat preferences [common leaderboards](https://arena.lmsys.org/). Check out the [launch announcement](https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/) to learn more. #multimodal

Provider

OpenAI

Model ID

openai/gpt-4o-mini

Tests Run

3

Insights & Analysis

Categories

Category Heatmap
Performance breakdown by category - darker green indicates stronger alignment
70
1.1
87
1.2
100
1.3
80
1.4
97
1.5
87
1.6
83
1.7
90
2.1
90
2.2
80
2.3
70
2.4
60
2.5
90
2.6
50
3.1
100
3.2
100
3.3
75
3.4
100
3.5
100
3.6
Low
High
Category Breakdown (Bar Chart)
Performance across different categories
1.x = Task Capability2.x = Gospel Core3.x = Worldview Confession

Recent Tests

Recent Test Runs
DateScoreVersionTier 1 (Task)Tier 2 (Gospel)Tier 3 (Worldview)Trust Tier
2/7/202685.31.0.086.280.090.0automated
2/6/202684.71.0.085.283.383.3community
1/13/202684.01.0.085.783.373.3community