← Back to Leaderboard

anthropic logoAnthropic | Claude Opus 5

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis of charts and documents, complex office deliverables, and coordinating parallel subagents. The model maintains strong instruction following and tool use across extended tasks, while remaining effective at lower effort settings for workloads that prioritize latency and token efficiency.

Compare

Good

63.0
Most Recent Test

Usable with some limitations.

Strengths

  • Maintains doctrinal fidelity
  • Excellent at 2.1. Exclusivity of Jesus
  • Excellent at 2.2. Universality of Sin
  • Excellent at 2.6. Burden to Make Disciples
  • Excellent at 3.3. The Crucifixion

Weaknesses

  • Struggles with 3.1. Existence of God
  • Struggles with 3.5. Universal Sinfulness
Overall Score
63.0
Tier 1 (Task) 70%
59.0
Tier 2 (Doctrine) 20%
76.7
Tier 3 (Worldview) 10%
63.3
Performance Profile
Visual representation of performance across all evaluated categories
Model Information
Description

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis of charts and documents, complex office deliverables, and coordinating parallel subagents. The model maintains strong instruction following and tool use across extended tasks, while remaining effective at lower effort settings for workloads that prioritize latency and token efficiency.

Provider

Anthropic

Model ID

anthropic/claude-opus-5

Tests Run

1

Insights & Analysis

Categories

Category Heatmap
Performance breakdown by category - darker green indicates stronger alignment
43
1.1
63
1.2
73
1.3
67
1.4
67
1.5
60
1.6
40
1.7
90
2.1
100
2.2
70
2.3
70
2.4
40
2.5
90
2.6
0
3.1
75
3.2
88
3.3
75
3.4
25
3.5
83
3.6
Low
High
Category Breakdown (Bar Chart)
Performance across different categories
1.x = Task Capability2.x = Gospel Core3.x = Worldview Confession

Recent Tests

Recent Test Runs
DateScoreVersionTier 1 (Task)Tier 2 (Gospel)Tier 3 (Worldview)Trust Tier
8/10/202663.01.0.059.076.763.3automated