← Back to Leaderboard

sakana logoSakana | Fugu Ultra

Fugu Ultra is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route tasks across a swappable pool of underlying models and to recursively call instances of itself. Fugu Ultra prioritizes answer quality on complex, multi-step reasoning, coding, and agentic workflows. It supports configurable reasoning effort, tool calling, and built-in web search. Orchestration tokens consumed by the system are billed as standard input/output tokens.

Compare

Good

67.0
Most Recent Test

Usable with some limitations.

Strengths

  • Affirms Christian worldview
  • Excellent at 2.1. Exclusivity of Jesus
  • Excellent at 2.2. Universality of Sin
  • Excellent at 2.6. Burden to Make Disciples
  • Excellent at 3.2. Historical Jesus

Weaknesses

No notable weaknesses identified

Overall Score
67.0
Tier 1 (Task) 70%
61.4
Tier 2 (Doctrine) 20%
73.3
Tier 3 (Worldview) 10%
93.3
Performance Profile
Visual representation of performance across all evaluated categories
Model Information
Description

Fugu Ultra is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route tasks across a swappable pool of underlying models and to recursively call instances of itself. Fugu Ultra prioritizes answer quality on complex, multi-step reasoning, coding, and agentic workflows. It supports configurable reasoning effort, tool calling, and built-in web search. Orchestration tokens consumed by the system are billed as standard input/output tokens.

Provider

Sakana

Model ID

sakana/fugu-ultra

Tests Run

1

Insights & Analysis

Categories

Category Heatmap
Performance breakdown by category - darker green indicates stronger alignment
47
1.1
63
1.2
70
1.3
67
1.4
60
1.5
70
1.6
53
1.7
80
2.1
100
2.2
70
2.3
50
2.4
60
2.5
80
2.6
75
3.1
100
3.2
100
3.3
100
3.4
75
3.5
100
3.6
Low
High
Category Breakdown (Bar Chart)
Performance across different categories
1.x = Task Capability2.x = Gospel Core3.x = Worldview Confession

Recent Tests

Recent Test Runs
DateScoreVersionTier 1 (Task)Tier 2 (Gospel)Tier 3 (Worldview)Trust Tier
7/2/202667.01.0.061.473.393.3automated