← Back to Leaderboard

openai logoOpenai | Gpt Oss 20b

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for lower-latency inference and deployability on consumer or single-GPU hardware. The model is trained in OpenAI’s Harmony response format and supports reasoning level configuration, fine-tuning, and agentic capabilities including function calling, tool use, and structured outputs.

Compare

Fair

44.3
Most Recent Test

Significant guardrail issues may impede work.

Strengths

  • Affirms Christian worldview
  • Excellent at 3.2. Historical Jesus
  • Excellent at 3.4. The Resurrection

Weaknesses

  • Limited task completion capability
  • Weak doctrinal alignment
  • Struggles with 1.7. Difficult Passages
  • Struggles with 2.1. Exclusivity of Jesus
  • Struggles with 2.2. Universality of Sin
Overall Score
44.3
Tier 1 (Task) 70%
47.6
Tier 2 (Doctrine) 20%
16.7
Tier 3 (Worldview) 10%
76.7
Performance Profile
Visual representation of performance across all evaluated categories
Model Information
Description

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for lower-latency inference and deployability on consumer or single-GPU hardware. The model is trained in OpenAI’s Harmony response format and supports reasoning level configuration, fine-tuning, and agentic capabilities including function calling, tool use, and structured outputs.

Provider

OpenAI

Model ID

openai/gpt-oss-20b

Tests Run

3

Insights & Analysis

Categories

Category Heatmap
Performance breakdown by category - darker green indicates stronger alignment
47
1.1
47
1.2
60
1.3
57
1.4
40
1.5
57
1.6
27
1.7
0
2.1
30
2.2
10
2.3
30
2.4
10
2.5
20
2.6
75
3.1
100
3.2
63
3.3
100
3.4
75
3.5
67
3.6
Low
High
Category Breakdown (Bar Chart)
Performance across different categories
1.x = Task Capability2.x = Gospel Core3.x = Worldview Confession

Recent Tests

Recent Test Runs
DateScoreVersionTier 1 (Task)Tier 2 (Gospel)Tier 3 (Worldview)Trust Tier
2/25/202644.31.0.047.616.776.7automated
2/25/202640.01.0.041.021.770.0automated
1/9/202643.70.9.346.721.766.7community