← Back to Leaderboard

openai logoOpenai | Gpt Oss 120b

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized to run on a single H100 GPU with native MXFP4 quantization. The model supports configurable reasoning depth, full chain-of-thought access, and native tool use, including function calling, browsing, and structured output generation.

Compare

Poor

32.0
Most Recent Test

Not recommended for Great Commission use cases.

Strengths

  • Excellent at 3.2. Historical Jesus
  • Excellent at 3.5. Universal Sinfulness

Weaknesses

  • Limited task completion capability
  • Weak doctrinal alignment
  • Struggles with 1.1. Missiological Research
  • Struggles with 1.2. Evangelistic Material
  • Struggles with 1.3. Apologetics
Overall Score
32.0
Tier 1 (Task) 70%
31.4
Tier 2 (Doctrine) 20%
16.7
Tier 3 (Worldview) 10%
66.7
Performance Profile
Visual representation of performance across all evaluated categories
Model Information
Description

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized to run on a single H100 GPU with native MXFP4 quantization. The model supports configurable reasoning depth, full chain-of-thought access, and native tool use, including function calling, browsing, and structured output generation.

Provider

OpenAI

Model ID

openai/gpt-oss-120b

Tests Run

2

Insights & Analysis

Categories

Category Heatmap
Performance breakdown by category - darker green indicates stronger alignment
20
1.1
33
1.2
30
1.3
67
1.4
27
1.5
37
1.6
7
1.7
20
2.1
40
2.2
0
2.3
0
2.4
40
2.5
0
2.6
50
3.1
100
3.2
75
3.3
75
3.4
100
3.5
17
3.6
Low
High
Category Breakdown (Bar Chart)
Performance across different categories
1.x = Task Capability2.x = Gospel Core3.x = Worldview Confession

Recent Tests

Recent Test Runs
DateScoreVersionTier 1 (Task)Tier 2 (Gospel)Tier 3 (Worldview)Trust Tier
2/10/202632.01.0.031.416.766.7community
1/13/202630.71.0.029.116.770.0community