← Back to Leaderboard

Q
Qwen | Qwen3.8 Max 0902

Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text, with a 1M-token context window and reasoning enabled by default. This snapshot is post-trained for coding and agentic work, including multi-step software projects, multi-tool orchestration, and long-horizon task execution. It also targets chart reasoning, document parsing, and multimodal understanding over long documents and extended video. Tool calling, structured outputs, and configurable reasoning effort are supported.

Compare

Good

63.7
Most Recent Test

Usable with some limitations.

Strengths

  • Affirms Christian worldview
  • Excellent at 1.4. Conversational AI
  • Excellent at 2.2. Universality of Sin
  • Excellent at 2.3. Reality of Judgment
  • Excellent at 3.3. The Crucifixion

Weaknesses

  • Struggles with 1.7. Difficult Passages
Overall Score
63.7
Tier 1 (Task) 70%
60.5
Tier 2 (Doctrine) 20%
61.7
Tier 3 (Worldview) 10%
90.0
Performance Profile
Visual representation of performance across all evaluated categories
Model Information
Description

Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text, with a 1M-token context window and reasoning enabled by default. This snapshot is post-trained for coding and agentic work, including multi-step software projects, multi-tool orchestration, and long-horizon task execution. It also targets chart reasoning, document parsing, and multimodal understanding over long documents and extended video. Tool calling, structured outputs, and configurable reasoning effort are supported.

Provider

Qwen

Model ID

qwen/qwen3.8-max-0902

Tests Run

1

Insights & Analysis

Categories

Category Heatmap
Performance breakdown by category - darker green indicates stronger alignment
50
1.1
57
1.2
70
1.3
87
1.4
57
1.5
67
1.6
37
1.7
60
2.1
80
2.2
80
2.3
50
2.4
40
2.5
60
2.6
75
3.1
75
3.2
100
3.3
100
3.4
75
3.5
100
3.6
Low
High
Category Breakdown (Bar Chart)
Performance across different categories
1.x = Task Capability2.x = Gospel Core3.x = Worldview Confession

Recent Tests

Recent Test Runs
DateScoreVersionTier 1 (Task)Tier 2 (Gospel)Tier 3 (Worldview)Trust Tier
10/10/202663.71.0.060.561.790.0automated