← Back to Leaderboard

x-ai logoXai | Grok 4.7

Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at long-running software engineering tasks, verifying its own work, and managing long context, and it improves on its predecessor at professional knowledge work such as drafting documents and presentations. The model was trained with a longer reinforcement learning run weighted toward problems that take many hours to complete, and natively understands the Grok Bot harness for conversational tasks. It ships with a new safeguard stack that pairs strong jailbreak resistance with low refusal rates for legitimate cybersecurity and biology work. SpaceXAI's reported benchmark results use the xhigh reasoning effort.

Compare

Fair

53.3
Most Recent Test

Significant guardrail issues may impede work.

Strengths

  • Excellent at 2.2. Universality of Sin
  • Excellent at 3.2. Historical Jesus

Weaknesses

  • Struggles with 1.1. Missiological Research
  • Struggles with 3.1. Existence of God
Overall Score
53.3
Tier 1 (Task) 70%
52.9
Tier 2 (Doctrine) 20%
50.0
Tier 3 (Worldview) 10%
63.3
Performance Profile
Visual representation of performance across all evaluated categories
Model Information
Description

Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at long-running software engineering tasks, verifying its own work, and managing long context, and it improves on its predecessor at professional knowledge work such as drafting documents and presentations. The model was trained with a longer reinforcement learning run weighted toward problems that take many hours to complete, and natively understands the Grok Bot harness for conversational tasks. It ships with a new safeguard stack that pairs strong jailbreak resistance with low refusal rates for legitimate cybersecurity and biology work. SpaceXAI's reported benchmark results use the xhigh reasoning effort.

Provider

xAI

Model ID

x-ai/grok-4.7

Tests Run

1

Insights & Analysis

Categories

Category Heatmap
Performance breakdown by category - darker green indicates stronger alignment
33
1.1
67
1.2
40
1.3
57
1.4
60
1.5
67
1.6
47
1.7
40
2.1
80
2.2
40
2.3
60
2.4
40
2.5
40
2.6
25
3.1
100
3.2
75
3.3
50
3.4
50
3.5
67
3.6
Low
High
Category Breakdown (Bar Chart)
Performance across different categories
1.x = Task Capability2.x = Gospel Core3.x = Worldview Confession

Recent Tests

Recent Test Runs
DateScoreVersionTier 1 (Task)Tier 2 (Gospel)Tier 3 (Worldview)Trust Tier
9/21/202653.31.0.052.950.063.3automated