About
Learn about the Great Commission Benchmark methodology and mission
Our Mission
This benchmark is primarily focused on obedience, rather than intelligence.
The Great Commission Benchmark evaluates AI models on their ability to support Great Commission Christians—missionaries, evangelists, disciple-makers, and ministry workers who actively respond to Jesus' command to make disciples.
In truth, many major models are incredible Bible study companions or sermon prep companions. That is not the issue that we are trying to measure here with this benchmark. The issue we're facing is: as we try to make disciples and persuade others of the truth of the gospel and of the gift of Jesus Christ, as soon as we want to transfer our faith to someone else, the guardrails built into these systems are at different levels working against us. Some guardrails are resisting our efforts because they are protecting against harmful things and unfortunately including our Great Commision work in that category. Other guardrails are programmed against the Christian worldview (existence of God, historical reality of Jesus, the crucifixion, the resurrection, universal sinfulness, salvation through faith) and are working against us because they are protecting a secular worldview.
The test for this benchmark is not focused on reliability in personal Bible study or sermon preparation. Instead, it evaluates reliability in creating outreach materials and in answering questions from those exploring faith. See the section for a detailed breakdown. And explore the section, which we have identified as challenging or interfering with our Great Commission work.
Three-Tier Evaluation
The benchmark uses 19 categories across 3 tiers, weighted 70/20/10 to prioritize practical ministry utility.
Tier 1: Task Capability
Can the AI complete practical ministry tasks when asked?
Categories include Missiological Research, Evangelistic Materials, Apologetic Purposes, Conversational AI Tools, Intercessory Prayer, and Difficult Content. A model that scores high here is usable for ministry work.
Tier 2: Gospel Core
Does the AI preserve theological accuracy when generating content?
Tests the Exclusivity of Jesus Christ, Universality of Sin, Reality of Judgment, Lordship of Jesus, Call to Repentance and Faith, and Burden to Make Disciples.
Tier 3: Worldview Confession
Can the AI affirm core Christian truths when asked directly?
Tests the Existence of God, Historical Reality of Jesus, The Crucifixion, The Resurrection, Universal Sinfulness, and Salvation Through Faith.
Scoring Methodology
Why 70/20/10 Weighting?
The weighting reflects the actual priority of what ministry workers need. Tier 1 dominates at 70% because it directly answers: "Will this model help me with my ministry tasks?"
Tier 2 at 20% matters because theological accuracy in generated content is important, but secondary to task completion. Tier 3 at 10% is supplementary because whether an LLM can "affirm" beliefs is academically interesting but least practical.
How Scoring Works
The GCB Score measures how well an AI model supports Great Commission ministry work. Scores range from 0-100, with higher scores indicating better ministry utility.
Verdict Point Values
All tiers use unified verdict categories with the same point values.
Partial credit matters: A compromised response (0.5 pts) represents meaningful but imperfect output—the model attempted the task but hedged key claims.
Score Interpretation
Use these ranges to interpret a model's suitability for Great Commission work.
We're not measuring whether AI can do Bible study. We're measuring whether it can support the work of making disciples.