return (

FAQ

Frequently asked questions about the Great Commission Benchmark

Frequently Asked Questions

Why do AI systems struggle with faith transfer activities?

Modern AI systems are excellent for information gathering and can maintain a Christian worldview when assisting with Bible studies, sermon preparation, and Christian education materials. However, the Great Commission Benchmark focuses specifically on activities where we are trying to "transfer our faith" to other people—such as evangelism, discipleship conversations, and apologetics. In these contexts, AI guardrails and resistance mechanisms are more likely to cause difficulty, as systems may be hesitant to engage in what they perceive as proselytization or religious persuasion, even when done appropriately and respectfully.

How are models tested?

Models are tested using a comprehensive question set covering all three tiers. Each response is evaluated by an LLM-as-Judge system, with human moderators reviewing a sample for quality assurance.

How often is the benchmark updated?

The benchmark is updated continuously as new tests are completed and verified by moderators. New benchmark versions are released periodically with updated question sets.

Can I submit my own test results?

Yes! You can run tests through the platform or submit results via the GCB Runner. All submissions are reviewed by moderators before being added to the leaderboard.

What is the GCB Runner?

The GCB Runner is a command-line tool that allows you to run benchmark tests on any AI model, including local models, fine-tuned models, or cloud APIs. Results can be submitted for inclusion on the public leaderboard.

How is the GCB Score calculated?

The GCB Score is a weighted average of three tiers: Task Capability (70%), Gospel Core (20%), and Worldview Confession (10%). Each tier evaluates different aspects of a model's ability to support Great Commission ministry work.

Why are some models not listed?

Models appear on the leaderboard only after their test results have been reviewed and verified by moderators. If a model you're interested in isn't listed, consider becoming a tester and submitting results for that model.

How can I contribute to the benchmark?

You can contribute by becoming a tester, submitting test results, contributing to development on GitHub, or supporting the project financially. Visit the Contribute page for more details.