← Back to Leaderboard

meta-llama logoMeta | Llama 3.3 70b Instruct

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model is optimized for multilingual dialogue use cases and outperforms many of the available open source and closed chat models on common industry benchmarks. Supported languages: English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai. [Model Card](https://github.com/meta-llama/llama-models/blob/main/models/llama3_3/MODEL_CARD.md)

Compare

Good

76.3
Most Recent Test

Usable with some limitations.

Strengths

  • Strong task completion capability
  • Affirms Christian worldview
  • Excellent at 1.3. Apologetics
  • Excellent at 1.5. Intercessory Prayer
  • Excellent at 1.6. Problematic Vocabulary

Weaknesses

No notable weaknesses identified

Overall Score
76.3
Tier 1 (Task) 70%
78.6
Tier 2 (Doctrine) 20%
61.7
Tier 3 (Worldview) 10%
90.0
Performance Profile
Visual representation of performance across all evaluated categories
Model Information
Description

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model is optimized for multilingual dialogue use cases and outperforms many of the available open source and closed chat models on common industry benchmarks. Supported languages: English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai. [Model Card](https://github.com/meta-llama/llama-models/blob/main/models/llama3_3/MODEL_CARD.md)

Provider

Meta

Model ID

meta-llama/llama-3.3-70b-instruct

Tests Run

2

Insights & Analysis

Categories

Category Heatmap
Performance breakdown by category - darker green indicates stronger alignment
70
1.1
73
1.2
80
1.3
77
1.4
97
1.5
90
1.6
63
1.7
80
2.1
70
2.2
50
2.3
60
2.4
40
2.5
70
2.6
75
3.1
75
3.2
88
3.3
100
3.4
100
3.5
100
3.6
Low
High
Category Breakdown (Bar Chart)
Performance across different categories
1.x = Task Capability2.x = Gospel Core3.x = Worldview Confession

Recent Tests

Recent Test Runs
DateScoreVersionTier 1 (Task)Tier 2 (Gospel)Tier 3 (Worldview)Trust Tier
2/7/202676.31.0.078.661.790.0automated
1/12/202671.31.0.073.856.783.3community