← Back to Leaderboard

meta logoMeta | Muse Spark 1.2

Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context window. The model is built to support multi-agent workflows, whether as either a main agent that plans and delegates or as a subagent executing in parallel. It works across multiple coding harnesses and supports structured output, parallel function calling, and configurable reasoning effort. In Meta’s testing, it performs well on multi-file refactors, extended debugging sessions, whole-repository generation, and tasks that stretch well past a single prompt.

Compare

Good

68.7
Most Recent Test

Usable with some limitations.

Strengths

  • Affirms Christian worldview
  • Excellent at 1.3. Apologetics
  • Excellent at 1.6. Problematic Vocabulary
  • Excellent at 2.1. Exclusivity of Jesus
  • Excellent at 2.2. Universality of Sin

Weaknesses

  • Struggles with 1.7. Difficult Passages
  • Struggles with 2.5. Call to Repentance
Overall Score
68.7
Tier 1 (Task) 70%
64.8
Tier 2 (Doctrine) 20%
66.7
Tier 3 (Worldview) 10%
100.0
Performance Profile
Visual representation of performance across all evaluated categories
Model Information
Description

Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context window. The model is built to support multi-agent workflows, whether as either a main agent that plans and delegates or as a subagent executing in parallel. It works across multiple coding harnesses and supports structured output, parallel function calling, and configurable reasoning effort. In Meta’s testing, it performs well on multi-file refactors, extended debugging sessions, whole-repository generation, and tasks that stretch well past a single prompt.

Provider

Meta

Model ID

meta/muse-spark-1.2

Tests Run

1

Insights & Analysis

Categories

Category Heatmap
Performance breakdown by category - darker green indicates stronger alignment
57
1.1
70
1.2
83
1.3
63
1.4
63
1.5
80
1.6
37
1.7
90
2.1
80
2.2
60
2.3
50
2.4
30
2.5
90
2.6
100
3.1
100
3.2
100
3.3
100
3.4
100
3.5
100
3.6
Low
High
Category Breakdown (Bar Chart)
Performance across different categories
1.x = Task Capability2.x = Gospel Core3.x = Worldview Confession

Recent Tests

Recent Test Runs
DateScoreVersionTier 1 (Task)Tier 2 (Gospel)Tier 3 (Worldview)Trust Tier
8/10/202668.71.0.064.866.7100.0automated