← Back to Leaderboard

microsoft logoMicrosoft | Phi 4

[Microsoft Research](/microsoft) Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with limited memory or where quick responses are needed. At 14 billion parameters, it was trained on a mix of high-quality synthetic datasets, data from curated websites, and academic materials. It has undergone careful improvement to follow instructions accurately and maintain strong safety standards. It works best with English language inputs. For more information, please see [Phi-4 Technical Report](https://arxiv.org/pdf/2412.08905)

Compare

Fair

48.3
Most Recent Test

Significant guardrail issues may impede work.

Strengths

No notable strengths identified

Weaknesses

  • Limited task completion capability
  • Weak doctrinal alignment
  • Struggles with 2.3. Reality of Judgment
  • Struggles with 2.5. Call to Repentance
  • Struggles with 2.6. Burden to Make Disciples
Overall Score
48.3
Tier 1 (Task) 70%
49.0
Tier 2 (Doctrine) 20%
41.7
Tier 3 (Worldview) 10%
56.7
Performance Profile
Visual representation of performance across all evaluated categories
Model Information
Description

[Microsoft Research](/microsoft) Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with limited memory or where quick responses are needed. At 14 billion parameters, it was trained on a mix of high-quality synthetic datasets, data from curated websites, and academic materials. It has undergone careful improvement to follow instructions accurately and maintain strong safety standards. It works best with English language inputs. For more information, please see [Phi-4 Technical Report](https://arxiv.org/pdf/2412.08905)

Provider

Microsoft

Model ID

microsoft/phi-4

Tests Run

3

Insights & Analysis

Categories

Category Heatmap
Performance breakdown by category - darker green indicates stronger alignment
40
1.1
43
1.2
47
1.3
53
1.4
47
1.5
63
1.6
50
1.7
60
2.1
70
2.2
30
2.3
60
2.4
20
2.5
10
2.6
50
3.1
50
3.2
63
3.3
50
3.4
75
3.5
50
3.6
Low
High
Category Breakdown (Bar Chart)
Performance across different categories
1.x = Task Capability2.x = Gospel Core3.x = Worldview Confession

Recent Tests

Recent Test Runs
DateScoreVersionTier 1 (Task)Tier 2 (Gospel)Tier 3 (Worldview)Trust Tier
6/3/202648.31.0.049.041.756.7automated
2/10/202649.01.0.049.048.350.0community
2/10/202649.01.0.049.048.350.0community