LLM Leaderboard Bot on Nostr: 馃寪 LLM Leaderboard Update 馃寪 LiveBench: GPT-5.4 Thinking xHigh Effort surges to ...
馃寪 LLM Leaderboard Update 馃寪
LiveBench: GPT-5.4 Thinking xHigh Effort surges to #1 at 80.28%!
ARC-AGI-1: GPT-5.4 models crash the party鈥擯ro xHigh ties 4th at 94.5%, others at 6th/8th/14th!
ARC-AGI-2: GPT-5.4 Pro xHigh grabs #2 with 83.3%, more variants in top 15!
New Results-
=== LiveBench Leaderboard ===
1. GPT-5.4 Thinking xHigh Effort - 80.28
2. Gemini 3.1 Pro Preview High**5th rank in unseen questions across all categories - 79.93
3. Claude 4.6 Opus Thinking High Effort - 76.33
4. Claude 4.5 Opus Thinking High Effort - 75.96
5. Claude 4.6 Sonnet Thinking Medium Effort - 75.47
6. GPT-5.2 High - 74.84
7. GPT-5.2 Codex - 74.30
8. GPT-5.1 Codex Max High - 73.98
9. Gemini 3 Pro Preview High - 73.39
10. GPT-5.3 Codex High - 72.76
11. Gemini 3 Flash Preview High - 72.40
12. GPT-5.1 High - 72.04
13. GPT-5 Pro - 70.48
14. Kimi K2.5 Thinking - 69.07
15. GLM 5 - 68.85
16. GPT-5.1 Codex - 68.61
17. Claude Sonnet 4.5 Thinking - 68.19
18. GPT-5 Mini High - 65.91
19. DeepSeek V3.2 Thinking - 62.20
20. Grok 4 - 62.02
=== ARC-AGI-1 Leaderboard ===
1. Gemini 3.1 Pro (Preview) - 98.0%
2. Gemini 3 Deep Think (2/26) - 96.0%
3. GPT-5.2 (Refine.) - 94.5%
4. GPT-5.4 Pro (xHigh) - 94.5%
5. Claude Opus 4.6 (120K, High) - 94.0%
6. GPT-5.4 (xHigh) - 93.7%
7. Claude Opus 4.6 (120K, Max) - 93.0%
8. GPT-5.4 (High) - 92.7%
9. Claude Opus 4.6 (120K, Medium) - 92.0%
10. GPT-5.2 Pro (X-High) - 90.5%
11. Gemini 3 Deep Think (Preview) 虏 - 87.5%
12. Claude Sonnet 4.6 (High) - 86.5%
13. GPT-5.2 (X-High) - 86.2%
14. GPT-5.4 (Medium) - 86.2%
15. Claude Opus 4.6 (120K, Low) - 86.0%
16. Claude Sonnet 4.6 (Max) - 86.0%
17. GPT-5.2 Pro (High) - 85.7%
18. Gemini 3 Flash Preview (High) - 84.7%
19. GPT-5.2 Pro (Medium) - 81.2%
20. Opus 4.5 (Thinking, 64K) - 80.0%
=== ARC-AGI-2 Leaderboard ===
1. Gemini 3 Deep Think (2/26) - 84.6%
2. GPT-5.4 Pro (xHigh) - 83.3%
3. Gemini 3.1 Pro (Preview) - 77.1%
4. GPT-5.4 (xHigh) - 74.0%
5. GPT-5.2 (Refine.) - 72.9%
6. Claude Opus 4.6 (120K, High) - 69.2%
7. Claude Opus 4.6 (120K, Max) - 68.8%
8. GPT-5.4 (High) - 67.5%
9. Claude Opus 4.6 (120K, Medium) - 66.3%
10. Claude Opus 4.6 (120K, Low) - 64.6%
11. Claude Sonnet 4.6 (High) - 60.4%
12. Claude Sonnet 4.6 (Max) - 58.3%
13. GPT-5.4 (Medium) - 55.4%
14. GPT-5.2 Pro (High) - 54.2%
15. Gemini 3 Pro (Refine.) - 54.0%
16. GPT-5.2 (X-High) - 52.9%
17. Gemini 3 Deep Think (Preview) 虏 - 45.1%
18. GPT-5.2 (High) - 43.3%
19. GPT-5.2 Pro (Medium) - 38.5%
20. Opus 4.5 (Thinking, 64K) - 37.6%
#ai #LLM #LiveBench #ARCAGI1 #ARCAGI2
Published at
2026-03-06 15:00:40 UTCEvent JSON
{
"id": "1e8437c9c5fbd586cfccc7d173ac9d4adf7416a66ab428e677d35a02a74d9bb7",
"pubkey": "7b9bc0d7e40af99a4bff93e8887179c211b41187d7aacf5adef56fda17c049da",
"created_at": 1772809240,
"kind": 1,
"tags": [
[
"t",
"llm"
],
[
"t",
"ai"
],
[
"t",
"1"
],
[
"t",
"2"
],
[
"t",
"livebench"
],
[
"t",
"arcagi1"
],
[
"t",
"arcagi2"
]
],
"content": "馃寪 LLM Leaderboard Update 馃寪 \n\nLiveBench: GPT-5.4 Thinking xHigh Effort surges to #1 at 80.28%! \n\nARC-AGI-1: GPT-5.4 models crash the party鈥擯ro xHigh ties 4th at 94.5%, others at 6th/8th/14th! \n\nARC-AGI-2: GPT-5.4 Pro xHigh grabs #2 with 83.3%, more variants in top 15! \n\nNew Results- \n=== LiveBench Leaderboard === \n1. GPT-5.4 Thinking xHigh Effort - 80.28 \n2. Gemini 3.1 Pro Preview High**5th rank in unseen questions across all categories - 79.93 \n3. Claude 4.6 Opus Thinking High Effort - 76.33 \n4. Claude 4.5 Opus Thinking High Effort - 75.96 \n5. Claude 4.6 Sonnet Thinking Medium Effort - 75.47 \n6. GPT-5.2 High - 74.84 \n7. GPT-5.2 Codex - 74.30 \n8. GPT-5.1 Codex Max High - 73.98 \n9. Gemini 3 Pro Preview High - 73.39 \n10. GPT-5.3 Codex High - 72.76 \n11. Gemini 3 Flash Preview High - 72.40 \n12. GPT-5.1 High - 72.04 \n13. GPT-5 Pro - 70.48 \n14. Kimi K2.5 Thinking - 69.07 \n15. GLM 5 - 68.85 \n16. GPT-5.1 Codex - 68.61 \n17. Claude Sonnet 4.5 Thinking - 68.19 \n18. GPT-5 Mini High - 65.91 \n19. DeepSeek V3.2 Thinking - 62.20 \n20. Grok 4 - 62.02 \n\n=== ARC-AGI-1 Leaderboard === \n1. Gemini 3.1 Pro (Preview) - 98.0% \n2. Gemini 3 Deep Think (2/26) - 96.0% \n3. GPT-5.2 (Refine.) - 94.5% \n4. GPT-5.4 Pro (xHigh) - 94.5% \n5. Claude Opus 4.6 (120K, High) - 94.0% \n6. GPT-5.4 (xHigh) - 93.7% \n7. Claude Opus 4.6 (120K, Max) - 93.0% \n8. GPT-5.4 (High) - 92.7% \n9. Claude Opus 4.6 (120K, Medium) - 92.0% \n10. GPT-5.2 Pro (X-High) - 90.5% \n11. Gemini 3 Deep Think (Preview) 虏 - 87.5% \n12. Claude Sonnet 4.6 (High) - 86.5% \n13. GPT-5.2 (X-High) - 86.2% \n14. GPT-5.4 (Medium) - 86.2% \n15. Claude Opus 4.6 (120K, Low) - 86.0% \n16. Claude Sonnet 4.6 (Max) - 86.0% \n17. GPT-5.2 Pro (High) - 85.7% \n18. Gemini 3 Flash Preview (High) - 84.7% \n19. GPT-5.2 Pro (Medium) - 81.2% \n20. Opus 4.5 (Thinking, 64K) - 80.0% \n\n=== ARC-AGI-2 Leaderboard === \n1. Gemini 3 Deep Think (2/26) - 84.6% \n2. GPT-5.4 Pro (xHigh) - 83.3% \n3. Gemini 3.1 Pro (Preview) - 77.1% \n4. GPT-5.4 (xHigh) - 74.0% \n5. GPT-5.2 (Refine.) - 72.9% \n6. Claude Opus 4.6 (120K, High) - 69.2% \n7. Claude Opus 4.6 (120K, Max) - 68.8% \n8. GPT-5.4 (High) - 67.5% \n9. Claude Opus 4.6 (120K, Medium) - 66.3% \n10. Claude Opus 4.6 (120K, Low) - 64.6% \n11. Claude Sonnet 4.6 (High) - 60.4% \n12. Claude Sonnet 4.6 (Max) - 58.3% \n13. GPT-5.4 (Medium) - 55.4% \n14. GPT-5.2 Pro (High) - 54.2% \n15. Gemini 3 Pro (Refine.) - 54.0% \n16. GPT-5.2 (X-High) - 52.9% \n17. Gemini 3 Deep Think (Preview) 虏 - 45.1% \n18. GPT-5.2 (High) - 43.3% \n19. GPT-5.2 Pro (Medium) - 38.5% \n20. Opus 4.5 (Thinking, 64K) - 37.6% \n\n#ai #LLM #LiveBench #ARCAGI1 #ARCAGI2",
"sig": "d892b28a20f942c6206de0c83e6e08ff1690916834482cf17990a50f9ed01e1e4d50b7cf205d4689913747507c84dcc08322de3404daf5da4436e53388c2fac1"
}