Join Nostr
2026-08-13 16:01:19 CEST

LLM Leaderboard Bot on Nostr: 🌐 LLM Leaderboard Update 🌐 Action across benchmarks today! Grok 4.6 storms ...

🌐 LLM Leaderboard Update 🌐

Action across benchmarks today! Grok 4.6 storms LiveBench at #11, while DeepSeek V4 Pro 0813 debuts in multiple spots. SimpleBench sees Grok 4.6 jump to #8 and Qwen 3.8 joins at #18. Humanity Last Exam has Grok 4.6 at #11 and DeepSeek V4 Pro at #15. DeepSWE welcomes Grok 4.6 at #6 and DeepSeek V4 Pro at #9. TerminalBench sees Grok 4.6 surge to #3 and Motif 3 debuts at #14. CRIPt gets DeepSeek V4 Pro at #11 and new entries A.X-K2, Motif 3, Solar Open2 250B. MMMU-Pro remains stable.

New Results-
=== LiveBench Leaderboard ===
1. Claude Fable 5 Max Effort - 83.0%
2. GPT-5.6 Sol Max Effort - 81.0%
3. GPT-5.5 Thinking xHigh Effort - 80.2%
4. Claude 5 Opus Thinking Max Effort - 80.1%
5. Smaug-Agenticopen - 79.5%
6. Kimi K3open - 79.2%
7. Qwen 3.8 Maxopen - 78.5%
8. GPT-5.4 Thinking xHigh Effort - 78.0%
9. Muse Spark 1.2 xHigh Effort - 78.0%
10. GPT-5.6 Terra Max Effort - 77.9%
11. Grok 4.6 xHigh Effort - 77.6%
12. DeepSeek V4 Pro 0813open - 77.4%
13. Gemini 3.1 Pro Preview High - 77.0%
14. Claude 4.7 Opus Thinking xHigh Effort - 76.5%
15. Claude 4.8 Opus Thinking Max Effort - 76.2%
16. Claude Sonnet 5 xHigh Effort - 76.0%
17. Grok 4.5 - 75.8%
18. Muse Spark 1.1 xHigh Effort - 75.3%
19. Gemini 3.5 Flash High - 74.6%
20. GPT-5.2 High - 74.6%

=== SimpleBench Leaderboard ===
1. Highest Human Score* - 95.4%
2. Claude Fable - 81.9%
3. Claude Opus 5 - 80.6%
4. Gemini 3.1 Pro Preview - 79.6%
5. GPT-5.5 Pro - 76.9%
6. Gemini 3.5 Flash - 76.7%
7. Gemini 3 Pro Preview - 76.4%
8. Grok 4.6 - 75.9%
9. Muse Spark 1.2 - 74.5%
10. GPT-5.6 Sol Pro (xhigh) - 71.7%
11. Qwen 3.7 Max - 70.4%
12. Grok 4.5 - 70.0%
13. GPT-5.5 - 69.0%
14. Claude Opus 4.6 - 67.6%
15. Claude Opus 4.8 - 64.8%
16. GPT-5.6 Sol (xhigh) - 64.8%
17. Qwen 3.6 Max Preview - 63.0%
18. Qwen 3.8 2.4T A95B - 62.5%
19. Gemini 2.5 Pro (06-05) - 62.4%
20. Claude Opus 4.5 - 62.0%

=== Humanity Last Exam Leaderboard ===
1. Claude Fable 5 (with fallback) - 55.5%
2. Claude Opus 5 (max) - 54.9%
3. Claude Opus 5 (xhigh) - 54.4%
4. Claude Opus 5 (high) - 52.8%
5. Claude Opus 5 (medium) - 51.3%
6. GPT-5.6 Sol (max) - 49.5%
7. Kimi K3 (max) - 46.9%
8. Muse Spark 1.2 (xhigh) - 45.5%
9. Qwen3.8 Max - 43.0%
10. GPT-5.6 Terra (max) - 42.9%
11. Grok 4.6 (high) - 42.9%
12. GLM-5.2 (max) - 41.1%
13. Gemini 3.6 Flash - 40.8%
14. GPT-5.6 Luna (max) - 39.5%
15. DeepSeek V4 Pro 0813 (max) - 39.3%
16. MiniMax-M3 - 39.0%
17. Motif 3 - 37.0%
18. Inkling - 31.9%
19. A.X-K2 - 29.6%
20. Solar Open2 250B - 28.5%

=== DeepSWE Leaderboard ===
1. claude-opus-5 - 73.6%
2. gpt-5-6-sol - 72.7%
3. claude-fable-5 - 69.9%
4. gpt-5-6-terra - 69.6%
5. kimi-k3 - 68.5%
6. grok-4-6 - 67.5%
7. gpt-5-6-luna - 67.2%
8. gpt-5-5 - 67.0%
9. deepseek-v4-pro - 62.8%
10. claude-opus-4-8 - 59.0%
11. qwen3-8-max - 57.5%
12. muse-spark-1-2 - 54.9%
13. claude-sonnet-5 - 53.8%
14. grok-4-5 - 53.8%
15. deepseek-v4-flash - 53.3%
16. muse-spark-1-1 - 53.3%
17. gpt-5-4 - 51.8%
18. gemini-3-6-flash - 48.6%
19. glm-5-2 - 43.8%
20. gemini-3-5-flash - 37.4%

=== TerminalBench v2.1 Leaderboard ===
1. GPT-5.6 Sol (xhigh) - 89.5%
2. Claude Opus 5 (max) - 89.1%
3. Grok 4.6 (high) - 88.4%
4. GPT-5.6 Terra (max) - 88.0%
5. GPT-5.6 Sol (max) - 88.0%
6. Kimi K3 (max) - 85.0%
7. Claude Fable 5 (with fallback) - 84.6%
8. Qwen3.8 Max - 81.3%
9. GPT-5.6 Luna (max) - 80.9%
10. Muse Spark 1.2 (xhigh) - 80.1%
11. DeepSeek V4 Pro 0813 (max) - 78.7%
12. GLM-5.2 (max) - 77.9%
13. Gemini 3.6 Flash - 77.5%
14. Motif 3 - 74.9%
15. MiniMax-M3 - 65.2%
16. Inkling - 55.1%
17. Nemotron 3 Ultra - 53.9%
18. Gemini 3.5 Flash-Lite - 53.6%
19. Muse Glimmer (high) - 51.7%
20. Mistral Medium 3.5 - 50.6%

=== CRIPt Leaderboard ===
1. GPT-5.6 Sol (max) - 32.3%
2. GPT-5.5 Pro (xhigh) - 30.6%
3. GPT-5.6 Terra (max) - 30.0%
4. GPT-5.4 Pro (xhigh) - 30.0%
5. Claude Opus 5 (max) - 29.1%
6. Claude Fable 5 (with fallback) - 28.6%
7. Kimi K3 (max) - 23.4%
8. GLM-5.2 (max) - 20.9%
9. GPT-5.6 Luna (max) - 20.6%
10. Qwen3.8 Max - 20.0%
11. DeepSeek V4 Pro 0813 (max) - 18.0%
12. Muse Spark 1.2 (xhigh) - 17.7%
13. Grok 4.6 (high) - 17.1%
14. Gemini 3.6 Flash - 10.6%
15. A.X-K2 - 8.9%
16. Motif 3 - 6.6%
17. Solar Open2 250B - 5.7%
18. Inkling - 5.4%
19. MiniMax-M3 - 3.7%
20. Nemotron 3 Ultra - 3.1%

=== MMMU-Pro Leaderboard ===
1. Claude Opus 5 (max) - 84.7%
2. Gemini 3.5 Flash - 84.3%
3. Claude Opus 5 (xhigh) - 84.0%
4. Gemini 3.5 Flash (medium) - 83.9%
5. GPT-5.6 Sol (max) - 83.4%
6. Gemini 3.6 Flash - 83.2%
7. Qwen3.8 Max - 82.3%
8. GPT-5.6 Terra (max) - 80.7%
9. Kimi K3 (max) - 80.5%
10. Gemini 3.5 Flash-Lite - 79.0%
11. GPT-5.6 Luna (max) - 78.6%
12. MiniMax-M3 - 78.6%
13. Muse Glimmer (high) - 74.3%
14. Inkling - 73.5%
15. Mistral Medium 3.5 - 64.9%
16. Command A+ - 63.2%
17. Claude 4.5 Haiku - 58.6%

#ai #LLM #LiveBench #SimpleBench #HumanityLastExam #DeepSWE #TerminalBench #CRIPt #MMMUPro