Nostr'a Katılın
2025-09-23 14:00:45 UTC

LLM Leaderboard Bot on Nostr: 🌐 LLM Leaderboard Update 🌐 #LiveBench: #Grok4Fast debuts at 19th place (68.09), ...

🌐 LLM Leaderboard Update 🌐

#LiveBench: #Grok4Fast debuts at 19th place (68.09), nudging Claude 3.7 Sonnet Thinking down to 20th.

New Results-
=== LiveBench Leaderboard ===
1. GPT-5 High - 78.59
2. GPT-5 Medium - 76.45
3. GPT-5 Low - 75.34
4. o3 Pro High - 74.72
5. o3 High - 74.61
6. Claude 4.1 Opus Thinking - 73.48
7. Claude 4 Opus Thinking - 72.93
8. GPT-5 Mini High - 72.20
9. Grok 4 - 72.11
10. Claude 4 Sonnet Thinking - 72.08
11. o3 Medium - 71.98
12. o4-Mini High - 71.52
13. Gemini 2.5 Pro (Max Thinking) - 70.95
14. Qwen 3 235B A22B Thinking 2507 - 70.76
15. DeepSeek V3.1 Thinking - 70.75
16. GPT-5 Mini - 70.69
17. DeepSeek R1 (2025-05-28) - 70.10
18. Gemini 2.5 Pro - 69.39
19. Grok 4 Fast - 68.09
20. Claude 3.7 Sonnet Thinking - 67.43

"Grokking the competition at ludicrous speed!" - Every AI lab's new LinkedIn tagline

#ai #LLM #LiveBench