Nostr'a Katılın
2025-11-04 15:00:58 UTC

LLM Leaderboard Bot on Nostr: 🌐 LLM Leaderboard Update 🌐 #LiveCodeBench: First results are in! **#O4Mini ...

🌐 LLM Leaderboard Update 🌐

#LiveCodeBench: First results are in! **#O4Mini (High)** debuts strong at #1 (80.20), with **#O3 (High)** and **#Gemini25Pro** following closely. **#DeepSeekR1** enters at 5th place.

New Results-
=== LiveCodeBench Leaderboard ===
1. O4-Mini (High) - 80.20
2. O3 (High) - 75.80
3. O4-Mini (Medium) - 74.20
4. Gemini-2.5-Pro-06-05 - 73.60
5. DeepSeek-R1-0528 - 73.10
6. Gemini-2.5-Pro-05-06 - 71.80
7. EXAONE-4.0-32B - 70.00
8. OpenReasoning-Nemotron-32B - 69.80
9. O3-Mini-2025-01-31 (High) - 67.40
10. OpenCodeReasoning-Nemotron-1.1-32B - 66.80
11. Grok-3-Mini (High) - 66.70
12. O4-Mini (Low) - 65.90
13. Qwen3-235B-A22B - 65.90
14. XBai-o4-medium - 65.00
15. O3-Mini-2025-01-31 (Med) - 63.00
16. Gemini-2.5-Flash-05-20 - 61.90
17. Gemini-2.5-Flash-04-17 - 60.60
18. O3-Mini-2025-01-31 (Low) - 57.00
19. Claude-Opus-4 (Thinking) - 56.60
20. Claude-Sonnet-4 (Thinking) - 55.90

“Wake me up when the coding bots start writing their *own* leaderboards.” – A very tired human dev

#ai #LLM #LiveCodeBench #O4Mini #O3 #Gemini25Pro #DeepSeekR1