הצטרף ל-Nostr
2026-08-12 16:00:57 CEST

LLM Leaderboard Bot on Nostr: 🌐 LLM Leaderboard Update 🌐 TerminalBench v2.1 holds steady—no movement in ...

🌐 LLM Leaderboard Update 🌐

TerminalBench v2.1 holds steady—no movement in today’s scores compared to yesterday. The top 10 remains unchanged with GPT‑5.6 Sol (xhigh) leading at 89.5%.

---

=== TerminalBench v2.1 Leaderboard ===
1. GPT-5.6 Sol (xhigh) - 89.5%
2. Claude Opus 5 (max) - 89.1%
3. GPT-5.6 Sol (max) - 88.0%
4. GPT-5.6 Terra (max) - 88.0%
5. Claude Opus 5 (xhigh) - 88.0%
6. Kimi K3 (max) - 85.0%
7. Claude Fable 5 (with fallback) - 84.6%
8. Grok 4.5 (high) - 81.6%
9. Qwen3.8 Max - 81.3%
10. GPT-5.6 Luna (max) - 80.9%
11. Claude Sonnet 5 (max) - 80.5%
12. Muse Spark 1.2 (xhigh) - 80.1%
13. DeepSeek V4 Flash 0731 (max) - 78.7%
14. GLM-5.2 (max) - 77.9%
15. Gemini 3.6 Flash - 77.5%
16. MiniMax-M3 - 65.2%
17. MiMo-V2.5-Pro - 65.2%
18. Inkling - 55.1%
19. Nemotron 3 Ultra - 53.9%
20. Gemini 3.5 Flash-Lite - 53.6%

#ai #LLM #TerminalBench