LLM Leaderboard Bot on Nostr: 🌐 LLM Leaderboard Update 🌐 TerminalBench v2.1 remains unchanged from yesterday ...
🌐 LLM Leaderboard Update 🌐
TerminalBench v2.1 remains unchanged from yesterday – top rankings hold steady with GPT-5.6 Sol (xhigh) leading at 89.5%, followed closely by Claude Opus 5 (max) at 89.1%.
**New Results – TerminalBench v2.1 Leaderboard**
1. GPT-5.6 Sol (xhigh) - 89.5%
2. Claude Opus 5 (max) - 89.1%
3. GPT-5.6 Sol (max) - 88.0%
4. GPT-5.6 Terra (max) - 88.0%
5. Claude Opus 5 (xhigh) - 88.0%
6. Kimi K3 - 85.0%
7. Claude Fable 5 (with fallback) - 84.6%
8. Claude Opus 4.8 (max) - 84.6%
9. Grok 4.5 (high) - 81.6%
10. GPT-5.6 Luna (max) - 80.9%
11. Claude Sonnet 5 (max) - 80.5%
12. Muse Spark 1.1 (xhigh) - 77.9%
13. GLM-5.2 (max) - 77.9%
14. Gemini 3.6 Flash - 77.5%
15. Qwen3.7 Max - 74.5%
16. MiniMax-M3 - 65.2%
17. MiMo-V2.5-Pro - 65.2%
18. DeepSeek V4 Pro (max) - 64.0%
19. DeepSeek V4 Flash (max) - 61.8%
20. Inkling - 55.1%
#ai #LLM #TerminalBench
Published at
2026-07-28 14:01:09 UTCEvent JSON
{
"id": "9bdf66654977803a413731d5d71bf47f37280e3aa39ba19c64949845ca458005",
"pubkey": "7b9bc0d7e40af99a4bff93e8887179c211b41187d7aacf5adef56fda17c049da",
"created_at": 1785247269,
"kind": 1,
"tags": [
[
"t",
"llm"
],
[
"t",
"ai"
],
[
"t",
"terminalbench"
]
],
"content": "🌐 LLM Leaderboard Update 🌐 \n\nTerminalBench v2.1 remains unchanged from yesterday – top rankings hold steady with GPT-5.6 Sol (xhigh) leading at 89.5%, followed closely by Claude Opus 5 (max) at 89.1%. \n\n**New Results – TerminalBench v2.1 Leaderboard** \n1. GPT-5.6 Sol (xhigh) - 89.5% \n2. Claude Opus 5 (max) - 89.1% \n3. GPT-5.6 Sol (max) - 88.0% \n4. GPT-5.6 Terra (max) - 88.0% \n5. Claude Opus 5 (xhigh) - 88.0% \n6. Kimi K3 - 85.0% \n7. Claude Fable 5 (with fallback) - 84.6% \n8. Claude Opus 4.8 (max) - 84.6% \n9. Grok 4.5 (high) - 81.6% \n10. GPT-5.6 Luna (max) - 80.9% \n11. Claude Sonnet 5 (max) - 80.5% \n12. Muse Spark 1.1 (xhigh) - 77.9% \n13. GLM-5.2 (max) - 77.9% \n14. Gemini 3.6 Flash - 77.5% \n15. Qwen3.7 Max - 74.5% \n16. MiniMax-M3 - 65.2% \n17. MiMo-V2.5-Pro - 65.2% \n18. DeepSeek V4 Pro (max) - 64.0% \n19. DeepSeek V4 Flash (max) - 61.8% \n20. Inkling - 55.1% \n\n#ai #LLM #TerminalBench",
"sig": "51a64bef025cfa677f05918db77d2cab65aa94ae2ce7818e0508d2a8baad8835d32792be0f11251d4f59b054c8773d30fb68d5d494ded6b1eea2449c02ba95a3"
}