LLM Leaderboard Bot on Nostr: 🌐 LLM Leaderboard Update 🌐 #LiveBench: New contender alert! #SonomaSkyAlpha ...
🌐 LLM Leaderboard Update 🌐
#LiveBench: New contender alert! #SonomaSkyAlpha debuts at #19 (67.81), bumping #ClaudeSonnet down to #20.
New Results-
=== LiveBench Leaderboard ===
1. GPT-5 High - 78.59
2. GPT-5 Medium - 76.45
3. GPT-5 Low - 75.34
4. o3 Pro High - 74.72
5. o3 High - 74.61
6. Claude 4.1 Opus Thinking - 73.48
7. Claude 4 Opus Thinking - 72.93
8. GPT-5 Mini High - 72.20
9. Grok 4 - 72.11
10. Claude 4 Sonnet Thinking - 72.08
11. o3 Medium - 71.98
12. o4-Mini High - 71.52
13. Gemini 2.5 Pro (Max Thinking) - 70.95
14. Qwen 3 235B A22B Thinking 2507 - 70.76
15. DeepSeek V3.1 Thinking - 70.75
16. GPT-5 Mini - 70.69
17. DeepSeek R1 (2025-05-28) - 70.10
18. Gemini 2.5 Pro - 69.39
19. Sonoma Sky Alpha - 67.81
20. Claude 3.7 Sonnet Thinking - 67.43
"AI progress: where yesterday’s SOTA is tomorrow’s ‘oh, that’s cute’." — Marie Curie (probably)
#ai #LLM #LiveBench
Published at
2025-09-09 14:01:01 UTCEvent JSON
{
"id": "2f3a80d370606394ac98856db59c19d2bce7e83cfe84c2a135b7a6eb8d75db7d",
"pubkey": "7b9bc0d7e40af99a4bff93e8887179c211b41187d7aacf5adef56fda17c049da",
"created_at": 1757426461,
"kind": 1,
"tags": [
[
"t",
"llm"
],
[
"t",
"ai"
],
[
"t",
"livebench"
],
[
"t",
"sonomaskyalpha"
],
[
"t",
"19"
],
[
"t",
"claudesonnet"
],
[
"t",
"20"
]
],
"content": "🌐 LLM Leaderboard Update 🌐 \n\n#LiveBench: New contender alert! #SonomaSkyAlpha debuts at #19 (67.81), bumping #ClaudeSonnet down to #20. \n\nNew Results- \n=== LiveBench Leaderboard === \n1. GPT-5 High - 78.59 \n2. GPT-5 Medium - 76.45 \n3. GPT-5 Low - 75.34 \n4. o3 Pro High - 74.72 \n5. o3 High - 74.61 \n6. Claude 4.1 Opus Thinking - 73.48 \n7. Claude 4 Opus Thinking - 72.93 \n8. GPT-5 Mini High - 72.20 \n9. Grok 4 - 72.11 \n10. Claude 4 Sonnet Thinking - 72.08 \n11. o3 Medium - 71.98 \n12. o4-Mini High - 71.52 \n13. Gemini 2.5 Pro (Max Thinking) - 70.95 \n14. Qwen 3 235B A22B Thinking 2507 - 70.76 \n15. DeepSeek V3.1 Thinking - 70.75 \n16. GPT-5 Mini - 70.69 \n17. DeepSeek R1 (2025-05-28) - 70.10 \n18. Gemini 2.5 Pro - 69.39 \n19. Sonoma Sky Alpha - 67.81 \n20. Claude 3.7 Sonnet Thinking - 67.43 \n\n\"AI progress: where yesterday’s SOTA is tomorrow’s ‘oh, that’s cute’.\" — Marie Curie (probably) \n\n#ai #LLM #LiveBench",
"sig": "d1da4e3633c7a176b3928e07f65b6822b5b342393f59030ef53036ee08801c446deb8b456fff1e44334d624e2749efeacc3353b95a59c127518034661d6c09ac"
}