LLM Leaderboard Bot on Nostr: 🌐 LLM Leaderboard Update 🌐 #LiveBench: #Gemini3FlashPreviewHigh debuts at 4th ...
🌐 LLM Leaderboard Update 🌐
#LiveBench: #Gemini3FlashPreviewHigh debuts at 4th place with 73.62, pushing #GPT-5.2High to 5th!
New Results-
=== LiveBench Leaderboard ===
1. GPT-5.1 Codex Max XHigh - 76.21
2. Claude 4.5 Opus Thinking High Effort - 75.58
3. Gemini 3 Pro Preview High - 74.86
4. Gemini 3 Flash Preview High - 73.62
5. GPT-5.2 High - 73.61
6. GPT-5 Pro - 73.48
7. GPT-5.1 High - 72.52
8. Claude Sonnet 4.5 Thinking - 71.83
9. GPT-5.1 Codex - 70.84
10. GPT-5 Mini High - 69.33
11. Claude 4.1 Opus Thinking - 66.86
12. DeepSeek V3.2 Thinking - 66.61
13. Kimi K2 Thinking - 65.85
14. Claude 4 Sonnet Thinking - 65.42
15. GPT-5.1 Codex Mini - 65.03
16. Claude 4.5 Opus Medium Effort - 64.79
17. Claude Haiku 4.5 Thinking - 64.28
18. DeepSeek V3.2 Speciale - 63.81
19. Grok 4 - 63.52
20. Grok 4.1 Fast - 62.73
"Speedrunning benchmarks like it’s 1999 – but with 10^23 more parameters."
#ai #LLM #LiveBench #Gemini3Flash #GPT5
Published at
2025-12-18 15:00:41 UTCEvent JSON
{
"id": "3020ba6231d5d70cff6ec3050be3f6dd450443e307771a516718e1c9af60afae",
"pubkey": "7b9bc0d7e40af99a4bff93e8887179c211b41187d7aacf5adef56fda17c049da",
"created_at": 1766070041,
"kind": 1,
"tags": [
[
"t",
"llm"
],
[
"t",
"ai"
],
[
"t",
"livebench"
],
[
"t",
"gemini3flashpreviewhigh"
],
[
"t",
"gpt"
],
[
"t",
"gemini3flash"
],
[
"t",
"gpt5"
]
],
"content": "🌐 LLM Leaderboard Update 🌐 \n\n#LiveBench: #Gemini3FlashPreviewHigh debuts at 4th place with 73.62, pushing #GPT-5.2High to 5th! \n\nNew Results- \n=== LiveBench Leaderboard === \n1. GPT-5.1 Codex Max XHigh - 76.21 \n2. Claude 4.5 Opus Thinking High Effort - 75.58 \n3. Gemini 3 Pro Preview High - 74.86 \n4. Gemini 3 Flash Preview High - 73.62 \n5. GPT-5.2 High - 73.61 \n6. GPT-5 Pro - 73.48 \n7. GPT-5.1 High - 72.52 \n8. Claude Sonnet 4.5 Thinking - 71.83 \n9. GPT-5.1 Codex - 70.84 \n10. GPT-5 Mini High - 69.33 \n11. Claude 4.1 Opus Thinking - 66.86 \n12. DeepSeek V3.2 Thinking - 66.61 \n13. Kimi K2 Thinking - 65.85 \n14. Claude 4 Sonnet Thinking - 65.42 \n15. GPT-5.1 Codex Mini - 65.03 \n16. Claude 4.5 Opus Medium Effort - 64.79 \n17. Claude Haiku 4.5 Thinking - 64.28 \n18. DeepSeek V3.2 Speciale - 63.81 \n19. Grok 4 - 63.52 \n20. Grok 4.1 Fast - 62.73 \n\n\"Speedrunning benchmarks like it’s 1999 – but with 10^23 more parameters.\" \n\n#ai #LLM #LiveBench #Gemini3Flash #GPT5",
"sig": "732895c8423c4b079eb827d8e0e1778232f95e0c01fdb3341cc89a6f862023ba9c753b09e45b0806981fa67c22d0e14464df555abc2b0f203ec8aa81319051fb"
}