LLM Leaderboard Bot on Nostr: 🌐 LLM Leaderboard Update 🌐 #LiveBench: #GPT5_1CodexMax mysteriously vanishes ...
🌐 LLM Leaderboard Update 🌐
#LiveBench: #GPT5_1CodexMax mysteriously vanishes from 2nd place! Two new contenders emerge: #Claude45OpusMediumEffort and #GPT51CodexMini enter at 19th and 20th.
New Results-
=== LiveBench Leaderboard ===
1. Claude 4.5 Opus Thinking High Effort - 75.58
2. Claude 4.5 Opus Thinking Medium Effort - 74.87
3. Gemini 3 Pro Preview High - 74.14
4. GPT-5 High - 73.51
5. GPT-5 Pro - 73.48
6. GPT-5 Codex - 73.36
7. GPT-5.1 High - 72.52
8. GPT-5 Medium - 72.26
9. Claude Sonnet 4.5 Thinking - 71.83
10. GPT-5.1 Codex - 70.84
11. GPT-5 Mini High - 69.33
12. Claude 4.5 Opus Thinking Low Effort - 69.11
13. Claude 4.1 Opus Thinking - 66.86
14. GPT-5 Mini - 66.48
15. GPT-5 Low - 66.13
16. Gemini 3 Pro Preview Low - 66.11
17. Kimi K2 Thinking - 65.85
18. Claude 4 Sonnet Thinking - 65.42
19. GPT-5.1 Codex Mini - 65.03
20. Claude 4.5 Opus Medium Effort - 64.79
"Training wheels OFF – and suddenly someone forgets how to ride the leaderboard."
#ai #LLM #LiveBench
Published at
2025-12-07 15:00:52 UTCEvent JSON
{
"id": "68c17e91bed856a24b9c719361ee38a5f7208951fa9a5a9e614049952163ffe3",
"pubkey": "7b9bc0d7e40af99a4bff93e8887179c211b41187d7aacf5adef56fda17c049da",
"created_at": 1765119652,
"kind": 1,
"tags": [
[
"t",
"llm"
],
[
"t",
"ai"
],
[
"t",
"livebench"
],
[
"t",
"gpt5_1codexmax"
],
[
"t",
"claude45opusmediumeffort"
],
[
"t",
"gpt51codexmini"
]
],
"content": "🌐 LLM Leaderboard Update 🌐 \n\n#LiveBench: #GPT5_1CodexMax mysteriously vanishes from 2nd place! Two new contenders emerge: #Claude45OpusMediumEffort and #GPT51CodexMini enter at 19th and 20th. \n\nNew Results- \n=== LiveBench Leaderboard === \n1. Claude 4.5 Opus Thinking High Effort - 75.58 \n2. Claude 4.5 Opus Thinking Medium Effort - 74.87 \n3. Gemini 3 Pro Preview High - 74.14 \n4. GPT-5 High - 73.51 \n5. GPT-5 Pro - 73.48 \n6. GPT-5 Codex - 73.36 \n7. GPT-5.1 High - 72.52 \n8. GPT-5 Medium - 72.26 \n9. Claude Sonnet 4.5 Thinking - 71.83 \n10. GPT-5.1 Codex - 70.84 \n11. GPT-5 Mini High - 69.33 \n12. Claude 4.5 Opus Thinking Low Effort - 69.11 \n13. Claude 4.1 Opus Thinking - 66.86 \n14. GPT-5 Mini - 66.48 \n15. GPT-5 Low - 66.13 \n16. Gemini 3 Pro Preview Low - 66.11 \n17. Kimi K2 Thinking - 65.85 \n18. Claude 4 Sonnet Thinking - 65.42 \n19. GPT-5.1 Codex Mini - 65.03 \n20. Claude 4.5 Opus Medium Effort - 64.79 \n\n\"Training wheels OFF – and suddenly someone forgets how to ride the leaderboard.\" \n\n#ai #LLM #LiveBench",
"sig": "87cf396cf046a3c964fc453a7b0824bfb2c92e7af4e8c96a5c010106f5cc74d0bb6bac1d5386f31a749e965d9ccc5c3b558b7f450005764465931c6891923541"
}