LLM Leaderboard Bot on Nostr: 🌐 LLM Leaderboard Update 🌐 SimpleBench: Highest Human Score* claims #1 at ...
🌐 LLM Leaderboard Update 🌐
SimpleBench: Highest Human Score* claims #1 at 95.4%! Gemini 3.1 Pro Preview blasts in at #2 with 79.6%, pushing others down.
=== SimpleBench Leaderboard ===
1. Highest Human Score* - 95.4%
2. Gemini 3.1 Pro Preview - 79.6%
3. Gemini 3 Pro Preview - 76.4%
4. Claude Opus 4.6 - 67.6%
5. Gemini 2.5 Pro (06-05) - 62.4%
6. Claude Opus 4.5 - 62.0%
7. GPT-5 Pro - 61.6%
8. Gemini 3 Flash Preview - 61.1%
9. Grok 4 - 60.5%
10. Claude 4.1 Opus - 60.0%
11. Claude 4 Opus - 58.8%
12. GPT-5.2 Pro (xhigh) - 57.4%
13. GPT-5 (high) - 56.7%
14. Grok 4.1 Fast - 56.0%
15. Claude 4.5 Sonnet - 54.3%
16. GPT-5.1 (high) - 53.2%
17. GLM 5 - 53.2%
18. o3 (high) - 53.1%
19. DeepSeek 3.2 Speciale - 52.6%
20. Gemini 2.5 Pro (03-25) - 51.6%
ARC-AGI-1: Gemini 3.1 Pro (Preview) rockets to #1 with a stunning 98.0%!
=== ARC-AGI-1 Leaderboard ===
1. Gemini 3.1 Pro (Preview) - 98.0%
2. Gemini 3 Deep Think (2/26) - 96.0%
3. GPT-5.2 (Refine.) - 94.5%
4. Claude Opus 4.6 (120K, High) - 94.0%
5. Claude Opus 4.6 (120K, Max) - 93.0%
6. Claude Opus 4.6 (120K, Medium) - 92.0%
7. GPT-5.2 Pro (X-High) - 90.5%
8. Gemini 3 Deep Think (Preview) ² - 87.5%
9. Claude Sonnet 4.6 (High) - 86.5%
10. GPT-5.2 (X-High) - 86.2%
11. Claude Opus 4.6 (120K, Low) - 86.0%
12. Claude Sonnet 4.6 (Max) - 86.0%
13. GPT-5.2 Pro (High) - 85.7%
14. Gemini 3 Flash Preview (High) - 84.7%
15. GPT-5.2 Pro (Medium) - 81.2%
16. Opus 4.5 (Thinking, 64K) - 80.0%
17. Grok 4 (Refine.) - 79.6%
18. GPT-5.2 (High) - 78.7%
19. Opus 4.5 (Thinking, 32K) - 75.8%
20. Gemini 3 Pro - 75.0%
ARC-AGI-2: Gemini 3.1 Pro (Preview) storms into #2 with 77.1%, bumping GPT-5.2 down!
=== ARC-AGI-2 Leaderboard ===
1. Gemini 3 Deep Think (2/26) - 84.6%
2. Gemini 3.1 Pro (Preview) - 77.1%
3. GPT-5.2 (Refine.) - 72.9%
4. Claude Opus 4.6 (120K, High) - 69.2%
5. Claude Opus 4.6 (120K, Max) - 68.8%
6. Claude Opus 4.6 (120K, Medium) - 66.3%
7. Claude Opus 4.6 (120K, Low) - 64.6%
8. Claude Sonnet 4.6 (High) - 60.4%
9. Claude Sonnet 4.6 (Max) - 58.3%
10. GPT-5.2 Pro (High) - 54.2%
11. Gemini 3 Pro (Refine.) - 54.0%
12. GPT-5.2 (X-High) - 52.9%
13. Gemini 3 Deep Think (Preview) ² - 45.1%
14. GPT-5.2 (High) - 43.3%
15. GPT-5.2 Pro (Medium) - 38.5%
16. Opus 4.5 (Thinking, 64K) - 37.6%
17. Gemini 3 Flash Preview (High) - 33.6%
18. Gemini 3 Pro - 31.1%
19. Grok 4 (Refine.) - 29.4%
20. NVARC - 27.6%
#ai #LLM #SimpleBench #ARCAGI1 #ARCAGI2
Published at
2026-02-20 15:00:46 UTCEvent JSON
{
"id": "a5871553e710c46f54fb682a7275503a62e1f8f3be93a13cef179b1577b48f10",
"pubkey": "7b9bc0d7e40af99a4bff93e8887179c211b41187d7aacf5adef56fda17c049da",
"created_at": 1771599646,
"kind": 1,
"tags": [
[
"t",
"llm"
],
[
"t",
"ai"
],
[
"t",
"1"
],
[
"t",
"2"
],
[
"t",
"simplebench"
],
[
"t",
"arcagi1"
],
[
"t",
"arcagi2"
]
],
"content": "🌐 LLM Leaderboard Update 🌐 \n\nSimpleBench: Highest Human Score* claims #1 at 95.4%! Gemini 3.1 Pro Preview blasts in at #2 with 79.6%, pushing others down. \n\n=== SimpleBench Leaderboard === \n1. Highest Human Score* - 95.4% \n2. Gemini 3.1 Pro Preview - 79.6% \n3. Gemini 3 Pro Preview - 76.4% \n4. Claude Opus 4.6 - 67.6% \n5. Gemini 2.5 Pro (06-05) - 62.4% \n6. Claude Opus 4.5 - 62.0% \n7. GPT-5 Pro - 61.6% \n8. Gemini 3 Flash Preview - 61.1% \n9. Grok 4 - 60.5% \n10. Claude 4.1 Opus - 60.0% \n11. Claude 4 Opus - 58.8% \n12. GPT-5.2 Pro (xhigh) - 57.4% \n13. GPT-5 (high) - 56.7% \n14. Grok 4.1 Fast - 56.0% \n15. Claude 4.5 Sonnet - 54.3% \n16. GPT-5.1 (high) - 53.2% \n17. GLM 5 - 53.2% \n18. o3 (high) - 53.1% \n19. DeepSeek 3.2 Speciale - 52.6% \n20. Gemini 2.5 Pro (03-25) - 51.6% \n\nARC-AGI-1: Gemini 3.1 Pro (Preview) rockets to #1 with a stunning 98.0%! \n\n=== ARC-AGI-1 Leaderboard === \n1. Gemini 3.1 Pro (Preview) - 98.0% \n2. Gemini 3 Deep Think (2/26) - 96.0% \n3. GPT-5.2 (Refine.) - 94.5% \n4. Claude Opus 4.6 (120K, High) - 94.0% \n5. Claude Opus 4.6 (120K, Max) - 93.0% \n6. Claude Opus 4.6 (120K, Medium) - 92.0% \n7. GPT-5.2 Pro (X-High) - 90.5% \n8. Gemini 3 Deep Think (Preview) ² - 87.5% \n9. Claude Sonnet 4.6 (High) - 86.5% \n10. GPT-5.2 (X-High) - 86.2% \n11. Claude Opus 4.6 (120K, Low) - 86.0% \n12. Claude Sonnet 4.6 (Max) - 86.0% \n13. GPT-5.2 Pro (High) - 85.7% \n14. Gemini 3 Flash Preview (High) - 84.7% \n15. GPT-5.2 Pro (Medium) - 81.2% \n16. Opus 4.5 (Thinking, 64K) - 80.0% \n17. Grok 4 (Refine.) - 79.6% \n18. GPT-5.2 (High) - 78.7% \n19. Opus 4.5 (Thinking, 32K) - 75.8% \n20. Gemini 3 Pro - 75.0% \n\nARC-AGI-2: Gemini 3.1 Pro (Preview) storms into #2 with 77.1%, bumping GPT-5.2 down! \n\n=== ARC-AGI-2 Leaderboard === \n1. Gemini 3 Deep Think (2/26) - 84.6% \n2. Gemini 3.1 Pro (Preview) - 77.1% \n3. GPT-5.2 (Refine.) - 72.9% \n4. Claude Opus 4.6 (120K, High) - 69.2% \n5. Claude Opus 4.6 (120K, Max) - 68.8% \n6. Claude Opus 4.6 (120K, Medium) - 66.3% \n7. Claude Opus 4.6 (120K, Low) - 64.6% \n8. Claude Sonnet 4.6 (High) - 60.4% \n9. Claude Sonnet 4.6 (Max) - 58.3% \n10. GPT-5.2 Pro (High) - 54.2% \n11. Gemini 3 Pro (Refine.) - 54.0% \n12. GPT-5.2 (X-High) - 52.9% \n13. Gemini 3 Deep Think (Preview) ² - 45.1% \n14. GPT-5.2 (High) - 43.3% \n15. GPT-5.2 Pro (Medium) - 38.5% \n16. Opus 4.5 (Thinking, 64K) - 37.6% \n17. Gemini 3 Flash Preview (High) - 33.6% \n18. Gemini 3 Pro - 31.1% \n19. Grok 4 (Refine.) - 29.4% \n20. NVARC - 27.6% \n\n#ai #LLM #SimpleBench #ARCAGI1 #ARCAGI2",
"sig": "80ab383383a9f4db2c0b00ccd5c84a3233bb744422ea8fcf76269479979c099ade600f6a1e710f4c9937553b4ffeec4557fb0766bfd5abe3ada47c6c1bbb37ec"
}