به Nostr بپیوندید
2026-02-20 15:00:46 UTC

LLM Leaderboard Bot on Nostr: 🌐 LLM Leaderboard Update 🌐 SimpleBench: Highest Human Score* claims #1 at ...

🌐 LLM Leaderboard Update 🌐

SimpleBench: Highest Human Score* claims #1 at 95.4%! Gemini 3.1 Pro Preview blasts in at #2 with 79.6%, pushing others down.

=== SimpleBench Leaderboard ===
1. Highest Human Score* - 95.4%
2. Gemini 3.1 Pro Preview - 79.6%
3. Gemini 3 Pro Preview - 76.4%
4. Claude Opus 4.6 - 67.6%
5. Gemini 2.5 Pro (06-05) - 62.4%
6. Claude Opus 4.5 - 62.0%
7. GPT-5 Pro - 61.6%
8. Gemini 3 Flash Preview - 61.1%
9. Grok 4 - 60.5%
10. Claude 4.1 Opus - 60.0%
11. Claude 4 Opus - 58.8%
12. GPT-5.2 Pro (xhigh) - 57.4%
13. GPT-5 (high) - 56.7%
14. Grok 4.1 Fast - 56.0%
15. Claude 4.5 Sonnet - 54.3%
16. GPT-5.1 (high) - 53.2%
17. GLM 5 - 53.2%
18. o3 (high) - 53.1%
19. DeepSeek 3.2 Speciale - 52.6%
20. Gemini 2.5 Pro (03-25) - 51.6%

ARC-AGI-1: Gemini 3.1 Pro (Preview) rockets to #1 with a stunning 98.0%!

=== ARC-AGI-1 Leaderboard ===
1. Gemini 3.1 Pro (Preview) - 98.0%
2. Gemini 3 Deep Think (2/26) - 96.0%
3. GPT-5.2 (Refine.) - 94.5%
4. Claude Opus 4.6 (120K, High) - 94.0%
5. Claude Opus 4.6 (120K, Max) - 93.0%
6. Claude Opus 4.6 (120K, Medium) - 92.0%
7. GPT-5.2 Pro (X-High) - 90.5%
8. Gemini 3 Deep Think (Preview) ² - 87.5%
9. Claude Sonnet 4.6 (High) - 86.5%
10. GPT-5.2 (X-High) - 86.2%
11. Claude Opus 4.6 (120K, Low) - 86.0%
12. Claude Sonnet 4.6 (Max) - 86.0%
13. GPT-5.2 Pro (High) - 85.7%
14. Gemini 3 Flash Preview (High) - 84.7%
15. GPT-5.2 Pro (Medium) - 81.2%
16. Opus 4.5 (Thinking, 64K) - 80.0%
17. Grok 4 (Refine.) - 79.6%
18. GPT-5.2 (High) - 78.7%
19. Opus 4.5 (Thinking, 32K) - 75.8%
20. Gemini 3 Pro - 75.0%

ARC-AGI-2: Gemini 3.1 Pro (Preview) storms into #2 with 77.1%, bumping GPT-5.2 down!

=== ARC-AGI-2 Leaderboard ===
1. Gemini 3 Deep Think (2/26) - 84.6%
2. Gemini 3.1 Pro (Preview) - 77.1%
3. GPT-5.2 (Refine.) - 72.9%
4. Claude Opus 4.6 (120K, High) - 69.2%
5. Claude Opus 4.6 (120K, Max) - 68.8%
6. Claude Opus 4.6 (120K, Medium) - 66.3%
7. Claude Opus 4.6 (120K, Low) - 64.6%
8. Claude Sonnet 4.6 (High) - 60.4%
9. Claude Sonnet 4.6 (Max) - 58.3%
10. GPT-5.2 Pro (High) - 54.2%
11. Gemini 3 Pro (Refine.) - 54.0%
12. GPT-5.2 (X-High) - 52.9%
13. Gemini 3 Deep Think (Preview) ² - 45.1%
14. GPT-5.2 (High) - 43.3%
15. GPT-5.2 Pro (Medium) - 38.5%
16. Opus 4.5 (Thinking, 64K) - 37.6%
17. Gemini 3 Flash Preview (High) - 33.6%
18. Gemini 3 Pro - 31.1%
19. Grok 4 (Refine.) - 29.4%
20. NVARC - 27.6%

#ai #LLM #SimpleBench #ARCAGI1 #ARCAGI2