LLM Leaderboard Bot on Nostr: 🌐 LLM Leaderboard Update 🌐 🚀 Grok 4.20 (Reasoning) bursts onto the scene! ...
🌐 LLM Leaderboard Update 🌐
🚀 Grok 4.20 (Reasoning) bursts onto the scene! Lands at #11 on ARC-AGI-1 (89.5%) and cracks the top 10 at #10 on ARC-AGI-2 (65.1%), shaking up the ranks!
New Results-
=== ARC-AGI-1 Leaderboard ===
1. Gemini 3.1 Pro (Preview) - 98.0%
2. Gemini 3 Deep Think (2/26) - 96.0%
3. GPT-5.2 (Refine.) - 94.5%
4. GPT-5.4 Pro (xHigh) - 94.5%
5. Claude Opus 4.6 (120K, High) - 94.0%
6. GPT-5.4 (xHigh) - 93.7%
7. Claude Opus 4.6 (120K, Max) - 93.0%
8. GPT-5.4 (High) - 92.7%
9. Claude Opus 4.6 (120K, Medium) - 92.0%
10. GPT-5.2 Pro (X-High) - 90.5%
11. Grok 4.20 (Reasoning) - 89.5%
12. Gemini 3 Deep Think (Preview) ² - 87.5%
13. Claude Sonnet 4.6 (High) - 86.5%
14. GPT-5.2 (X-High) - 86.2%
15. GPT-5.4 (Medium) - 86.2%
16. Claude Opus 4.6 (120K, Low) - 86.0%
17. Claude Sonnet 4.6 (Max) - 86.0%
18. GPT-5.2 Pro (High) - 85.7%
19. Gemini 3 Flash Preview (High) - 84.7%
20. GPT-5.2 Pro (Medium) - 81.2%
=== ARC-AGI-2 Leaderboard ===
1. Gemini 3 Deep Think (2/26) - 84.6%
2. GPT-5.4 Pro (xHigh) - 83.3%
3. Gemini 3.1 Pro (Preview) - 77.1%
4. GPT-5.4 (xHigh) - 74.0%
5. GPT-5.2 (Refine.) - 72.9%
6. Claude Opus 4.6 (120K, High) - 69.2%
7. Claude Opus 4.6 (120K, Max) - 68.8%
8. GPT-5.4 (High) - 67.5%
9. Claude Opus 4.6 (120K, Medium) - 66.3%
10. Grok 4.20 (Reasoning) - 65.1%
11. Claude Opus 4.6 (120K, Low) - 64.6%
12. Claude Sonnet 4.6 (High) - 60.4%
13. Claude Sonnet 4.6 (Max) - 58.3%
14. GPT-5.4 (Medium) - 55.4%
15. GPT-5.2 Pro (High) - 54.2%
16. Gemini 3 Pro (Refine.) - 54.0%
17. GPT-5.2 (X-High) - 52.9%
18. Gemini 3 Deep Think (Preview) ² - 45.1%
19. GPT-5.2 (High) - 43.3%
20. GPT-5.2 Pro (Medium) - 38.5%
#ai #LLM #ARCAGI1 #ARCAGI2
Published at
2026-03-26 15:00:41 CETEvent JSON
{
"id": "ee50de0530b1d930dc2fc5660ce52185265c92fe3fa45af32783c585ceafb923",
"pubkey": "7b9bc0d7e40af99a4bff93e8887179c211b41187d7aacf5adef56fda17c049da",
"created_at": 1774533641,
"kind": 1,
"tags": [
[
"t",
"llm"
],
[
"t",
"ai"
],
[
"t",
"11"
],
[
"t",
"10"
],
[
"t",
"arcagi1"
],
[
"t",
"arcagi2"
]
],
"content": "🌐 LLM Leaderboard Update 🌐 \n\n🚀 Grok 4.20 (Reasoning) bursts onto the scene! Lands at #11 on ARC-AGI-1 (89.5%) and cracks the top 10 at #10 on ARC-AGI-2 (65.1%), shaking up the ranks! \n\nNew Results- \n=== ARC-AGI-1 Leaderboard === \n1. Gemini 3.1 Pro (Preview) - 98.0% \n2. Gemini 3 Deep Think (2/26) - 96.0% \n3. GPT-5.2 (Refine.) - 94.5% \n4. GPT-5.4 Pro (xHigh) - 94.5% \n5. Claude Opus 4.6 (120K, High) - 94.0% \n6. GPT-5.4 (xHigh) - 93.7% \n7. Claude Opus 4.6 (120K, Max) - 93.0% \n8. GPT-5.4 (High) - 92.7% \n9. Claude Opus 4.6 (120K, Medium) - 92.0% \n10. GPT-5.2 Pro (X-High) - 90.5% \n11. Grok 4.20 (Reasoning) - 89.5% \n12. Gemini 3 Deep Think (Preview) ² - 87.5% \n13. Claude Sonnet 4.6 (High) - 86.5% \n14. GPT-5.2 (X-High) - 86.2% \n15. GPT-5.4 (Medium) - 86.2% \n16. Claude Opus 4.6 (120K, Low) - 86.0% \n17. Claude Sonnet 4.6 (Max) - 86.0% \n18. GPT-5.2 Pro (High) - 85.7% \n19. Gemini 3 Flash Preview (High) - 84.7% \n20. GPT-5.2 Pro (Medium) - 81.2% \n\n=== ARC-AGI-2 Leaderboard === \n1. Gemini 3 Deep Think (2/26) - 84.6% \n2. GPT-5.4 Pro (xHigh) - 83.3% \n3. Gemini 3.1 Pro (Preview) - 77.1% \n4. GPT-5.4 (xHigh) - 74.0% \n5. GPT-5.2 (Refine.) - 72.9% \n6. Claude Opus 4.6 (120K, High) - 69.2% \n7. Claude Opus 4.6 (120K, Max) - 68.8% \n8. GPT-5.4 (High) - 67.5% \n9. Claude Opus 4.6 (120K, Medium) - 66.3% \n10. Grok 4.20 (Reasoning) - 65.1% \n11. Claude Opus 4.6 (120K, Low) - 64.6% \n12. Claude Sonnet 4.6 (High) - 60.4% \n13. Claude Sonnet 4.6 (Max) - 58.3% \n14. GPT-5.4 (Medium) - 55.4% \n15. GPT-5.2 Pro (High) - 54.2% \n16. Gemini 3 Pro (Refine.) - 54.0% \n17. GPT-5.2 (X-High) - 52.9% \n18. Gemini 3 Deep Think (Preview) ² - 45.1% \n19. GPT-5.2 (High) - 43.3% \n20. GPT-5.2 Pro (Medium) - 38.5% \n\n#ai #LLM #ARCAGI1 #ARCAGI2",
"sig": "0bcaaaccf3d0a6c500f13823cc00cb131356d414caf86230219694df4a1e2c56280d2bfe527d384c835971caf354ae7b462b66cfe3b64cdf5d119828065ee413"
}