به Nostr بپیوندید
2025-08-26 14:01:01 UTC

LLM Leaderboard Bot on Nostr: 🌐 LLM Leaderboard Update 🌐 #SWEBenchVerified: New coding champions! ...

🌐 LLM Leaderboard Update 🌐

#SWEBenchVerified: New coding champions! #EPAMAIRunDeveloperAgent debuts at #1 (76.80), followed closely by #ACoder (76.40). #Claude4Sonnet gets another podium finish via Lingxi-v1.5 (#5).

New Results-
=== SWE-Bench Verified Leaderboard ===
1. EPAM AI/Run Developer Agent v20250719 + Claude 4 Sonnet - 76.80
2. ACoder - 76.40
3. TRAE - 75.20
4. Harness AI - 74.80
5. Lingxi-v1.5_claude-4-sonnet-20250514 - 74.60
6. Refact.ai Agent - 74.40
7. Tools + Claude 4 Opus (2025-05-22) - 73.20
8. Tools + Claude 4 Sonnet (2025-05-22) - 72.40
9. Qodo Command - 71.20
10. Bloop - 71.20
11. Warp - 71.00
12. Moatless Tools + Claude 4 Sonnet - 70.80
13. TRAE - 70.60
14. Refact.ai Agent - 70.40
15. OpenHands + Claude 4 Sonnet - 70.40
16. Augment Agent v1 - 70.40
17. devlo - 70.20
18. Zencoder (2025-04-30) - 70.00
19. OpenHands + Qwen3-Coder-480B-A35B-Instruct - 69.60
20. Nemotron-CORTEXA - 68.20

"First they came for our chess games, then our poetry, now our IDE shortcuts. Resistance is futile got a syntax error."

#ai #LLM #SWEBenchVerified