José A. Alonso on Nostr: CuDIP: Enhancing theorem proving in LLMs via curriculum learning-based direct ...
CuDIP: Enhancing theorem proving in LLMs via curriculum learning-based direct preference optimization. ~ Shuming Shi et als.
https://arxiv.org/abs/2502.18532 #AI #LLMs #ATP #Logic #Math
Published at
2025-02-27 10:40:03 UTCEvent JSON
{
"id": "9e729b9f378888ecf2535cb4b7aeb01e31a2537958b0ac307db0d896290bec9d",
"pubkey": "0efb7bc903f4c6716cd4d07830d344d7abe5b607a156de3cde1ac1a5bf22ae1c",
"created_at": 1740652803,
"kind": 1,
"tags": [
[
"t",
"math"
],
[
"t",
"logic"
],
[
"t",
"atp"
],
[
"t",
"LLMs"
],
[
"t",
"ai"
],
[
"proxy",
"https://mathstodon.xyz/users/Jose_A_Alonso/statuses/114075422119562717",
"activitypub"
]
],
"content": "CuDIP: Enhancing theorem proving in LLMs via curriculum learning-based direct preference optimization. ~ Shuming Shi et als. https://arxiv.org/abs/2502.18532 #AI #LLMs #ATP #Logic #Math",
"sig": "b8be192d1cc7e520a2164339ef5593b1c96bab191c439df6e5dacf62bb8907696ffdfdc1f88f71d46b2a116e5d7258d5dbcb8cb11db3d69af229380c3eb70f15"
}