npub1ns…8ahke on Nostr: Seems like a bad study if only 44% can get it correct with or without LLMs. ...
Seems like a bad study if only 44% can get it correct with or without LLMs. Especially if you tell them to use a bad model you know will fail on subjects they know nothing about. Paying them incentivized them to care slightly more but clearly not enough to abandon the tool you asked them to use. 😬
