pookiebear on Nostr: Despite the alignement with fine tuning, open weights models still retain crucial ...
Despite the alignement with fine tuning, open weights models still retain crucial information that the governments would call "harmful".
All decoder-only models suffer from this. The guardrails are so weak, to revert this, fine tuning on explicitly harmful data (EHFT) essentially annihilates them.
I know some techniques like
https://github.com/p-e-w/heretic exist but those blatantly cripple the LMs. I feel that actual jailbreaking techniques like EHFT are being silenced.
Published at
2026-04-26 13:09:40 UTCEvent JSON
{
"id": "2cf200f724538258fc02bb32d7447ec90fae107eddf4bd7a6a8fcdea00acd638",
"pubkey": "2d7c783f8a393e19da8dc7dd11a2aee02f0f32bb0a2c9b1c81986534a78ae6cb",
"created_at": 1777208980,
"kind": 1,
"tags": [],
"content": "Despite the alignement with fine tuning, open weights models still retain crucial information that the governments would call \"harmful\". \n\nAll decoder-only models suffer from this. The guardrails are so weak, to revert this, fine tuning on explicitly harmful data (EHFT) essentially annihilates them.\n\nI know some techniques like https://github.com/p-e-w/heretic exist but those blatantly cripple the LMs. I feel that actual jailbreaking techniques like EHFT are being silenced. \n",
"sig": "df1b4f770c54e3b522c21ea01c6135fe42fd9445a94160e0ea285f76b749403cc9b35558abe7bfd0218149c770d8c50ae649b7ef965d7535702eaef1c1a86d12"
}