I did not use 5.3 too much so I cannot comment. My main test now, is to see ML training performance.
With 5.4 I made it modify an LM to a different architecture and purpose. It needed tons of hand-holding and sucked.
With 5.5 I gave it a bit simpler task, where it could use more off-the-shelf solutions, but model stability was still hard.
One thing I noticed is that 5.4 struggled figuring out an environment quirk, while 5.5 easily solved it.
I’ll test 5.6 on the same task I gave to 5.5 soon, and maybe after that the 5.4 task.

