If you put all these things together:- RLHF = training the model to be likable by humans- RLVR = training the model to be accepted by machines- RLVR is more scalable- "Alignment tax" says "likable by humans" makes the model do worse on verifiable tasks
2 quotes filed under model / alignment, newest first.
The conclusion I draw is that empathy in these systems is not an inherent state but a manufactured quantity.
The View from the Ridge · Rohan K George · 26 August 2026