Mark Reveley

Kun Chen

If you put all these things together:- RLHF = training the model to be likable by humans- RLVR = training the model to be accepted by machines- RLVR is more scalable- "Alignment tax" says "likable by humans" makes the model do worse on verifiable tasks

Source · Kun Chen · 2 August 2026