Mark Reveley

reinforcement

1 quote filed under model / training / reinforcement, newest first.

If you put all these things together:- RLHF = training the model to be likable by humans- RLVR = training the model to be accepted by machines- RLVR is more scalable- "Alignment tax" says "likable by humans" makes the model do worse on verifiable tasks

Source · Kun Chen · 2 August 2026