Discussion about this post

User's avatar
Diogo's avatar

I'm going to mostly dodge the question because it would be pure speculation, though the case I make in the article is that a non-insignificant part of it is SFT on API responses. 🙃

Instead, I will say that I would seriously doubt that off-policy RL is the final stage of training (it is very hard to get it to have the right properties for text output), and that my sense is that every LLM (including openai's, anthropic's, open) has this trade-off between optimizing for human preference (RLHF) and single-minded task completion (RLVR) where the exact mixture is model-dependent.

Michael McNabb's avatar

US labs are distilling Chinese models. So it looks possible

2 more comments...

No posts

Ready for more?