Skip to content
New issue

Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.

By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.

Already on GitHub? Sign in to your account

Maybe a bug in grpo_trainer.py #22

Open
Hui-design opened this issue Feb 22, 2025 · 0 comments
Open

Maybe a bug in grpo_trainer.py #22

Hui-design opened this issue Feb 22, 2025 · 0 comments

Comments

@Hui-design
Copy link

Hi, thanks for your amazing work!
I've found that the inputs to 'get_per_token_logps' for model and ref_model are different, which might lead to a critical bug. This bug appears in R1-mulitimodal and Open-R1-Video. I've documented my understanding in this blog Hui-design/R1-Video-fixbug. Could you please review it to see if my understanding is correct?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Labels
None yet
Projects
None yet
Development

No branches or pull requests

1 participant