RL checkpoints of Octopus-8B and baselines of paper: Learning Self-Correction in Vision–Language Models via Rollout Augmentation
Create images in seconds. No sign-up, no paywall, no setup.