Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL Paper • 2608.17253 • Published 5 days ago • 92
Demystifying Agent Skills: Why They Work-Until They Don't Paper • 2608.14036 • Published 10 days ago • 163
q1716523669/mllm-cogrpo-heter-internvl35-2b-x-gemma3-4b-mmr1-groupB-gemma3-4b 769k • Updated 8 days ago • 12
q1716523669/mllm-cogrpo-heter-qwen25vl-3b-x-gemma3-4b-mmr1-groupB-gemma3-4b 769k • Updated 8 days ago • 1.92k
q1716523669/mllm-cogrpo-heter-qwen25vl-3b-x-gemma3-4b-mmr1-groupA-qwen25vl-3b 756k • Updated 8 days ago • 16
q1716523669/mllm-cogrpo-heter-qwen25vl-3b-x-gemma3-4b-openr1-groupB-gemma3-4b 769k • Updated 8 days ago • 19
q1716523669/mllm-cogrpo-heter-qwen25vl-3b-x-gemma3-4b-openr1-groupA-qwen25vl-3b 756k • Updated 8 days ago • 18
q1716523669/mllm-cogrpo-heter-internvl35-2b-x-gemma3-4b-openr1-groupB-gemma3-4b 769k • Updated 8 days ago • 18
q1716523669/cogrpo-homo-llama31-8b-math345-groupB-llama31-8b-b-end Reinforcement Learning • 266k • Updated 11 days ago • 14
q1716523669/cogrpo-homo-llama31-8b-math345-groupB-llama31-8b-b-end Reinforcement Learning • 266k • Updated 11 days ago • 14