SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning Paper • 2608.14277 • Published 14 days ago • 35
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning Paper • 2608.14290 • Published 14 days ago • 33
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Paper • 2608.08160 • Published 20 days ago • 29
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering Paper • 2607.28568 • Published 29 days ago • 185
HumanCLAW: Can Vision-Language Models Act Through a Body? Paper • 2607.27180 • Published 30 days ago • 76
PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails Paper • 2607.05910 • Published Jul 7 • 38
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published Jun 23 • 159
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? Paper • 2606.24530 • Published Jun 23 • 66
EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions Paper • 2606.23654 • Published Jun 22 • 80
MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling Paper • 2606.13473 • Published Jun 11 • 96
FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents Paper • 2606.12087 • Published Jun 10 • 79
ComBench: A Benchmark for Rigorous Proof Reasoning and Constructive Realization in Olympiad-Level Combinatorics Paper • 2606.10479 • Published Jun 9 • 20
SceneCode: Executable World Programs for Editable Indoor Scenes with Articulated Objects Paper • 2605.19587 • Published May 19 • 10
Draft-OPD: On-Policy Distillation for Speculative Draft Models Paper • 2605.29343 • Published May 28 • 37
COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation Paper • 2605.31264 • Published May 29 • 131