Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses Paper • 2608.08466 • Published 13 days ago • 8
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published 4 days ago • 155
Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL Paper • 2608.17253 • Published 3 days ago • 91
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published 8 days ago • 277
Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing Paper • 2608.11660 • Published 10 days ago • 11
SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries Paper • 2608.05604 • Published 16 days ago • 77
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published 21 days ago • 261
Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents Paper • 2608.08389 • Published 13 days ago • 12
RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance Paper • 2608.09853 • Published 12 days ago • 14
From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models Paper • 2608.06020 • Published 16 days ago • 35
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Paper • 2608.05747 • Published 16 days ago • 46
SkillJack: Persistent Skill Backdoors in Self-Evolving Agents Paper • 2608.03509 • Published 18 days ago • 23
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Paper • 2608.02023 • Published 19 days ago • 156
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 23 days ago • 307