PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents Paper • 2608.19861 • Published 10 days ago • 9
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives Paper • 2608.13552 • Published 17 days ago • 46
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published 29 days ago • 263
SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding Paper • 2608.05137 • Published 20 days ago • 27
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Paper • 2608.05747 • Published 24 days ago • 46
Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance Paper • 2608.00782 • Published 29 days ago • 16
PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents Paper • 2608.04003 • Published 25 days ago • 34
Progressive Agent Skill Generation via Reinforcement Learning Paper • 2608.01678 • Published 27 days ago • 60