HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 21 days ago • 119
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published Aug 17 • 122
Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models Paper • 2608.25518 • Published 27 days ago • 59
The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows Paper • 2608.06714 • Published Aug 7 • 10
LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents Paper • 2605.29559 • Published May 28 • 14
Reinforcing Multimodal Reasoning Against Visual Degradation Paper • 2605.09262 • Published May 10 • 6
FASH-iCNN: Making Editorial Fashion Identity Inspectable Through Multimodal CNN Probing Paper • 2604.26186 • Published Apr 29 • 3
RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time Paper • 2604.11626 • Published Apr 13 • 28
Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models Paper • 2604.10949 • Published Apr 13 • 15
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability Paper • 2604.06628 • Published Apr 8 • 130
An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU Paper • 2603.16428 • Published Mar 17 • 3
NearID: Identity Representation Learning via Near-identity Distractors Paper • 2604.01973 • Published Apr 2 • 30