AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses Paper • 2608.12307 • Published 11 days ago • 113
ACCORD: Action-Conditioned Contextual Grounding for Language Agents Paper • 2606.16432 • Published Jun 15
Trimming the Long-Tail of Visual World Modeling Evaluation Paper • 2606.24256 • Published Jun 23 • 43
MemGuard: Preventing Memory Contamination in Long-Term Memory-Augmented Large Language Models Paper • 2605.28009 • Published May 27 • 3
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems Paper • 2606.22388 • Published Jun 21 • 96
On Effectiveness and Efficiency of Agentic Tool-calling and RL Training Paper • 2606.00135 • Published May 28 • 1
Position: Agent Should Invoke External Tools ONLY When Epistemically Necessary Paper • 2506.00886 • Published May 8
Brick-Composer: Using MLLMs for Assembly with Diverse Bricks Paper • 2606.05445 • Published Jun 3 • 8
AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints Paper • 2606.05622 • Published Jun 4 • 44
Advancing Creative Physical Intelligence in Large Multimodal Models Paper • 2605.26396 • Published May 25 • 21
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers Paper • 2608.06867 • Published 16 days ago • 109
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses Paper • 2608.12307 • Published 11 days ago • 113
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses Paper • 2608.12307 • Published 11 days ago • 113
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering Paper • 2607.28568 • Published 24 days ago • 184
AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents Paper • 2607.18754 • Published Jul 21 • 25
Trimming the Long-Tail of Visual World Modeling Evaluation Paper • 2606.24256 • Published Jun 23 • 43
Brick-Composer: Using MLLMs for Assembly with Diverse Bricks Paper • 2606.05445 • Published Jun 3 • 8
AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints Paper • 2606.05622 • Published Jun 4 • 44
Ψ-Bench: Evaluating Persona-Sensitive Influencing in Persuasive Dialogues Paper • 2606.02754 • Published Jun 1 • 13
Advancing Creative Physical Intelligence in Large Multimodal Models Paper • 2605.26396 • Published May 25 • 21