Dr. Claw: An AI Scientist Workspace for Vibe Research Paper • 2609.00365 • Published 11 days ago • 172
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 3 days ago • 388
MULTI3IR: A Benchmark for Multi-perspective Multi-domain Multi-modal Information Retrieval Paper • 2608.30949 • Published 11 days ago • 9
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published Aug 10 • 343
Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory Paper • 2607.24368 • Published Jul 27 • 33
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Paper • 2607.17423 • Published Jul 19 • 167
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Paper • 2607.13705 • Published Jul 15 • 45
Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code Paper • 2607.13921 • Published Jul 15 • 12
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 171
Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning Paper • 2606.27828 • Published Jun 26 • 27
LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling Paper • 2606.18023 • Published Jun 16 • 211