SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 4 days ago • 95
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 4 days ago • 96
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 4 days ago • 139
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement Paper • 2609.14857 • Published 7 days ago • 209
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening Paper • 2609.18708 • Published 5 days ago • 65
ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments Paper • 2609.19134 • Published 5 days ago • 82
LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence Paper • 2609.17488 • Published 6 days ago • 419
ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents Paper • 2609.17523 • Published 6 days ago • 30
Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction Paper • 2609.13285 • Published 13 days ago • 77
PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models Paper • 2609.14973 • Published 7 days ago • 171
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 7 days ago • 239
Continual Learning Mechanisms Compose for Long-Horizon Memorization Paper • 2609.06986 • Published 14 days ago • 367
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search Paper • 2609.13356 • Published 10 days ago • 256
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation Paper • 2609.11638 • Published 11 days ago • 691
DataFlex-RL: An Evaluation Platform for RLVR Data Policies Paper • 2609.06107 • Published 16 days ago • 163