papers
updated
AI for Auto-Research: Roadmap & User Guide
Paper
• 2605.18661
• Published • 71
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data
Paper
• 2605.18287
• Published • 15
MixSD: Mixed Contextual Self-Distillation for Knowledge Injection
Paper
• 2605.16865
• Published • 10
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
Paper
• 2603.28069
• Published • 9
VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization
Paper
• 2602.09934
• Published • 1
Flash-WAM: Modality-Aware Distillation for World Action Models
Paper
• 2606.05254
• Published • 7
Latent Reasoning with Normalizing Flows
Paper
• 2606.06447
• Published • 8
Discrete-WAM: Unified Discrete Vision-Action Token Editing for World-Policy Learning
Paper
• 2606.05645
• Published • 2
Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation?
Paper
• 2606.04811
• Published • 17
Cosmos 3: Omnimodal World Models for Physical AI
Paper
• 2606.02800
• Published • 142
Qwen-Image-Flash: Beyond Objective Design
Paper
• 2606.03746
• Published • 37
Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation
Paper
• 2606.02684
• Published • 17
GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors
Paper
• 2606.05160
• Published • 9
Unlocking Feature Learning in Gated Delta Networks at Scale
Paper
• 2606.04048
• Published • 2
SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
Paper
• 2605.31148
• Published • 3
On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters
Paper
• 2606.02437
• Published • 241
PhysBrain 1.0 Technical Report
Paper
• 2605.15298
• Published • 145
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments
Paper
• 2605.30280
• Published • 146
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
Paper
• 2605.27365
• Published • 146
Paper
• 2605.03269
• Published • 128
Robots Need More than VLA and World Models
Paper
• 2606.06556
• Published • 31
World Pilot: Steering Vision-Language-Action Models with World-Action Priors
Paper
• 2606.12403
• Published • 27
Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning
Paper
• 2606.11683
• Published • 30
On Subquadratic Architectures: From Applications to Principles
Paper
• 2606.12364
• Published • 24
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models
Paper
• 2606.11324
• Published • 172
World Model Self-Distillation: Training World Models to Solve General Tasks
Paper
• 2606.12072
• Published • 15
Paper
• 2606.13392
• Published • 154
WEAVER, Better, Faster, Longer: An Effective World Model for Robotic Manipulation
Paper
• 2606.13672
• Published • 4
Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack
Paper
• 2606.14409
• Published • 15
μ_0: A Scalable 3D Interaction-Trace World Model
Paper
• 2606.13769
• Published • 11
World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible
Paper
• 2606.13652
• Published • 16
World Value Models for Robotic Manipulation
Paper
• 2606.24742
• Published • 7
EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies
Paper
• 2606.20092
• Published • 6
From Foundation to Application: Improving VLA Models in Practice
Paper
• 2607.06403
• Published • 20
GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation
Paper
• 2607.02642
• Published • 40
OpenCoF: Learning to Reason Through Video Generation
Paper
• 2607.08763
• Published • 29
Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing
Paper
• 2607.07953
• Published • 16
RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies
Paper
• 2607.04434
• Published • 15
Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation
Paper
• 2607.07608
• Published • 57
Let RGB Be the Language of Vision
Paper
• 2607.12450
• Published • 14
GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch
Paper
• 2607.13960
• Published • 28
SPEAR: A Simulator for Photorealistic Embodied AI Research
Paper
• 2607.06701
• Published • 8
Paper
• 2607.16051
• Published • 77
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Paper
• 2607.15330
• Published • 73
On-Policy Delta Distillation
Paper
• 2607.15161
• Published • 38
Understanding Reasoning from Pretraining to Post-Training
Paper
• 2607.16097
• Published • 29
See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models
Paper
• 2607.11498
• Published • 7
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
Paper
• 2607.17977
• Published • 198
Group Entropy-Controlled Policy Optimization
Paper
• 2607.16850
• Published • 29
Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning
Paper
• 2607.14183
• Published • 68
Distilled Reinforcement Learning for LLM Post-training
Paper
• 2607.17247
• Published • 9
JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models
Paper
• 2607.16074
• Published • 8
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
Paper
• 2607.19191
• Published • 309
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
Paper
• 2607.19064
• Published • 77
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report
Paper
• 2607.18367
• Published • 60
H^2SD: Hybrid Hindsight Self-Distillation
Paper
• 2607.18955
• Published • 7
Appearance Pointers -- Multimodal Region Control of Diffusion Transformers
Paper
• 2607.19344
• Published • 4
Masked Visual Actions for Unified World Modeling
Paper
• 2607.19343
• Published • 9
SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments
Paper
• 2607.20207
• Published • 4
Self-Supervised Learning of Structured Dynamics from Videos
Paper
• 2607.21576
• Published • 19
Visual Contrastive Self-Distillation
Paper
• 2607.21556
• Published • 52
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
Paper
• 2607.21553
• Published • 39
TableVerse: A Large-scale Tabletop Dataset with Real-world Grounded Layouts for Generalizable Manipulation
Paper
• 2607.21017
• Published • 7
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning
Paper
• 2607.21653
• Published • 32
Scaling Native Multimodal Pre-Training From Scratch
Paper
• 2607.22043
• Published • 28
Kimi K3: Open Frontier Intelligence
Paper
• 2607.24653
• Published • 486
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey
Paper
• 2607.21655
• Published • 193
Data Pyramid for Embodied Manipulation
Paper
• 2607.24744
• Published • 36
UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
Paper
• 2607.23373
• Published • 7
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
Paper
• 2607.27205
• Published • 139
StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation
Paper
• 2607.26754
• Published • 18
PhiZero: A World Model Built Around Physical Language
Paper
• 2607.28624
• Published • 167
Multi-Head Attention Residuals
Paper
• 2607.27230
• Published • 11
Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers
Paper
• 2607.28611
• Published • 21
Paper
• 2607.27201
• Published • 103
RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models
Paper
• 2607.26991
• Published • 8
ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow
Paper
• 2607.27924
• Published • 14
N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation
Paper
• 2607.23783
• Published • 47
N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
Paper
• 2607.23782
• Published • 76
DiffusionGemma Technical Report
Paper
• 2608.00146
• Published • 33
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
Paper
• 2608.05000
• Published • 52
Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data
Paper
• 2608.02580
• Published • 22
BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation
Paper
• 2608.05042
• Published • 7