Token Warping Helps MLLMs Look from Nearby Viewpoints
Abstract
Token-level warping in vision-language models demonstrates superior stability and semantic coherence for viewpoint transformation compared to pixel-wise methods, achieving better visual reasoning performance.
Can warping tokens, rather than pixels, help multimodal large language models (MLLMs) understand how a scene appears from a nearby viewpoint? While MLLMs perform well on visual reasoning, they remain fragile to viewpoint changes, as pixel-wise warping is highly sensitive to small depth errors and often introduces geometric distortions. Drawing on theories of mental imagery that posit part-level structural representations as the basis for human perspective transformation, we examine whether image tokens in ViT-based MLLMs serve as an effective substrate for viewpoint changes. We compare forward and backward warping, finding that backward token warping, which defines a dense grid on the target view and retrieves a corresponding source-view token for each grid point, achieves greater stability and better preserves semantic coherence under viewpoint shifts. Experiments on our proposed ViewBench benchmark demonstrate that token-level warping enables MLLMs to reason reliably from nearby viewpoints, consistently outperforming all baselines including pixel-wise warping approaches, spatially fine-tuned MLLMs, and a generative warping method.
Community
CVPR 2026
Paper: https://arxiv.org/abs/2604.02870
Project Page: https://token-warping-mllm.github.io/
Code: https://github.com/KAIST-Visual-AI-Group/Token-Warping-MLLM
Interesting
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models (2026)
- DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding (2026)
- SeGPruner: Semantic-Geometric Visual Token Pruner for 3D Question Answering (2026)
- Contrastive Language-Colored Pointmap Pretraining for Unified 3D Scene Understanding (2026)
- Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding (2026)
- Cog3DMap: Multi-View Vision-Language Reasoning with 3D Cognitive Maps (2026)
- Spa3R: Predictive Spatial Field Modeling for 3D Visual Reasoning (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2604.02870 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper