DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination

CoRL 2026

Yankai Fu1,2*, Ning Chen1,2*, Junkai Zhao2†, Heng Zhang1,
Guocai Yao2, Pengwei Wang2, Zhongyuan Wang2, Shanghang Zhang1,2✉
1State Key Laboratory of Multimedia Information Processing, School of Computer Science,
Peking University; 2Beijing Academy of Artificial Intelligence
*Equal contribution, Project leader, Corresponding author
Project video

Abstract

Overview of DeCAL, its multimodal inputs, tasks, and results

Dexterous manipulation involves contact-rich and fine-grained interactions with the physical world, posing significant challenges for existing vision-language-action models because of severe visual occlusions and complex contact dynamics. While recent work has incorporated tactile sensing into robotic manipulation, most approaches still rely on homogeneous multimodal fusion, without adaptive tactile integration or explicit modeling of physical dynamics.

We present DeCAL, a physically-grounded dexterous vision-language-action model that unifies understanding, imagination, and action generation. Built on a Mixture-of-Transformers architecture, DeCAL introduces Adaptive Visuo-Tactile Fusion to regulate tactile interaction through contact-aware gating, and Visuo-Tactile Latent Co-Imagination to jointly model future visual and tactile dynamics. Across six real-world tasks, DeCAL reaches a 71% average success rate and an 83.4% progress success rate, while maintaining strong out-of-distribution generalization.

71%Average success rate
83.4%Progress success rate
0.27sPer action chunk

Method

Three collaborative experts share knowledge through directional joint attention, turning touch and vision into physically-grounded action.

DeCAL architecture with understanding, generation, and action experts
Architecture overview: scene understanding, visuo-tactile dynamics foresight, and factorized action generation.

DeCAL is built upon a MoT architecture that unifies scene understanding, visuo-tactile dynamics foresight, and action generation. The Action Expert employs Factorized Flow Matching to decouple arm and hand motion, enabling better coordination and dexterous manipulation.

Collaborative MoT

An Understanding Expert, a Generation Expert, and an Action Expert exchange information through asymmetric cross-modal joint attention. Later experts can attend to all preceding knowledge while keeping each capability specialized.

Adaptive Visuo-Tactile Fusion

Per-finger tactile deformation maps become local and global tokens. A contact-aware gate raises the contribution of tactile evidence during physical interaction and suppresses it when vision should dominate.

Latent Co-Imagination

The world-model expert predicts future visual and tactile latents, including force, deformation, and raw tactile signals. This forward-looking supervision embeds implicit physical dynamics into the policy.

Factorized Flow Matching

Arm and hand actions start from independent noise distributions, then are jointly processed for coordinated motion. Decoupling their dynamics preserves fine-grained hand control without losing bimanual coordination.

Built for contact-rich dexterity

The real-world platform combines two 6-DoF UR5 arms, two 22-DoF SharpaWave hands, three Intel RealSense D435 cameras, and high-resolution vision-based tactile sensing on every fingertip.

  • RGB Head and wrist camera views
  • Touch Raw tactile images and deformation maps
  • Force Per-finger 6-DoF force and torque
Dual-arm dexterous robot system and teleoperation hardware

Real-World Demos and Results

Six contact-rich tasks, 100 demonstrations per task, and 20 real-world evaluation trials for every method.

Six real-world tasks and tactile representation analysis

We analyze the global tokens extracted by the tactile encoder using t-SNE. For each task, we sample frames from distinct contact phases, covering diverse interaction patterns such as twisting, insertion, and wiping. As shown in the figure above, the learned global tactile representations form clear clusters across different contact modes, indicating that the encoder captures meaningful and structured physical interaction patterns.

Wipe Vase

100% SR

Erase Whiteboard

80% SR

Assemble Parts

65% SR

Twist Cap

80% SR

Pipetting

60% SR

Screw Light Bulb

40% SR
GR00T N1.632
InternVLA-A133
ViTacFormer48
DECO56
DeCAL71
+15.0Points over the strongest baseline
20Evaluation trials per task
6 × 100Demonstrations across six tasks
Success rate (%) across six real-world tasks
MethodWipeEraseAssembleTwistPipettingScrewAvg.
GR00T N1.630503010601032
InternVLA-A16550155353033
ViTacFormer80451565552548
DECO90603570453556
DeCAL (Ours)100806580604071

Generalization

DeCAL remains robust when the scene, illumination, or manipulated object moves beyond the training distribution.

Unseen background

60%

Cluttered scene

70%

Unseen lighting

70%

Unseen object

75%
Generalization success rates compared with ViTacFormer and DECO

Touch anchors the policy when appearance changes.

DeCAL leads all evaluated baselines across background, clutter, lighting, and object shifts. Its largest margin appears on the unseen-object setting, where geometric and force cues become especially important.

  • Background60%
  • Cluttered70%
  • Lighting70%
  • Object75%

Adaptive Tactile Gating

The contact-aware gate regulates when tactile evidence should influence the policy during dexterous interaction.

Tactile gate values over testing episodes for Assemble Parts and Twist Cap

Contact activates tactile feedback.

Across Assemble Parts and Twist Cap testing episodes, the gate remains low during non-contact motion, allowing visual observations to dominate. Once meaningful physical contact occurs, the gate value rises sharply and increases the contribution of tactile information to action generation.

With all other components enabled, removing tactile gating reduces success on both evaluated tasks.

Ablation Results (Success rate)
TaskWithout gateWith gateGain
Assemble Parts50%65%+15%
Twist Cap70%80%+10%

Citation

BibTeX
@article{fu2026decal,
              title={DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination},
              author={Yankai Fu, Ning Chen, Junkai Zhao, Heng Zhang, Guocai Yao, Pengwei Wang, Zhongyuan Wang, Shanghang Zhang},
              journal={arXiv preprint arXiv:2609.09119},
              year={2026}
            }