Researchers from the Intelligent Space Robotics lab (Skoltech) presented HapticVLA — a VLA model that relies on tactile sensors during training but operates without them at inference, acting only on vision and robot state. In real-world experiments on tasks with fragile objects, the model achieved an average success rate of 86.7%, outperforming baseline VLAs, including models that received tactile data directly at inference. Code, weights, and datasets are open; the work was accepted to IROS 2026.



What happened
The first version of the HapticVLA paper appeared on arXiv on 16.03.2026 (arXiv:2603.15257), with an updated revision dated 02.08.2026; the work was accepted to IROS 2026. The method consists of three sequential stages: first, tactile rewards are computed offline with penalties for excessive grasp force and object slipping; then the base SmolVLA model is fine-tuned using Safety-Aware Reward-Weighted Flow Matching (SA-RWFM); finally, Tactile Distillation distills a compact tactile token into a VLA student that predicts actions only from vision and manipulator state. Validation was not done in simulation: the model performed three real-world tasks with objects of varying fragility — a plastic can, waffles, and eggs — and outperformed baseline VLAs in average success rate, including those that received tactile data directly at inference.
Context
VLA models (Vision-Language-Action) translate an image and command directly into robot actions, and the classic solution for gentle operations is tactile sensors operating at inference. HapticVLA inverts this scheme: the tactile modality is moved from inference to the training stage, where it serves as a teacher and reward signal, and its meaning is then distilled into the model weights. The novelty of the method lies in the order of data usage, not in a new architecture: the pattern of "expensive sensor during training — no sensor in production" is non-trivial for VLA models. It is also significant that the setup is built from open components: the Crab platform with two SO-101 manipulators within the LeRobot ecosystem, which significantly reduces the cost of independent verification.
Why this matters for industry
For industry, the key point is that "sense of touch" ceases to be a property of hardware and becomes a property of the model. Tactile sensors are expensive and do not transfer well between platforms; if gentle behavior is distilled into the weights of a VLA model (SA-RWFM plus Tactile Distillation), tactile-aware operations can be deployed on cheap standard manipulators like the SO-101 without tactile hardware on board. For Physical AI teams, this reduces pilot CAPEX and improves reproducibility of contact manipulation across different robots, while value shifts from sensors to training data and distillation methods. Open code, weights, and datasets further reduce the cost of verification and refutation by other teams — a rare property for contact robotics.
Why this matters for users
Readers already have access to a full set for independent work: an open LeRobot fork with the Crab platform, model weights, datasets, and a project page with a digital twin of the setup. The approach can be studied and replicated on SO-101 without purchasing tactile sensors — build the setup according to the digital twin and run the SA-RWFM and Tactile Distillation pipeline on your own objects, including fragile ones. A separate signal for students: the work was done in six months in the AI Robotics master's program at Yandex and Skoltech — a clear example of the level of results achievable within an academic program.
What is still unknown / limitations
The phrasing "first VLA model" is the authors' positioning from the abstract, and the novelty should be checked against prior work. Results were obtained on a desktop two-manipulator setup and have not yet been re-verified by third parties; direct transfer to a production environment, such as packaging fragile goods, without adaptation is not possible. Finally, this is a research pipeline, not a finished product: there is no API or service, and the 86.7% metric needs to be confirmed by independent replications on other platforms and objects.
Sources
- HapticVLA: Contact-Rich Manipulation via Vision-Language-Action Model without Inference-Time Tactile Sensing (arXiv:2603.15257 paper)
- HapticVLA project page (code, weights, datasets, and digital twin of the setup)
- HapticVLA GitHub repository (Advanced-Robotic-Manipulation/crab, work accepted to IROS 2026)
Author
Look at AI, editorial team
