VisTacAlign

Co-Training Dexterous Policies on Tactile Human and Robot Demonstrations

Anonymous Authors

Under review, ICRA 2027

VisTacAlign70% lift
Co-trained on robot demonstrations plus human demonstrations with aligned touch.
Robot-only baseline50% lift
The same 26 robot demonstrations, without the human data.

Drill task: grasp, pull the trigger, and lift while holding it. Adding 26 aligned human demonstrations raises success from 50% to 70%.

Human demonstrations are cheap, but a robot can only learn from them if they look, move and feel like its own. VisTacAlign aligns human demonstrations to the robot in vision, touch and kinematics, so a few robot demonstrations go much further on force-critical dexterous tasks.

10 → 80%
Lego task success, robot-only vs co-trained
50 → 70%
Drill lift success with aligned human data
0 → 70%
Plucking large strawberries
92 → 58%
Domain classifier on tactile signals, chance is 50%

Why human demonstrations?

Human demonstrationglove
A tactile glove and a tracker. No robot needed, and two to three times faster per demonstration.
Robot teleoperationrobot
The same grasp, teleoperated. The embodiment the policy runs on, but slow to collect.
But human demonstrations without aligned touch are worse than no human data at all.

How does VisTacAlign work?

VisionRp2d
The human hand is erased and the robot's own hand is rendered in its place, then stereo is re-run.
Touchsignal space
Glove forces are mapped onto the robot's force distribution, in the robot's own units.
Component 1
Kinematic retargeting

Glove hand poses mapped to the robot hand, with a learned fingertip correction calibrated once against a headset as pseudo ground truth.

Component 2
Tactile alignment

A static map carries human glove forces into the robot's own taxel units, fitted from unpaired demonstrations.

Component 3
Visual alignment (Rp2d)

The human hand is removed and a robot-textured mesh is composited back into both stereo views, then stereo is re-run.

Component 4
3D-tactile transformer

A 23.5M-parameter policy over point-cloud, proprioceptive and tactile tokens, trained with consistency flow matching.

The full pipeline and policy architecture
Pipeline: human and robot streams are aligned in vision, touch and kinematics, fed to a diffusion transformer, and deployed.
The co-training pipeline. Top: human demonstrations with a tactile glove (blue) and teleoperated robot demonstrations (green) become a shared co-training dataset. Each stream provides object-centric point clouds, per-finger tactile activations and action trajectories. Middle: a diffusion transformer encodes point-cloud, proprioception and tactile tokens as context and denoises 64-step action chunks. Bottom: the co-trained policies are deployed on the robot across all three tasks.
Fixing the glove's fingertips (13.7 mm → 7.1 mm error)

Motion-capture gloves miss pinch grasps by about a centimetre. We calibrate once against a headset's hand tracker, fitting a small network that nudges each joint by at most 20° with bone lengths fixed.

Held-out evaluation over five temporal folds and two held-out recordings.
MetricRaw gloveCorrected
Median fingertip error13.7 mm7.1 mm
Thumb–index error at pinches11.8 mm3.4 mm
Raw glove tracking leaves a gap between thumb and index during a pinch; the correction closes it.
Pinch correction. Raw glove tracking (red) leaves a gap between the thumb and index fingertips during a pinch. Our correction (green) snaps the retargeted thumb into contact so the robot hand closes the grasp.

How do we align touch?

The glove and the robot fingertips share nothing: different layout, gain and range. We reduce each to one force per finger and fit a single static map that carries the human distribution onto the robot's, in the robot's own units. Zero contact stays zero, and the policy and robot data are never transformed.

Glove and robot taxel activations during the same drill grasp, and the mapped force distributions.
Tactile alignment. During the same drill grasp the glove (blue) and the robot hand (green) produce taxel activations with different layouts, sensitivities and ranges (left). Each is reduced to a per-finger force and the human distribution is mapped onto the robot's (right). Raw human forces are broadly spread relative to the robot. Mapped forces overlap the robot in the joint thumb–index density and in the per-finger marginals.
Held-out episodes, five folds, map refitted per fold. Chance accuracy for the classifier is 50%.
MeasureBeforeAfter
KS distance to robot marginal (thumb)0.330.08
KS distance to robot marginal (index)0.420.08
Human-vs-robot force classifier92%58%
Watch the raw human and robot signals side by side
The tactile gap. The glove (top) and the robot hand (bottom) during the same grasp. Taxel layouts differ, load is shared differently between fingers, and peak forces differ by up to 3×.

How do we align vision?

The human hand is erased from both stereo views and the robot's hand mesh, painted with real robot pixels, is composited in its place. Stereo is then run again, so the inserted hand carries the same sensor noise as real robot point clouds instead of looking pasted in.

Both streams lifted to 3D; the human hand is replaced with a posed, coloured robot-hand mesh and re-stereoed.
Visual alignment. Both streams are lifted to 3D and share the same crops. For human demonstrations (blue), the segmented hand is replaced by a posed robot-hand mesh, coloured and reprojected into the image (red), so the resulting point cloud matches what the robot (green) observes at deployment.
Separability is the AUC of a roughness test against real sensed points (0.5 = indistinguishable). Chamfer and F-score compare the human hand point cloud with a robot hand point cloud at the nearest joint configuration.
MeasureBeforeOurs
Separability from sensed points (pasted mesh → composited)0.990.67–0.76
Chamfer distance (glove point cloud → composited)13.3 mm10.2 mm
F-score at 10 mm (glove point cloud → composited)0.400.59

How well does it work?

Three real-world tasks where success depends on force. 26H10R means 26 human and 10 robot demonstrations.

Drill: human data helps, but only with touch

Rollouts are at the top of the page. From 10 robot demonstrations, adding 26 touchless human demonstrations lowers success to 30%. The same demonstrations with aligned touch raise it to 55%.

Subtask success rate (%), 20 rollouts per policy. Lift is overall task success. Obs. is the visual observation, point cloud (PC) or RGB. Human tact. indicates whether the human demonstrations also carry aligned tactile.
DemosObs.Human tact. ApproachActivateLift
26RRGB–1003535
26RPC–906550
26H26RPCyes807070
10RPC–856535
26H10RPCno956530
26H10RPCyes1005555

Strawberry: touch regulates force

Pluck a strawberry held by a magnet: hard enough to beat the magnet, gentle enough not to crush it. Without touch the policy fails on large fruit entirely.

Co-train + tactileOurs
Grips hard enough to overcome the magnet without crushing the fruit.
Robot only, no touchBaseline
Robot only, no tactile input. It under-grips small fruit and crushes larger ones.
Plucking success (%) per fruit size, ten rollouts per cell.
Fruit sizeRobot onlyCo-train + tactile
Small (20–25 mm)6090
Medium (30–35 mm)7080
Large (40–45 mm)070
Grip force by fruit size
Mean peak grasp force by fruit size: the robot-only policy under-grips small fruit and over-grips large; the co-trained policy stays near the demonstration force.
Mean peak grasp force (thumb + index) in taxel units, for three fruit sizes, ten trials per size for each policy. The dashed line is the force at which the demonstrations hold the fruit. The robot-only policy under-grips small fruit and over-grips medium and large, while the co-trained tactile policy stays in a narrow band across all sizes.

Lego: the recipe scales, and needs visual alignment

Pick, align, push until it snaps, retreat. Human data takes success from 10% to 80%, and without the visual alignment the push and retreat stages collapse.

With visual alignmentOurs
Human point clouds rendered with the robot hand (Rp2d). Co-training takes Lego from 10% to 80% success.
No Rp2dAblation
Raw human point clouds. The policy still picks up the brick, but success collapses at the push and retreat.
Subtask success rate (%) across 20 trials. Top block: a robot-only baseline and the full method with a growing number of robot demonstrations. Bottom block: ablations removing the visual alignment or the tactile input.
DemosTactileRp2d PickupAlignPushRetreat
30Ryes–70101010
50H10Ryesyes90504040
50H20Ryesyes60606060
50H30Ryesyes100808080
50H10Ryesno100501010
50H20Ryesno1001003025
50H30Rnoyes70707070

Hardware

Benchtop setup: tactile glove and wrist tracker for humans, tactile robot hand on a 7-DoF arm, stereo camera observing the workspace.
Setup. Human demonstrations are collected with a motion-capture glove fitted with our capacitive tactile glove. Robot demonstrations use a 17-DoF tactile hand on a 7-DoF arm. Trackers on the glove, the camera mount and the table provide a shared frame, and a stereo camera observes the workspace.

Q&A

Why align in signal space instead of a learned latent space?

Three reasons. The aligned quantity is still a per-finger force in the robot's units, so it can be checked directly against the robot distribution with standard statistics, which is exactly what the tables above do. The map is fitted once and deterministically from a few thousand unpaired frames, with one hyperparameter and no encoder to train. And because it is static and acts frame by frame, it cannot shift the timing of contact events.

We do not compare head-to-head against a learned latent alignment. Our case rests on interpretability and on guarantees that hold by construction, and quantifying that trade-off is the most immediate item of future work.

Why re-run stereo instead of just pasting the robot hand into the point cloud?

Because a pasted mesh is not a measurement. Its points are noise-free, uniformly dense, and cover surfaces a single-view sensor could never see. A simple roughness test separates them from real sensed points almost perfectly, at 0.99 area under the curve, which means a policy can learn that shortcut too. Running stereo on the composite drops that to 0.67–0.76.

Do you still need robot demonstrations?

Yes. The tactile map is fitted per task from unpaired force samples, so a short robot recording of the task's contacts is still needed. The result is not that human data replaces robot data, but that it substantially raises what a small number of robot demonstrations is worth.

What are the main limitations?

The tactile map is fitted per task. Reducing a fingertip to one normal force discards the taxel layout and shear, which likely matters for in-hand manipulation and slip. The tasks are short-horizon and largely quasi-static. And we do not compare against alignment in a learned latent space.

How fast does the policy run?

The policy has 23.5M parameters. At deployment, stereo runs at 20 Hz, the glove at 60 Hz and the robot taxels at 30 Hz. We sample 10 flow steps and execute the first 32 of a 64-step chunk, about 1.1 seconds, before replanning, which takes 269 ms on a consumer GPU.

Does the aligned human data change how hard the robot grips?

That is the clearest single effect. On the strawberry task the robot-only policy under-grips small fruit and over-grips large fruit, and never plucks the largest size. The co-trained policy holds a narrow force band across all three sizes and succeeds on every one. Grasp force is precisely the quantity the tactile alignment transfers from the human hand.