REAL-TIME VLA · PAPER COMPANION · 2026
From inference
to execution.
Model inference optimization and system-level deployment evaluation for vision-language-action policies on physical robots.
THE RESEARCH QUESTION
What happens after
the model responds?
VLA policies infer action chunks at a lower rate than a robot executes them. The resulting timing gap creates stale observations, discontinuous handovers, and hidden system latency. This report studies the complete path from model computation to physical control, under one reproducible deployment protocol.
Inference speed and task success come from separate evaluations; combining the faster sampler with an execution strategy changes task success.
RESULTS
Measured on real
bimanual manipulation.
Six execution strategies are compared on a long-horizon T-shirt folding task using two Agilex Piper arms. Every method uses the same base policy, hardware, task definition, and success criteria.
| Method | Success | Mean time | Throughput |
|---|---|---|---|
| Legato |
96.7%
|
73.56 s | 47.31 h−1 |
| VLASH |
93.3%
|
77.73 s | 43.22 h−1 |
| Temporal Smoothing |
76.7%
|
93.93 s | 29.38 h−1 |
| Training-time RTC |
63.3%
|
129.37 s | 17.62 h−1 |
| Naive Asynchronous |
63.3%
|
124.13 s | 18.37 h−1 |
| Inference-time RTC |
60.0%
|
133.73 s | 16.15 h−1 |
Scroll horizontally to view the full table →
RESEARCH CONTEXT
From model latency
to system timing.
Real-time VLA research spans asynchronous inference, inter-chunk continuity, adaptive horizons, fast action generation, continuous-time representations, and system-level deployment. This report focuses on the final transition: making those ideas measurable on a physical robot.
SYSTEM
A distributed runtime
built for traceability.
Policy inference, action publication, and robot control run at independent rates. Each action chunk carries its timing and provenance through the loop.
LATENCY CALIBRATION
Measure the loop,
not just the model.
A visual time code, an end-effector ArUco marker, joint commands, image timestamps, and proprioceptive feedback are recorded together. Their phase relationships estimate exposure, readout, motion-response, and feedback delays.

EXECUTION STRATEGIES
Eight ways to
close the timing gap.
Eight execution approaches are illustrated below. Six are compared in the physical-robot evaluation under a common deployment interface. Select an image to view it at full size.








INFERENCE EFFICIENCY
Two stages.
Less waiting.
Flow Matching integration is not equally informative at every step. The proposed schedule keeps a large early move and a short terminal refinement. It reduces model-side inference time, with a task-success tradeoff when combined with the tested execution strategies.


PHYSICAL SETUP
Designed for
long-horizon tasks.

We evaluate continuous execution on a long-horizon T-shirt folding task. The same protocol is repeated across three garment conditions.
Tail Gap, Switch Gap, velocity, acceleration, and tracking error expose behavior at action-chunk boundaries.









TRAJECTORY DIAGNOSTICS
Continuity shows up
at the handover.
Local position, velocity, and acceleration traces expose the execution behavior that task success alone cannot show. Shaded regions mark action-chunk transitions.






OPEN MATERIALS
Read it. Run it.
Build on it.
@misc{wu2026realtimevlasstageawaretwostep,
title={Toward Real-Time VLAs: Stage-Aware Two-Step Flow Denoising and System-Level Evaluation},
author={Di Wu and Rongtian Shen and Ping Liu and Yan Shen and Zhenhan Yin and Shun Zuo and Xuhua Chen and He Zheng and Lingfeng Zhang and Jianglin Zhang and Tao Zhang},
year={2026},
eprint={2609.39822},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2609.39822},
}