ICRA 2027
Dual-Stream Memory for Vision-Language-Action Policies
Recurrent execution context in one token. Retrieved spatial evidence above a relevance threshold. 49.22% mean TSR on RoboMME, 62.5% on four Piper tasks.
Abstract
Memory-dependent manipulation requires a policy to track execution state and recover visual evidence that is no longer observable. We propose dual-stream memory for vision-language-action policies, assigning these requirements to complementary representations. A task-agnostic Gated DeltaNet compresses history into one temporal token, supervised by progress prediction and action-conditioned feature reconstruction. A task-conditioned retriever retains frames above a fixed relevance threshold; linear multi-scale fusion encodes each frame into sixteen spatial tokens. We then freeze the temporal compressor and jointly optimize retrieval and policy learning. Across sixteen RoboMME tasks, the complete model achieves 49.22% mean success versus 44.51% for the published FrameSamp+Modul reference, improving fourteen tasks and all four suite means. On four real-robot tasks, mean success reaches 62.5% versus 37.5%, with the largest gains on repeated actions requiring execution-stage discrimination. These results support combining execution context with selective visual access, while failures under cup swaps expose persistent identity tracking as an unresolved requirement.
Two streams, one policy
A task-agnostic Gated DeltaNet compresses history into one temporal token. Frozen SigLIP retrieval keeps frames above LSE logit 0.6 and writes sixteen spatial tokens per frame. Stage 2 freezes GDN and trains retrieval with the π0.5 policy.
Results
The complete model is 49.22% mean TSR on sixteen RoboMME tasks versus 44.51% for FrameSamp+Modul. Reference gained most (+8.98). Combined gains are not additive.
Suite tables Component additions, RoboMME suite means, and Piper TSR
| Component / suite | Base | Addition | Δ (pp) |
|---|---|---|---|
| GDN / Counting | 65.22 | 72.17±4.25 | +6.94 |
| Retrieval / Permanence | 25.11 | 29.33±0.76 | +4.22 |
| Linear fusion / Reference | 36.33 | 40.83±3.62 | +4.50 |
Suite means when each component is added to FrameSamp+Modul. TSR (%) averages three evaluation seeds. Δ is computed before rounding. FrameSamp+Modul values are the published RoboMME reference.
| Suite | FrameSamp+Modul | Ours | Δ (pp) |
|---|---|---|---|
| Counting | 65.22 | 70.83 | +5.61 |
| Permanence | 25.11 | 27.75 | +2.64 |
| Reference | 36.33 | 45.31 | +8.98 |
| Imitation | 51.39 | 53.00 | +1.61 |
| Overall | 44.51 | 49.22 | +4.71 |
Complete-model RoboMME suite means versus the published FrameSamp+Modul reference. Means weight tasks equally. Δ is computed from the reported task means.
| Method | Pick×3 | Swing×2 | RePick | Shell | Avg. |
|---|---|---|---|---|---|
| π0.5 base | 20.0 | 6.7 | 3.3 | 30.0 | 15.0 |
| NativeMEM | 20.0 | 30.0 | 40.0 | 36.7 | 31.7 |
| FrameSamp+Modul | 26.7 | 36.7 | 33.3 | 53.3 | 37.5 |
| Ours | 80.0 | 93.3 | 36.7 | 40.0 | 62.5 |
Piper TSR (%): 30 rollouts per task and method. Avg. weights the four tasks equally. The largest gains are on Pick×3 and Swing×2. Shell declined from 53.3% to 40.0%.
Training diagnostics Progress head, sixteen-slot history reconstruction, and temporal retrieval
Piper rollouts
Four tasks and four methods, sixteen clips. Overhead frames illustrate the settings, not evaluation outcomes.
Overhead frames show the four task settings, not policy-evaluation outcomes.
BibTeX
@inproceedings{anonymous2027dualstream,
title={Dual-Stream Memory for Vision-Language-Action Policies},
year={2027},
note={Under double-anonymous review}
}