ICRA 2027 Double-anonymous review

ICRA 2027

Dual-Stream Memory for Vision-Language-Action Policies

Anonymous Authors

Recurrent execution context in one token. Retrieved spatial evidence above a relevance threshold. 49.22% mean TSR on RoboMME, 62.5% on four Piper tasks.

Teaser: dual-stream memory with RoboMME and Piper mean task success bars.
Dual-stream memory combines recurrent execution context with retrieved spatial evidence. Bars show mean TSR across sixteen RoboMME tasks and four Piper tasks. Orange boxes mark the visible target and its covering cup; memory symbols are schematic.
PAPER

Abstract

Memory-dependent manipulation requires a policy to track execution state and recover visual evidence that is no longer observable. We propose dual-stream memory for vision-language-action policies, assigning these requirements to complementary representations. A task-agnostic Gated DeltaNet compresses history into one temporal token, supervised by progress prediction and action-conditioned feature reconstruction. A task-conditioned retriever retains frames above a fixed relevance threshold; linear multi-scale fusion encodes each frame into sixteen spatial tokens. We then freeze the temporal compressor and jointly optimize retrieval and policy learning. Across sixteen RoboMME tasks, the complete model achieves 49.22% mean success versus 44.51% for the published FrameSamp+Modul reference, improving fourteen tasks and all four suite means. On four real-robot tasks, mean success reaches 62.5% versus 37.5%, with the largest gains on repeated actions requiring execution-stage discrimination. These results support combining execution context with selective visual access, while failures under cup swaps expose persistent identity tracking as an unresolved requirement.

ARCHITECTURE

Two streams, one policy

A task-agnostic Gated DeltaNet compresses history into one temporal token. Frozen SigLIP retrieval keeps frames above LSE logit 0.6 and writes sixteen spatial tokens per frame. Stage 2 freezes GDN and trains retrieval with the π0.5 policy.

Dual-stream architecture: retrieval, GDN compression, and staged training.
(a) GDN compresses history into one temporal token; threshold retrieval and multi-scale fusion produce 16mt spatial tokens. (b) Stage 2 trains retrieval and the policy with GDN frozen. (c) Stage 1 pretrains GDN; auxiliary heads are then removed.
EVIDENCE

Results

The complete model is 49.22% mean TSR on sixteen RoboMME tasks versus 44.51% for FrameSamp+Modul. Reference gained most (+8.98). Combined gains are not additive.

Three-panel results: component additions, RoboMME suite means, and Piper task success.
(a) Adding one component to FrameSamp+Modul. (b) Complete-model suite means versus the published FrameSamp+Modul reference. (c) Piper TSR, 30 rollouts per cell. Ours is terracotta.
Suite tables Component additions, RoboMME suite means, and Piper TSR
Component additions to FrameSamp+Modul
Component / suite Base Addition Δ (pp)
GDN / Counting 65.22 72.17±4.25 +6.94
Retrieval / Permanence 25.11 29.33±0.76 +4.22
Linear fusion / Reference 36.33 40.83±3.62 +4.50

Suite means when each component is added to FrameSamp+Modul. TSR (%) averages three evaluation seeds. Δ is computed before rounding. FrameSamp+Modul values are the published RoboMME reference.

Complete model RoboMME suite means
Suite FrameSamp+Modul Ours Δ (pp)
Counting 65.22 70.83 +5.61
Permanence 25.11 27.75 +2.64
Reference 36.33 45.31 +8.98
Imitation 51.39 53.00 +1.61
Overall 44.51 49.22 +4.71

Complete-model RoboMME suite means versus the published FrameSamp+Modul reference. Means weight tasks equally. Δ is computed from the reported task means.

Piper task success rates
Method Pick×3 Swing×2 RePick Shell Avg.
π0.5 base 20.0 6.7 3.3 30.0 15.0
NativeMEM 20.0 30.0 40.0 36.7 31.7
FrameSamp+Modul 26.7 36.7 33.3 53.3 37.5
Ours 80.0 93.3 36.7 40.0 62.5

Piper TSR (%): 30 rollouts per task and method. Avg. weights the four tasks equally. The largest gains are on Pick×3 and Swing×2. Shell declined from 53.3% to 40.0%.

Training diagnostics Progress head, sixteen-slot history reconstruction, and temporal retrieval
Current-progress predictions versus ground truth for eight tasks, two episodes each.
Current-progress predictions for the first eight tasks, two episodes per task. Progress-head MSE is 0.03430 over sixteen tasks.
Sixteen-slot visual history comparisons with green match boxes and red mismatch boxes.
Sixteen-slot visual-history comparisons for InsertPeg, PatternLock, and SwingXtimes. Green boxes mark scene matches and red boxes mark mismatches. These examples illustrate action-conditioned reconstruction, not a dataset-level retrieval metric.
Temporal retrieval on PickHighlight, VideoUnmask, ButtonUnmask, and StopCube.
Temporal retrieval on four training episodes: PickHighlight, VideoUnmask, ButtonUnmask, and StopCube. Gold lines mark keyframe centers. Inference thresholds LSE logits at 0.6, not the dashed visualization line.
PHYSICAL

Piper rollouts

Four tasks and four methods, sixteen clips. Overhead frames illustrate the settings, not evaluation outcomes.

Pick×3 setting: apple and yellow tray.
(a) Pick×3
Swing×2 setting: apple between two trays.
(b) Swing×2
RePick setting: four objects in a row.
(c) RePick
Shell setting: cups on the table.
(d) Shell

Overhead frames show the four task settings, not policy-evaluation outcomes.

π0.5
NativeMEM
FrameSamp+Modul
Ours
Pick×3
Fail
Fail
Success
Success
Swing×2
Fail
Fail
Fail
Success
RePick
Fail
Fail
Fail
Success
Shell
Fail
Fail
Fail
Success
CITE

BibTeX

@inproceedings{anonymous2027dualstream,
  title={Dual-Stream Memory for Vision-Language-Action Policies},
  year={2027},
  note={Under double-anonymous review}
}