Multimodal imitation learning requires diverse executable futures under the same observation and consistent behavior across replanning cycles. We present Conditional Trajectory Peaks (CTP), a single-pass policy framework that jointly predicts complete action-chunk candidates, probability masses, and trajectory scales. Distribution-Aware Peak Specialization (DAPS) specializes trajectory peaks using trajectory-level posterior responsibilities and mass- and scale-modulated overlap constraints. Evidence-Gated Trajectory Belief Transport (ETBT) maintains cross-chunk consistency through geometric correspondence between exchangeable candidate sets, while allowing current policy evidence to override historical constraints. CTP achieves a coverage score of 91.40% on Push-T; success rates of 100.0%, 79.72%, and 84.44% on D3IL Avoiding, Aligning, and Sorting-2, respectively. On LIBERO, CTP achieves an average success rate of 97.25%. In real-world dual-arm experiments, CTP preserves both placement modes in a two-plate task, succeeding in all 50 trials. On bottle uprighting and pen placement into a holder, it maintains success rates comparable to π0.5 while reducing policy inference latency from 218.24 ms to 75.80 ms. These results demonstrate that single-pass trajectory modeling can combine multimodal behavior, closed-loop consistency, and efficient inference.
Motivation
Four linked challenges in multimodal closed-loop control
Figure 1. Core challenges in multimodal closed-loop manipulation. (a) The same observation admits multiple feasible futures, while the conditional mean may be infeasible. (b) Diffusion and flow-matching policies generate action chunks iteratively, increasing online decision latency. (c) Naive single-pass multi-candidate prediction can still suffer from training-time mode collapse and inconsistent execution caused by mode switching during receding-horizon replanning.
Method
Trajectory-level multimodality without iterative generation
Given an observation history, CTP models the conditional distribution of an entire H-step action chunk rather than predicting each action independently. A shared policy backbone produces K complete trajectory centers, their probability masses, and bounded residual scales in a single forward pass. One discrete component therefore explains the full planning horizon, giving every candidate a trajectory-level behavioral identity while retaining NFE = 1.
CTP separates the two challenges of multimodal closed-loop control. During training, DAPS prevents exchangeable candidates from collapsing onto the same dominant solution. During receding-horizon execution, ETBT associates candidates geometrically across consecutive replanning cycles and preserves a behavior only while it remains compatible with current observations.
Method overview. DAPS specializes exchangeable trajectory peaks during training; ETBT transports behavioral belief using trajectory correspondence at inference. Candidate generation remains NFE = 1.
Representation
Conditional trajectory peaks
CTP defines a finite conditional mixture over complete action chunks. Each peak predicts a trajectory center μ, probability mass π, and residual scale σ. The representation supports either direct time-domain chunks or a compact trajectory basis, without predefined behavioral anchors.
Training
DAPS
Trajectory-level posterior responsibilities assign demonstrations using fit, probability mass, and uncertainty together. A distribution-aware overlap constraint then separates high-probability peaks that redundantly explain the same data region, while unsupported candidates may reduce their mass instead of being forced into artificial modes.
Inference
ETBT
After executing part of a chunk, ETBT matches every previous unexecuted suffix to current candidate prefixes. It transports behavioral belief through these permutation-equivariant correspondences, fuses history as a soft prior with current policy evidence, and releases that prior when correspondence becomes unreliable.
Quantitative results
Strong closed-loop performance with single-pass inference
Matched Push-T training and evaluation protocol; higher is better.
Qualitative results
Experiments by benchmark
Videos are grouped by benchmark, following the organization used in the paper.
D3IL
Multimodal low-dimensional control
100%Avoiding
79.72%Aligning
84.44%Sorting-2
Avoiding, Aligning and Sorting-2 are the three tasks reported in the main quantitative table; Sorting-4 is included as an additional D3IL rollout. Every video is generated by its corresponding evaluated checkpoint on a formally successful context.
CTP is integrated into π0.5, predicts four 50-step candidates, executes 10 actions and replans with trajectory belief decoding. Representative successful rollouts are shown for all four LIBERO suites.
On an RTX 4090 at batch size 1, CTP reaches 97.25% average success at 75.80 ms per policy call, versus 96.75% and 218.24 ms for the matched π0.5 baseline.
Real-robot experiments
Behavioral modes that survive deployment
A dual-arm study evaluates both multimodal preservation and policy-call efficiency.
Two-plate placement. CTP selects the left plate 31 times and the right plate 19 times across 50 trials, with all 50 placements successful.Success–latency comparison. CTP matches π0.5 on bottle uprighting and remains comparable on pen placement while reducing policy latency from 218.24 ms to 75.80 ms.
Multimodal two-plate placement
All eight displayed rollouts begin from the same object placement. The opposite successful outcomes therefore reflect conditional policy multimodality rather than variation in the initial state; the side counters summarize the full 50-trial evaluation.
8 rolloutsidentical start pose
Eight representative rollouts from an identical initial object placement; the side counters report the complete 50-trial evaluation.
Representative real-robot rollouts
Successful executions on bottle uprighting and pen placement into a holder, using the deployed CTP policy.