EmbodiedModel中文 ↗

In-Context Visual Guidance with Magic-Brain

Language-following remains a persistent challenge for Vision-Language-Action (VLA) models. Our analysis of failure cases suggests that many failures do not originate from the motor actions themselves, but from failures to correctly follow language instructions. In contrast, we observed that VLA models respond more reliably to visual guidance. Our solution in Magic-Brain is therefore to use the VLA model as a manipulation policy and introduce an additional visual-guidance layer upstream. A high-level VLM Agent first localizes the target with a bounding box (bbox); a segmentation model then refines the box into a three-color mask, which is overlaid onto the observation image. In this way, semantic guidance is converted into visual guidance. The mask colors also encode temporal information, providing an additional form of short-term memory augmentation. On MolmoSpaces-Bench, this modification increased the grasping success rate of magicVLA from 70% to 94.4%.

This article explains the motivation behind this design, how it is implemented, and where it fits within the Magic-Brain framework.

1. Visual Guidance: Showing the VLA What to Grasp

1.1. Why Language Instructions Are Not Enough

A language-only instruction requires the VLA to solve two tasks that are comparatively difficult for it: first, understand the word "apple"; second, identify which pixels in the image correspond to that apple. VLA models have become highly capable at learning actions, but grounding noun phrases to image regions remains one of the most weakly supervised parts of the training signal. Our failure analysis supports this observation: in most failed grasping cases, the problem was not unstable grasp execution, but selecting the wrong object or placing it at the wrong location.

The In-Context VLA study provides more systematic evidence: allowing a VLA to freely generate textual chain-of-thought (CoT) reasoning can instead degrade low-level control. The reasons are concrete: the generated rationale may not align with objects in the visual scene due to insufficient grounding; text-generation latency interrupts the control loop; and when training objectives conflict, the policy may learn to "describe the action" rather than "execute the action well." Their conclusion is that what a VLA lacks is not the ability to "speak," but direct access to already grounded information such as "where the target is." This is consistent with our design intuition. Rather than requiring the VLA to learn object localization from text by itself, we separate responsibilities: target localization is assigned to a VLM that is strong at visual understanding, while action execution is assigned to the VLA. Our additional step is to render "where the target is" directly into the observation image, so that the VLA obtains the answer visually without having to parse the corresponding text.

1.2. Agent-Generated Bounding Boxes and CV-Generated Masks

The pipeline consists of two stages. First, the high-level VLM Agent reads the current observation image and task description, and uses in-context learning to output the target object's bbox. Several few-shot examples are included in the prompt, with no model fine-tuning required. Second, the bbox is passed to a downstream SAM-like segmentation model, which refines the coarse rectangle into a pixel-level mask.

Magic-Brain research figure

Visual Guidance Pipeline

Figure 1. Visual guidance pipeline: environment observation → Agent ICL outputs bbox → CV segmentation model outputs mask → three-color composition → guided frame is fed into magicVLA.

Why use this design? There are two main reasons. First, the idea of overlaying visual markers on the image and using them as the VLA input interface has already been validated. VP-VLA allows a System 2 VLM planner to overlay target locations and placement regions on the observation image using crosshairs and bounding boxes, while the System 1 controller consumes these visual prompts, producing improvements of 5% and 8.3% on Robocasa-GR1-Tabletop and SimplerEnv, respectively.¹ Second, the effectiveness of in-context learning in VLA settings has also been demonstrated: Retrieval-VLA combines retrieval augmentation with in-context learning, enabling a pretrained VLA to adapt to new tasks without fine-tuning.² From an engineering perspective, the bbox also serves as a clean contract-based interface: the upstream VLM can be replaced with a stronger model, and the downstream segmentation model can be upgraded independently, without either component needing to know the implementation details of the other. VP-VLA similarly adopts a "VLM-generated box + segmentation refinement" design.

1.3. Semantic Design of the Three-Color Mask

The segmentation mask is not passed to the VLA in its raw form. Instead, it is color-coded according to semantics using three colors and overlaid on the observation image.

Magic-Brain research figure

In our tabletop experiments, highly saturated colors with low visual ambiguity relative to the scene produced more stable visual cues. We therefore selected green, yellow, and silver-white. Each color carries exactly one semantic meaning, so the VLA does not need to infer combinations of colors.

For the three-color semantics to be effective, the VLA must actually learn to interpret these colors. Rather than expecting the pretrained model to infer their meanings automatically, we perform mask-based post-training on magicVLA. The model is further trained on demonstrations containing three-color guided frames, thereby encoding the semantics "green means grasp the target, silver-white means place here, and yellow represents historical trajectories" into the policy parameters. This follows the same general principle as VP-VLA, which introduces an auxiliary visual-grounding objective during training so that the policy learns to exploit visual prompts. Visual guidance provides a pixel-level prompt, while post-training ensures that the model can interpret that prompt; the two components form complementary parts of the same design.²

1.4. Pixelizing Temporal Information

The green mask replaces the object-identification role that would otherwise be carried by a semantic instruction. By directly highlighting the object to be grasped in the visual observation, it eliminates the need to align the instruction with the corresponding object pixels.

The silver-white mask replaces semantic identification of the destination. By directly visualizing the placement location, it reduces the need for the VLA to infer and align the intended target position from language.

The yellow trajectory provides additional information about motion state and short-term history. VLA policies typically observe only a single frame or a short temporal window and do not maintain explicit memory, making information such as "where the object was just now" or "whether it has moved" difficult to recover. We render object positions and end-effector trajectories from the previous few frames into the current frame, with older trajectories gradually fading over time. In this way, temporal information is encoded spatially: history becomes pixels.

Magic-Brain research figure

Three-Color Mask Visualization

Figure 2. Example of three-color mask composition: green indicates the current interaction target, yellow indicates historical trajectories that decay over time, and silver-white indicates the placement region.

2. VLA Post-Training: Enabling Visual-Guidance Following

2.1. Motivation and Design Objectives

Once the semantics of the three-color mask are defined, the mapping between color channels and semantic categories is fixed as a visual protocol used by subsequent data generation and model training. The remaining core problem is how to enable the VLA to understand and consume this protocol. Specifically, when observing an image overlaid with the three-color mask, the policy should interpret each color channel as the corresponding semantic information (green: current interaction/grasp target; silver-white: target placement region; yellow: historical trajectory and temporal information) and generate an action sequence consistent with the visual guidance.

Importantly, we do not modify the VLA network architecture. Instead, mask semantics are encoded entirely in the input observation, and data-driven post-training is used to establish a conditional mapping from color channel → semantic information → action distribution. This design provides three benefits. First, the mask protocol is decoupled from the model architecture, allowing the semantic design to be adjusted without changing the network structure. Second, post-training is performed incrementally on top of the pretrained VLA's general manipulation capabilities, while mixing original data to minimize degradation of its existing language-following ability. Third, visual and language guidance share the same policy network and can therefore be naturally combined: language specifies task semantics, while the mask provides the target, placement region, and temporal information.

2.2. Visual-Guidance Data Generation (Mixed Data Pipeline)

Based on the protocol above, we use a mixed-data pipeline to synthesize visually guided training samples from existing demonstration data. The procedure is as follows:

(1) Trajectory replay and mask synthesis. Existing real-robot and simulation demonstrations are replayed frame by frame. Following the three-color mask semantics, the corresponding mask patterns are synthesized onto observations from each camera view: the current interaction/grasp target is rendered in green, the target placement region in silver-white, and historical trajectory and temporal information in yellow. The masks are projected using camera intrinsic and extrinsic parameters and aligned with the original RGB observations to ensure that their spatial locations remain geometrically consistent with the scene.

(2) Multi-source mixing. Synthetic samples and original data are mixed at a predefined ratio to form the final mixed dataset:

• Mask-guided data: teaches the conditional mapping from "visual guidance → action" and constitutes the key new data introduced at this stage;

• Original data without masks: prevents the model from overfitting to the assumption that "a mask is always present," while preserving the general manipulation and language-following capabilities acquired during pretraining.

(3) Robustness augmentation. To prevent the model from merely "memorizing" fixed mask patterns, we introduce systematic perturbations during synthesis, including random jitter in mask position and shape, random color-channel dropout (forcing the model to produce reasonable actions even when part of the semantic information is missing), and diversified sampling of historical trajectory patterns. We also mix in a certain proportion of "conflict samples," in which the language instruction and mask guidance point to different targets, so that the model can learn the intended priority behavior when the two sources of guidance conflict.

2.3. Capability Targets and Evaluation

After post-training, the VLA is expected to support visual-guidance following. We evaluate this capability along three dimensions:

1. Guidance-following accuracy: given a language instruction plus a three-color mask, can the model correctly approach the green target region, place the object in the silver-white target region, and use the yellow historical trajectory to maintain awareness of short-term motion state? Performance is measured using task success rate and trajectory deviation;

2. Online redirection: during rollout, the mask is dynamically updated (e.g., by moving the current target or target placement region) to test whether the model can respond to changes in visual guidance in real time without resetting the task;

3. Unguided regression test: the model is evaluated on the original tasks without masks to verify that its general manipulation and language-following capabilities do not degrade after post-training.

Together, mask semantic design (the visual protocol) and VLA post-training (protocol consumption) form a closed loop: the former defines "how a human/system expresses spatial constraints to the robot," while the latter ensures that "the robot can correctly interpret and execute those constraints." This provides a foundation for subsequent visually guided task orchestration and human-robot interaction.

2.4. Evaluation Results

We conduct an ablation comparison on MolmoSpaces-Bench. The baseline uses the same magicVLA policy and the same task set; the only variable is whether the observation contains the visual-guidance frame:

• ms-pick (single-stage grasping): success rate increases from 70% to 94.4%, an absolute improvement of 24.4 percentage points;

• pick-and-place (grasp followed by placement): success rate increases from 41% to 84.7%, an absolute improvement of 43.7 percentage points.

Magic-Brain research figure

MolmoSpaces-Bench Evaluation Results

Figure 3. Ablation results on MolmoSpaces-Bench: visual guidance improves the success rate by +24.4 pp and +43.7 pp on the two task categories, respectively.

Across the two task categories evaluated here, pick-and-place, which includes both grasping and placement stages, shows a larger absolute improvement. One possible explanation is that the placement stage must answer the question "where should the object be placed?", and this type of region-level grounding is more challenging for language-only instructions. For example, a "basket" may correspond to a spatial placement region rather than a single point-like object, whereas the silver-white mask can directly visualize the target placement region in the observation. In this experiment, visual guidance therefore provides a more pronounced benefit during the placement stage.

3. Overview of the Magic-Brain Framework

3.1. Three-Layer Architecture

Magic-Brain is our embodied-agent framework and adopts a three-layer architecture: apps → gateway → agent_core. The apps layer implements task-facing functional Agents, including navigation, grasping, interaction, and sentinel Agents. The sentinel Agent is primarily responsible for proactive and event-triggered tasks. Together, these four types of Agents integrate the robot's algorithmic capabilities, organizing navigation, grasping, interaction, and proactive perception/response into a unified embodied-agent system.

Magic-Brain research figure

Magic-Brain Three-Layer Architecture

Figure 4. Three-layer architecture of Magic-Brain.

The gateway layer is responsible for connection management, event broadcasting, and authentication, serving as the infrastructure layer for cross-robot embodiment and large-scale deployment. agent_core is the core execution engine of the Agent framework and implements key Agent mechanisms including Plan-and-Act, tool management, Memory, and Context. With this layered design, a single agent_core can support multiple functional Agents in the apps layer.

3.2. Planning, Interruption, and Replanning

The main loop is driven by Plan-and-Act and follows a closed-loop process of "plan → act → observe → replan." Subtasks are organized as a directed acyclic graph (DAG). Each SubTask has one of five states: READY, BLOCKED, RUNNING, DONE, or FAILED. Dependency checks ensure that the task graph remains acyclic. At runtime, batches of READY subtasks are computed dynamically, and mutually independent tasks are executed concurrently using asyncio.gather.

Interruptions are the norm in embodied scenarios. We separate InterruptBus (PAUSE/ABORT/ESTOP) from SteeringBus (STEER/FOLLOWUP) as orthogonal channels: the former expresses "stop," while the latter expresses "change the plan." Because their semantics differ, they do not share the same channel. For actuation safety, tools declared with has_safety=True pass through a five-stage tool pipeline that includes SafetyHook and Gate.evaluate arbitration. A subtask in the FAILED or DEVIATED state triggers replanning, with max_replans=3. Checkpoints are restored at subtask granularity rather than step granularity because embodied actions are not reversible: once a robotic arm has poured out water, the physical action cannot be rolled back; what can be re-executed is the task.

Returning to the question posed at the beginning: we did not make the VLA "smarter"; we simply made the relevant information directly visible to it. In embodied systems, "seeing" is often not only a question of perception-model capacity, but also of how information reaches the model. Language specifies "what to do," while pixels specify "which one" and "where." Assigning each information channel to the model best suited to process it is the division of labor we have found most effective so far.

References

1. Wang, Z., Chen, Y., Liu, Y., Ye, J., Chen, P., Lu, C., Liu, S., Yu, B., Jia, J. “VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models.” arXiv:2603.22003, 2026. Proposes overlaying visual markers (crosshairs and bounding boxes) on observation images as an input interface for VLA models. A System 2 VLM planner generates the markers, a System 1 controller consumes them, and an auxiliary visual-grounding objective is introduced during training.

2. Yang, J., Huang, W., Zhang, J., Hu, M., Guo, H. “In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use.” arXiv:2608.05738, 2026. Reports that free-form textual CoT can degrade low-level VLA control and argues that the VLA should consume already grounded information rather than generate language. The method uses in-context post-training together with agentic tool use.

3. Zhang, Y., Wang, R., Lin, J., Wang, Z., Qi, X. “Retrieval-VLA: Training-Free In-Context Adaptation for Vision-Language-Action Models.” CVPR 2026. Combines retrieval augmentation with in-context learning to enable pretrained VLA models to adapt to new tasks without gradient-based updates.

4. Kim, et al. “MolmoSpaces.” arXiv:2602.11337, 2026. An embodied-AI evaluation benchmark; MolmoSpaces-Bench includes pick and pick-and-place tasks.