Decide what to remember
A vision–language model reads the instruction and proposes the identities, progress and procedures that may matter later.
Watch the cube–container relation.
A simple memory layer. A longer view of the task.
1CMU · 2Tulane · 3NYU · 4Princeton · 5Columbia
*Equal contribution
01 / The idea
A current image cannot tell a robot which object was hidden, how many actions it has completed, or which route it saw earlier. SimpleARM keeps this task-relevant history available for control.
A vision–language model reads the instruction and proposes the identities, progress and procedures that may matter later.
Frozen perception and event tools verify observations and update compact memory as the robot and its surroundings move.
A history-dependent subgoal triggers the relevant read. Memory resolves the referent; current perception locates it for the frozen policy.
Three forms of task state
The added memory is episode-scoped and training-free. The pretrained grounded-subgoal composer, robot policy and perception models stay frozen.
02 / Memory in action
Follow a real ButtonUnmaskSwap episode. The colors disappear and the covers move, but the memory keeps track of what is underneath.
“First press both buttons on the table, then pick up the container hiding the green cube, finally pick up another container hiding the red cube.”
t = 0The green and red cubes are visible. The instruction selects these identities for memory.
Active operationDetection · color verification
Seed 7, episode 47, from a qualitative replication of the final configuration. The episode succeeds at step 596. These are recorded front-camera crops; the callouts and memory summaries explain the documented trace. A composer proposal is an intermediate target, not a separate no-memory rollout. Pixel verification, DINO identity features and optical flow were active here; SAM was unavailable on this rollout host.
03 / Across the benchmark
Across all 16 RoboMME tasks, SimpleARM reaches 67.17% mean success. The largest gains appear when control needs persistent relations or references to the past.
Success (%) · one tick = one percentage point
Swipe to see the complete chart →
SimpleARM: mean over 3 seeds. Baselines: published RoboMME results ↗.
| Suite | Task | No memory | MemER | FrameSamp | SimpleARM |
|---|
System-level comparison with published baselines. Gains vary by task: Counting remains a limitation, including 0% on StopCube. The added memory receives no training; the pretrained policies use RoboMME’s existing checkpoints.
04 / What makes the difference?
Remove a memory component or change how it is read, then compare with its matched full-method control. Every ablation covers all 800 episodes.
Full method Ablation · success (%)
Swipe to see the complete chart →
800 same-host, same-seed, same-task, same-episode pairs per row · seed 11. Each bead between endpoints = 1 pp of difference.
Removing route memory lowers full-benchmark success by 11.50 pp; removing demonstration references lowers it by 9.13 pp.
Replacing structured retrieval with a VLM decision at every proposed subgoal lowers success by 11.88 pp.
Changing the source to received frames gives 0.00 pp aggregate change. The table does not show that every component helps every task.
| Intervention | Full (%) | Ablation (%) | Δ (pp) |
|---|
The matched control varies by row because outcomes are matched to their actual execution host. These are single-seed interventions; component effects overlap and should not be added together.
05 / A closer look at visual history
We also fine-tune and evaluate RoboMME’s FrameSamp + ModuL with the same Recent-N strategy at training and test time.
Recent32 has the highest observed mean of these three settings. This is a separate family of trained models from SimpleARM.
Success (%) · one rung = 1 pp
Whiskers: ±1 sample SD across three evaluation seeds. One training seed (42) and one fixed final checkpoint per setting.
Recent4/16/32 each use evaluation seeds 7, 11 and 23, for 7,200 episodes in total. Some evaluation rounds ran on different hosts, so host effects are not isolated. Recent8 has one completed seed7 evaluation: 211/800 (26.375%). Its evaluated checkpoint used a 70k EMA warm start and reset optimizer; it is not included in the three-seed chart. Earlier inference-only Recent interventions on Uniform-trained weights are a separate experiment.
06 / Conclusion
SimpleARM turns interaction history into compact, task-relevant state. A memory plan decides what to retain, frozen tools maintain it as the world changes, and selective retrieval brings it back when a subgoal needs the past.
This simple, training-free memory layer helps frozen robot policies act on identities, progress and procedures that the current image alone cannot reveal.
If you find SimpleARM useful, please cite our paper.
@misc{zhang2026simpleagenticmemorygeneralist,
title={Simple Agentic Memory for Generalist Robot Policies},
author={Yuyou Zhang and Yunbei Zhang and Miao Li and Janet Wang and Zijian Jin and Shilong Liu and Ding Zhao},
year={2026},
eprint={2609.36595},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2609.36595},
}