Memory as Plans: World-Action Modeling with Memory-Grounded Planning
Abstract
MaP-WAM improves non-Markovian robotic manipulation by separating memory-grounded planning from plan-conditioned execution, using compact episodic segment records and progress-calibrated action chunks to maintain fixed inference latency.
Mainstream robotic policies often adopt a Markovian formulation, but many complex real-world manipulation tasks are inherently non-Markovian, requiring long-horizon memory beyond the current observation. Existing memory mechanisms often rely on language summaries, growing visual windows, or their combinations, and may therefore lose fine-grained visual evidence or face a trade-off between history coverage and execution efficiency. We introduce MaP-WAM, a Memory-as-Plans framework that decomposes memory-dependent world-action modeling into memory-grounded planning and plan-conditioned execution, and uses long-term multimodal episodic context as planning-time evidence rather than repeatedly conditioning the executor on the full history. MaP-WAM represents memory as completed segment records containing language instructions and sparse visual context, and converts this episodic memory into compact plans comprising the next segment-level language plan and corresponding visual guidance. A World-Action-Progress (WAP) model executes each plan over an unknown duration by jointly predicting action chunks and corresponding execution progress at inference time, calibrating predicted progress through plan-observation alignment for adaptive segment transitions and closed-loop context updates. MaP-WAM keeps the executor context length fixed, while structured attention further enables key-value caching in both planning and execution. MaP-WAM achieves state-of-the-art performance on RMBench with an 83.3% success rate and attains 78.0% success on real-robot tasks, while maintaining approximately constant executor inference latency as task history grows.
Community
Project Page: https://sizhezhao.github.io/projects/MaP-WAM/
GitHub: https://github.com/aipixel/MaP-WAM
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory (2026)
- StageWAM: Joint-Embedding Stage Prediction for World-Action Models in Robot Manipulation (2026)
- Foresight Without Seeing: Latent Futures for World Action Models (2026)
- World Tokens: Enhancing Embodied Policies with Training-Time World Modeling (2026)
- Latent Action as Intention Enables Efficient Future Imagination for World Action Models (2026)
- SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation (2026)
- JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.11561 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper