Energy-Efficient Edge Deployment of Action-Aware Embodied Agents through Memory-Guided Model Compression
Keywords:
embodied AI, edge computing, model compression, action-aware memory, energy efficiency, deployment sustainabilityAbstract
The proliferation of embodied artificial intelligence agents operating in physical and virtual environments demands increasingly sophisticated models that can perceive, reason, and act in real time. Deploying such action-aware agents on resource-constrained edge devices introduces severe energy and latency constraints that traditional cloud-centric paradigms cannot satisfy without sacrificing autonomy or privacy. This paper presents a system-level investigation into energy-efficient edge deployment of embodied agents through memory-guided model compression. We argue that action-aware episodic and working memory structures, originally developed to improve agent performance in interactive domains, can be repurposed as powerful guidance signals for compressing deep neural policies and perception stacks. By exposing the salience of network pathways with respect to temporally extended action sequences, these memory modules enable a new class of compression strategies that preserve task-critical representations while aggressively pruning, quantizing, and distilling components of negligible behavioral impact. The paper examines the architectural trade-offs involved in coupling memory-augmented agent architectures with compression pipelines designed for heterogeneous edge hardware. Detailed discussions cover the infrastructure for deployment, the interplay between compression fidelity and energy consumption, robustness under distributional shift, fairness implications of compressed decision-making, and governance frameworks for sustainable and responsible edge intelligence. We propose a holistic reference architecture that integrates online memory tracking, multi-objective compression policies, and continuous model lifecycle management. The analysis reveals that memory-guided compression can reduce per-inference energy by up to an order of magnitude compared to uniform compression baselines while maintaining interactive task performance, but only if the system design carefully navigates the tensions between forgetting, catastrophic interference, and the need for rapid adaptation. The paper concludes with forward-looking perspectives on policy, standardization, and the co-design of edge hardware and lightweight memory substrates for a new generation of sustainable embodied agents.
References
1. Han, S., Mao, H., & Dally, W. J. (2016). Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding. In International Conference on Learning Representations.
2. Howard, A. G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., ... & Adam, H. (2017). MobileNets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861.
3. Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., ... & Hassabis, D. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529-533.
4. Parisotto, E., & Salakhutdinov, R. (2018). Neural map: Structured memory for deep reinforcement learning. In International Conference on Learning Representations.
5. Hinton, G., Vinyals, O., & Dean, J. (2015). Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531.
6. Chen, T., Moreau, T., Jiang, Z., Zheng, L., Yan, E., Cowan, M., ... & Krishnamurthy, A. (2018). TVM: An automated end-to-end optimizing compiler for deep learning. In USENIX Symposium on Operating Systems Design and Implementation.
7. Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics.
8. Patterson, D., Gonzalez, J., Le, Q., Liang, C., Munguia, L. M., Rothchild, D., ... & Dean, J. (2021). Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350.
9. Molchanov, P., Tyree, S., Karras, T., Aila, T., & Kautz, J. (2017). Pruning convolutional neural networks for resource efficient inference. In International Conference on Learning Representations.
10. Deng, J., Dong, W., Socher, R., Li, L. J., Li, K., & Fei-Fei, L. (2009). ImageNet: A large-scale hierarchical image database. In IEEE Conference on Computer Vision and Pattern Recognition.
11. Xiong, Z., Song, Y., Kang, H., Yan, Q., Jiang, L., Yang, J., ... & Jacobs, N. (2026). ActWorld: From Explorable to Interactive World Model via Action-Aware Memory. arXiv preprint arXiv:2606.17730.
12. Zhuang, Z., Tan, M., Zhuang, B., Liu, J., Guo, Y., Wu, Q., ... & Huang, J. (2018). Discrimination-aware channel pruning for deep neural networks. In Advances in Neural Information Processing Systems.
13. Li, Y., Zhuang, B., Shen, J., Wang, J., & Huang, T. S. (2020). When NAS meets fairness: Exploring the interplay between architecture search and fairness. In European Conference on Computer Vision.
14. Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., & Chen, L. C. (2018). MobileNetV2: Inverted residuals and linear bottlenecks. In IEEE Conference on Computer Vision and Pattern Recognition.
15. Tan, M., & Le, Q. V. (2019). EfficientNet: Rethinking model scaling for convolutional neural networks. In International Conference on Machine Learning.
16. Redmon, J., & Farhadi, A. (2018). YOLOv3: An incremental improvement. arXiv preprint arXiv:1804.02767.
17. Radosavovic, I., Xiao, T., Vinyals, O., & Teh, Y. W. (2022). Masked visual modeling meets reinforcement learning for efficient and general purpose agents. arXiv preprint arXiv:2206.15029.
18. Kang, D., Emmons, J., Abuzaid, F., Bailis, P., & Zaharia, M. (2017). NoScope: Optimizing neural network queries over video streams at scale. Proceedings of the VLDB Endowment, 10(11), 1586-1597.
19. Courbariaux, M., Hubara, I., Soudry, D., El-Yaniv, R., & Bengio, Y. (2016). Binarized neural networks: Training deep neural networks with weights and activations constrained to +1 or -1. arXiv preprint arXiv:1602.02830.
20. Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Engineering Systems and Digital Innovation

This work is licensed under a Creative Commons Attribution 4.0 International License.