Resource-Aware Diffusion Model Compression for Mobile and Embedded Multimodal Intelligence Systems

Authors

  • Malcolm J. Perez Department of Computer Science, Colorado State University, Fort Collins, CO, USA.
  • Seamor Parekh School of Computing, Clemson University, Clemson, SC, USA.
  • Ankit R. Rha Department of Computer Science, University of Central Florida, Orlando, FL, USA.
  • Martins M. Parzez School of Electrical Engineering and Computer Science, Oregon State University, Corvallis, OR, USA.

Keywords:

diffusion models, model compression, mobile AI, edge computing, multimodal intelligence, resource-aware systems, sustainable computing, fairness, governance

Abstract

Diffusion models have emerged as powerful generative frameworks capable of producing high-fidelity multimodal content that spans text, images, speech, and sensor streams. Their deployment on mobile and embedded platforms, however, is severely constrained by memory footprints, energy budgets, and real-time latency requirements that conflict with the iterative denoising process at their core. This paper presents a system-level analysis of resource-aware diffusion model compression, moving beyond isolated algorithmic contributions to examine the structural trade-offs, architectural implications, and socio-technical dimensions of shrinking these models for edge-native multimodal intelligence. We survey a continuum of compression paradigms—quantization, pruning, knowledge distillation, neural architecture search, and sampling trajectory reduction—and evaluate them through the lenses of hardware-software co-design, on-device orchestration, sustainable computing, fairness, and governance. The discussion foregrounds infrastructure-level considerations such as heterogeneous accelerator utilization, memory hierarchies, and federal learning pipelines that distribute compression across devices. By integrating case illustrations from mobile vision-language systems, on-device audio synthesis, and sensor-conditioned scene generation, we identify persistent tensions between fidelity, latency, and equity. We argue that compression must be treated not as a post-hoc optimization but as a first-class design constraint that shapes the entire training-inference-operations lifecycle. The paper concludes by outlining policy implications and future research directions for robust, inclusive, and energy-proportional diffusion intelligence at the extreme edge.

References

1. Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33, 6840–6851.

2. Rombach, R., Blattmann, A., Lorenz, D., Esser, P., & Ommer, B. (2022). High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 10684–10695).

3. Hinton, G., Vinyals, O., & Dean, J. (2015). Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531.

4. Han, S., Mao, H., & Dally, W. J. (2015). Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding. arXiv preprint arXiv:1510.00149.

5. Zhu, C., Han, S., Mao, H., & Dally, W. J. (2016). Trained ternary quantization. arXiv preprint arXiv:1612.01064.

6. Li, Y., Wang, Y., Chen, X., Li, Z., & Wang, Y. (2023). SnapFusion: Text-to-image diffusion model on mobile devices within two seconds. Advances in Neural Information Processing Systems, 36.

7. Li, X., Liu, Y., Lian, L., Yang, H., Dong, Z., Kang, D., Zhang, S., & Keutzer, K. (2023). Q-Diffusion: Quantizing diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 17535–17544).

8. Salimans, T., & Ho, J. (2022). Progressive distillation for fast sampling of diffusion models. arXiv preprint arXiv:2202.00512.

9. Howard, A. G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., & Adam, H. (2017). MobileNets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861.

10. Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., & Chen, L. C. (2018). MobileNetV2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 4510–4520).

11. Tan, M., & Le, Q. V. (2019). EfficientNet: Rethinking model scaling for convolutional neural networks. In International Conference on Machine Learning (pp. 6105–6114). PMLR.

12. Hooker, S., Courville, A., Clark, G., Dauphin, Y., & Frome, A. (2019). What do compressed deep neural networks forget? arXiv preprint arXiv:1911.05248.

13. Chen, C., Wang, C., Li, Y., Wan, Z., Geng, M., Xiao, J., ... & Peng, Y. (2026). JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators. arXiv preprint arXiv:2606.28421.

14. Patterson, D., Gonzalez, J., Le, Q., Liang, C., Munguia, L.-M., Rothchild, D., So, D., Texier, M., & Dean, J. (2021). Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350.

15. Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 3645–3650).

16. Satyanarayanan, M. (2017). The emergence of edge computing. Computer, 50(1), 30–39.

17. Baltrusaitis, T., Ahuja, C., & Morency, L.-P. (2019). Multimodal machine learning: A survey and taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(2), 423–443.

18. Bonawitz, K., Ivanov, V., Kreuter, B., Marcedone, A., McMahan, H. B., Patel, S., Ramage, D., Segal, A., & Seth, K. (2017). Practical secure aggregation for privacy-preserving machine learning. In Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security (pp. 1175–1191).

19. Mittelstadt, B. (2019). Principles alone cannot guarantee ethical AI. Nature Machine Intelligence, 1(11), 501–507.

20. David, R., Duke, J., Jain, A., Janapa Reddi, V., Jeffries, N., Li, J., Kreeger, N., Nappier, I., Natraj, M., Regev, S., Rhodes, R., Wang, T., & Warden, P. (2021). TensorFlow Lite Micro: Embedded machine learning on TinyML systems. Proceedings of Machine Learning and Systems, 3, 800–811.

Downloads

Published

2026-07-15