Open Access

A Classification of Architectural Trade-Offs in Deploying Generative Models to Mobile Applications Under Device Resource Constraints

4 Mobile Applications Developer, Mobilesource Corp., Miami, United States of America

Abstract

Generative models have entered consumer mobile applications as a routine product feature, and their computational profile departs from that of the discriminative models that preceded them on the device. Image synthesis models exceed the memory and arithmetic budgets of a smartphone session, forcing the processing pipeline to span the device boundary. This review proposes a classification of the architectural trade-offs that govern such divided pipelines. Five axes organize the design space: the execution locus of each pipeline stage, the temporal contract the application makes with the user, the boundary that user data crosses, the governance of output quality, and the economics of a single invocation. The classification rests on a criterion that receives limited treatment in the deployment literature. The execution locus of a stage follows from how often a user invokes it within a session, the marginal cost of a remote call, and the technical feasibility of local execution; the latter enters the decision only after the frequency question has been settled. A stage repeated dozens of times within a single session belongs on the device, even when a server could perform it faster. The review also maps the mechanisms that substitute for automated assessment of generative output quality and identifies the absence of a computable proxy for aesthetic acceptability as an open problem. Observations from engineering practice in consumer mobile applications with generative image features supply illustrations for each axis. Those observations are descriptive and carry no controlled measurement.

Keywords

References

Cai, G., Tian, R., Yang, L., Jia, Y., Li, L., & Wang, J. (2026). Efficient inference for edge large language models: A survey. Tsinghua Science and Technology, 31(3), 1365–1380. https://doi.org/10.26599/TST.2025.9010166
Duan, S., Wang, D., Ren, J., Lyu, F., Zhang, Y., Wu, H., & Shen, X. (2023). Distributed artificial intelligence empowered by end-edge-cloud computing: A survey. IEEE Communications Surveys & Tutorials, 25(1), 591–624. https://doi.org/10.1109/COMST.2022.3218527
Ganesh, P., Tran, C., Shokri, R., & Fioretto, F. (2025). The data minimization principle in machine learning. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (pp. 3075–3093). Association for Computing Machinery. https://doi.org/10.1145/3715275.3732195
Hartwig, S., Engel, D., Sick, L., Kniesel, H., Payer, T., Poonam, P., Glöckler, M., Bäuerle, A., & Ropinski, T. (2025). A survey on quality metrics for text-to-image generation. IEEE Transactions on Visualization and Computer Graphics, 31(10), 9464–9483. https://doi.org/10.1109/TVCG.2025.3585077
Hu, H., Huang, Y., Chen, Q., Zhuo, T. Y., & Chen, C. (2023). A first look at on-device models in iOS apps. ACM Transactions on Software Engineering and Methodology, 33(1), Article 26. https://doi.org/10.1145/3617177
Lee, H. M., Yadav, D., Lee, S., Govindarazan, K., Chen, C., & Sundar, S. S. (2025). While we wait... How users perceive waiting times and generation cues during AI image generation. In Extended Abstracts of the 2025 CHI Conference on Human Factors in Computing Systems (Article 602). Association for Computing Machinery. https://doi.org/10.1145/3706599.3719725
Luccioni, A. S., Jernite, Y., & Strubell, E. (2024). Power hungry processing: Watts driving the cost of AI deployment? In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (pp. 85–99). Association for Computing Machinery. https://doi.org/10.1145/3630106.3658542
Matsubara, Y., Levorato, M., & Restuccia, F. (2022). Split computing and early exiting for deep learning applications: Survey and research challenges. ACM Computing Surveys, 55(5), Article 90. https://doi.org/10.1145/3527155
Semerikov, S. O., Vakaliuk, T. A., Kanevska, O. B., Ostroushko, O. A., & Kolhatin, A. O. (2025). Edge intelligence unleashed: A survey on deploying large language models in resource-constrained environments. Journal of Edge Computing, 4(2), 179–233. https://doi.org/10.55056/jec.1000
Tan, F. F.-Y., Messerschmidt, M. A., Yin, W., & Nov, O. (2026). The impact of response latency and task type on human-LLM interaction and perception. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (Article 14). Association for Computing Machinery. https://doi.org/10.1145/3772318.3790716
Terekhov, V. (2024a). Approaches to optimizing power consumption of Android applications: Practical methods and tools. American Research Journal of Computer Science and Information Technology, 7(1). https://www.arjonline.org/papers/arjcsit/v7-i1/16.pdf
Terekhov, V. (2024b). Implementation of machine learning in Android applications. International Journal of Computer (IJC), 53(1), 72–79. https://ijcjournal.org/InternationalJournalOfComputer/article/view/2306
Tian, Y., Liu, Y., Wang, S., & Kwong, S. (2025). Quality assessment for text-to-image generation: A survey. IEEE MultiMedia, 32(2), 44–52. https://doi.org/10.1109/MMUL.2025.3538862
Yang, L., Zhang, Z., Song, Y., Hong, S., Xu, R., Zhao, Y., Shao, Y., Zhang, W., Cui, B., & Yang, M.-H. (2024). Diffusion models: A comprehensive survey of methods and applications. ACM Computing Surveys, 56(4), Article 105. https://doi.org/10.1145/3626235

Similar Articles

1-10 of 87

You may also start an advanced similarity search for this article.