Optimizing Resource Allocation for Elastic Tensorflow Container Scheduling on Multi-core Servers
DOI:
https://doi.org/10.19153/cleiej.28.6.2Keywords:
Multi-Core Server, Orchestrators, Elastic Tensorflow, Container, DockerAbstract
In this paper, we delve into the necessity of applying coupled application-container elasticity in the context of resource allocation within multi-tenant scenarios, specifically targeting CNN training procedures on multi-core architectures. Our focus is primarily on elucidating the process of endowing a parallel application, exemplified by Tensorflow, with dynamic thread elasticity support. Additionally, we emphasize the criticality of harnessing this functionality in tandem with container resource management by container orchestrators to avert undesirable resource oversubscription and under-utilization effects, all the while optimizing core occupancy. The empirical findings comprise a wide range of resource allocation and dynamic re-allocation strategies, and reinforce the compelling case for elasticity support not only at the container level but also at the application level. These findings underscore the substantial improvements achieved in terms of time-to-solution, resource utilization, and productivity across diverse workloads.
References
F. Rossi, M. Nardelli, and V. Cardellini, “Horizontal and vertical scaling of container-based applica-
tions using reinforcement learning,” in 2019 IEEE 12th International Conference on Cloud Computing
(CLOUD). IEEE, 2019, pp. 329–338.
E. Casalicchio, “Container orchestration: A survey,” Systems Modeling: Methodologies and Tools, pp.
–235, 2019.
S. Abraham, A. K. Paul, R. I. S. Khan, and A. R. Butt, “On the use of containers in high performance
computing environments,” in 2020 IEEE 13th International Conference on Cloud Computing (CLOUD).
IEEE, 2020, pp. 284–293.
S. Shekhar, H. Abdel-Aziz, A. Bhattacharjee, A. Gokhale, and X. Koutsoukos, “Performance
interference-aware vertical elasticity for cloud-hosted latency-sensitive applications,” in 2018 IEEE 11th
International Conference on Cloud Computing (CLOUD). IEEE, 2018, pp. 82–89.
Y. Al-Dhuraibi, F. Paraiso, N. Djarallah, and P. Merle, “Autonomic vertical elasticity of docker con-
tainers with elasticdocker,” in 2017 IEEE 10th international conference on cloud computing (CLOUD).
IEEE, 2017, pp. 472–479.
E. Preeth, F. J. P. Mulerickal, B. Paul, and Y. Sastri, “Evaluation of docker containers based on
hardware utilization,” in 2015 international conference on control communication & computing India
(ICCC). IEEE, 2015, pp. 697–700.
H. T. Ciptaningtyas, B. J. Santoso, and M. F. Razi, “Resource elasticity controller for docker-based web
applications,” in 2017 11th International Conference on Information & Communication Technology and
System (ICTS). IEEE, 2017, pp. 193–196.
Y. H. Oh, S. Kim, Y. Jin, S. Son, J. Bae, J. Lee, Y. Park, D. U. Kim, T. J. Ham, and J. W. Lee,
“Layerweaver: Maximizing resource utilization of neural processing units via layer-wise scheduling,” in
IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE,
, pp. 584–597.
W. Zheng, M. Tynes, H. Gorelick, Y. Mao, L. Cheng, and Y. Hou, “Flowcon: Elastic flow configuration
for containerized deep learning applications,” in Proceedings of the 48th International Conference on
Parallel Processing, 2019, pp. 1–10.
Docker. [Online]. Available: https://www.docker.com/
R. Doukha, S. A. Mahmoudi, M. Zbakh, and P. Manneback, “Deployment of containerized deep learning
applications in the cloud,” in 2020 5th International Conference on Cloud Computing and Artificial
Intelligence: Technologies and Applications (CloudTech). IEEE, 2020, pp. 1–6.
Docker swarm. [Online]. Available: https://docs.docker.com/engine/swarm/
Kubernetes. [Online]. Available: https://kubernetes.io/es/
Marathon. [Online]. Available: https://mesosphere.github.io/marathon/
Apache mesos. [Online]. Available: https://mesos.apache.org/
Cloudify. [Online]. Available: https://cloudify.co/
H. Zhao, W. Cui, Q. Chen, J. Leng, K. Yu, D. Zeng, C. Li, and M. Guo, “Coda: Improving resource
utilization by slimming and co-locating dnn and cpu jobs,” in 2020 IEEE 40th International Conference
on Distributed Computing Systems (ICDCS). IEEE, 2020, pp. 853–863.
J. Mohan, A. Phanishayee, J. Kulkarni, and V. Chidambaram, “Looking beyond {GPUs} for {DNN}
scheduling on {Multi-Tenant} clusters,” in 16th USENIX Symposium on Operating Systems Design and
Implementation (OSDI 22), 2022, pp. 579–596.
Y. Peng, Y. Bao, Y. Chen, C. Wu, and C. Guo, “Optimus: an efficient dynamic resource scheduler for
deep learning clusters,” in Proceedings of the Thirteenth EuroSys Conference, 2018, pp. 1–14.
W. Xiao, R. Bhardwaj, R. Ramjee, M. Sivathanu, N. Kwatra, Z. Han, P. Patel, X. Peng, H. Zhao,
Q. Zhang et al., “Gandiva: Introspective cluster scheduling for deep learning,” in 13th USENIX Sym-
posium on Operating Systems Design and Implementation (OSDI 18), 2018, pp. 595–610.
J. Han, M. M. Rafique, L. Xu, A. R. Butt, S.-H. Lim, and S. S. Vazhkudai, “Marble: A multi-gpu aware
job scheduler for deep learning on hpc systems,” in 2020 20th IEEE/ACM International Symposium on
Cluster, Cloud and Internet Computing (CCGRID). IEEE, 2020, pp. 272–281.
S. Wang, O. J. Gonzalez, X. Zhou, T. Williams, B. D. Friedman, M. Havemann, and T. Woo, “An
efficient and non-intrusive gpu scheduling framework for deep learning training systems,” in SC20:
International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE,
, pp. 1–13.
H. Zhang, L. Stafman, A. Or, and M. J. Freedman, “Slaq: quality-driven scheduling for distributed
machine learning,” in Proceedings of the 2017 Symposium on Cloud Computing, 2017, pp. 390–404.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 Leandro Libutti

This work is licensed under a Creative Commons Attribution 4.0 International License.
CLEIej is supported by its home institution, CLEI, and by the contribution of the Latin American and international researchers community, and it does not apply any author charges whatsoever for submitting and publishing. Since its creation in 1998, all contents are made publicly accesibly. The current license being applied is a (CC)-BY license (effective October 2015; between 2011 and 2015 a (CC)-BY-NC license was used).