Optimizing Resource Allocation for Elastic Tensorflow Container Scheduling on Multi-core Servers

Authors

  • Leandro Libutti III LIDI, National University of La Plata, Argentina

DOI:

https://doi.org/10.19153/cleiej.28.6.2

Keywords:

Multi-Core Server, Orchestrators, Elastic Tensorflow, Container, Docker

Abstract

In this paper, we delve into the necessity of applying coupled application-container elasticity in the context of resource allocation within multi-tenant scenarios, specifically targeting CNN training procedures on multi-core architectures. Our focus is primarily on elucidating the process of endowing a parallel application, exemplified by Tensorflow, with dynamic thread elasticity support. Additionally, we emphasize the criticality of harnessing this functionality in tandem with container resource management by container orchestrators to avert undesirable resource oversubscription and under-utilization effects, all the while optimizing core occupancy. The empirical findings comprise a wide range of resource allocation and dynamic re-allocation strategies, and reinforce the compelling case for elasticity support not only at the container level but also at the application level. These findings underscore the substantial improvements achieved in terms of time-to-solution, resource utilization, and productivity across diverse workloads.

References

F. Rossi, M. Nardelli, and V. Cardellini, “Horizontal and vertical scaling of container-based applica-

tions using reinforcement learning,” in 2019 IEEE 12th International Conference on Cloud Computing

(CLOUD). IEEE, 2019, pp. 329–338.

E. Casalicchio, “Container orchestration: A survey,” Systems Modeling: Methodologies and Tools, pp.

–235, 2019.

S. Abraham, A. K. Paul, R. I. S. Khan, and A. R. Butt, “On the use of containers in high performance

computing environments,” in 2020 IEEE 13th International Conference on Cloud Computing (CLOUD).

IEEE, 2020, pp. 284–293.

S. Shekhar, H. Abdel-Aziz, A. Bhattacharjee, A. Gokhale, and X. Koutsoukos, “Performance

interference-aware vertical elasticity for cloud-hosted latency-sensitive applications,” in 2018 IEEE 11th

International Conference on Cloud Computing (CLOUD). IEEE, 2018, pp. 82–89.

Y. Al-Dhuraibi, F. Paraiso, N. Djarallah, and P. Merle, “Autonomic vertical elasticity of docker con-

tainers with elasticdocker,” in 2017 IEEE 10th international conference on cloud computing (CLOUD).

IEEE, 2017, pp. 472–479.

E. Preeth, F. J. P. Mulerickal, B. Paul, and Y. Sastri, “Evaluation of docker containers based on

hardware utilization,” in 2015 international conference on control communication & computing India

(ICCC). IEEE, 2015, pp. 697–700.

H. T. Ciptaningtyas, B. J. Santoso, and M. F. Razi, “Resource elasticity controller for docker-based web

applications,” in 2017 11th International Conference on Information & Communication Technology and

System (ICTS). IEEE, 2017, pp. 193–196.

Y. H. Oh, S. Kim, Y. Jin, S. Son, J. Bae, J. Lee, Y. Park, D. U. Kim, T. J. Ham, and J. W. Lee,

“Layerweaver: Maximizing resource utilization of neural processing units via layer-wise scheduling,” in

IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE,

, pp. 584–597.

W. Zheng, M. Tynes, H. Gorelick, Y. Mao, L. Cheng, and Y. Hou, “Flowcon: Elastic flow configuration

for containerized deep learning applications,” in Proceedings of the 48th International Conference on

Parallel Processing, 2019, pp. 1–10.

Docker. [Online]. Available: https://www.docker.com/

R. Doukha, S. A. Mahmoudi, M. Zbakh, and P. Manneback, “Deployment of containerized deep learning

applications in the cloud,” in 2020 5th International Conference on Cloud Computing and Artificial

Intelligence: Technologies and Applications (CloudTech). IEEE, 2020, pp. 1–6.

Docker swarm. [Online]. Available: https://docs.docker.com/engine/swarm/

Kubernetes. [Online]. Available: https://kubernetes.io/es/

Marathon. [Online]. Available: https://mesosphere.github.io/marathon/

Apache mesos. [Online]. Available: https://mesos.apache.org/

Cloudify. [Online]. Available: https://cloudify.co/

H. Zhao, W. Cui, Q. Chen, J. Leng, K. Yu, D. Zeng, C. Li, and M. Guo, “Coda: Improving resource

utilization by slimming and co-locating dnn and cpu jobs,” in 2020 IEEE 40th International Conference

on Distributed Computing Systems (ICDCS). IEEE, 2020, pp. 853–863.

J. Mohan, A. Phanishayee, J. Kulkarni, and V. Chidambaram, “Looking beyond {GPUs} for {DNN}

scheduling on {Multi-Tenant} clusters,” in 16th USENIX Symposium on Operating Systems Design and

Implementation (OSDI 22), 2022, pp. 579–596.

Y. Peng, Y. Bao, Y. Chen, C. Wu, and C. Guo, “Optimus: an efficient dynamic resource scheduler for

deep learning clusters,” in Proceedings of the Thirteenth EuroSys Conference, 2018, pp. 1–14.

W. Xiao, R. Bhardwaj, R. Ramjee, M. Sivathanu, N. Kwatra, Z. Han, P. Patel, X. Peng, H. Zhao,

Q. Zhang et al., “Gandiva: Introspective cluster scheduling for deep learning,” in 13th USENIX Sym-

posium on Operating Systems Design and Implementation (OSDI 18), 2018, pp. 595–610.

J. Han, M. M. Rafique, L. Xu, A. R. Butt, S.-H. Lim, and S. S. Vazhkudai, “Marble: A multi-gpu aware

job scheduler for deep learning on hpc systems,” in 2020 20th IEEE/ACM International Symposium on

Cluster, Cloud and Internet Computing (CCGRID). IEEE, 2020, pp. 272–281.

S. Wang, O. J. Gonzalez, X. Zhou, T. Williams, B. D. Friedman, M. Havemann, and T. Woo, “An

efficient and non-intrusive gpu scheduling framework for deep learning training systems,” in SC20:

International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE,

, pp. 1–13.

H. Zhang, L. Stafman, A. Or, and M. J. Freedman, “Slaq: quality-driven scheduling for distributed

machine learning,” in Proceedings of the 2017 Symposium on Cloud Computing, 2017, pp. 390–404.

Downloads

Published

2026-03-13