Benchmarking GPU Passthrough Performance on Docker for AI Cloud System
DOI:
https://doi.org/10.47709/brilliance.v5i2.6794Keywords:
Artificial Intelligence, Docker Container, GPU, GPU Passthrough, NVIDIA CudaAbstract
The use of artificial intelligence (AI), which depends only on CPU resources, tends to result in longer execution times or CPU time. Especially when handling large amounts or complex workloads. To overcome that issue, the use of a graphics processing unit (GPU) becomes a significant support. GPUs can significantly speed up AI inferences through their parallel architecture. One recent approach to integrating GPUs into an AI system is called GPU passthrough. Either natively (native environment), or through Docker environment. However, until recently, the efficiency and results between those methods have remained unexplored, particularly in local cloud environment.This study aimed to compare GPU performance between native and Docker environment using a 10.000 x 10.000 matrix multiplication workload with the TensorFlow frameworks. Execution time and GPU performance measured using the nvidia-smi tool. Data is recorded automatically in CSV format. The researcher used the NVIDIA CUDA environment to ensure full compatibility with GPU acceleration.The result demonstrated that GPU processing in native environment had faster average time, as in 1.52 seconds. In another case, GPU passthrough in docker environment demonstrated higher GPU utilization, as in 86.2% but had a longer execution time.These findings indicate that GPU overhead occurred in docker environment due to the containerization layer. On the contrary, the native environment resulted in shorter execution time, even though it did not maximize the GPU utilization. These results provide valuable basis data for technical decision-making in GPU-based AI deployment in a limited environment.
References
Belkhiri, A., & Dagenais, M. (2024). Analyzing GPU Performance in Virtualized Environments: A Case Study. Future Internet, 16(3). https://doi.org/10.3390/fi16030072
Chang, C. H., Yang, C. T., Lee, J. Y., Lai, C. L., & Kuo, C. C. (2020). On construction and performance evaluation of a virtual desktop infrastructure with GPU accelerated. IEEE Access, 8, 170162–170173. https://doi.org/10.1109/ACCESS.2020.3023924
Choi, H. S., Kim, Y., Lee, J., & Kim, Y. (2021). Empirical performance evaluation of communication libraries for MULTI-GPU based distributed deep learning in a container environment. KSII Transactions on Internet and Information Systems, 15(3), 911–931. https://doi.org/10.3837/tiis.2021.03.006
?isar, P., Erlenvajn, D., & Maravi? ?isar, S. (2018). Implementation of software-defined networks using open-source environment. In Tehnicki Vjesnik (Vol. 25, pp. 222–230). Strojarski Facultet. https://doi.org/10.17559/TV-20160928094756
Kumar, M., & Kaur, G. (2022). Study of container-based JupyterLab and AI Framework on HPC with GPU usage. 2022 International Conference on Smart Generation Computing, Communication and Networking, SMART GENCON 2022. https://doi.org/10.1109/SMARTGENCON56628.2022.10084107
Kurkure, U., Sivaraman, H., & Vu, L. (2017). Machine Learning Using Virtualized GPUs in Cloud Environments. In J. M. Kunkel, R. Yokota, M. Taufer, & J. Shalf (Eds.), High Performance Computing (pp. 591–604). Springer International Publishing.
Lingayat, A., Badre, R. R., & Gupta, A. K. (2018). Integration of linux containers in openstack: An introspection. Indonesian Journal of Electrical Engineering and Computer Science, 12(3), 1094–1105. https://doi.org/10.11591/ijeecs.v12.i3.pp1094-1105
NVIDIA. (n.d.-a). Nvidia Container Toolkit - Architecture Overview. Retrieved August 4, 2025, from https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/arch-overview.html
NVIDIA. (n.d.-b). Nvidia SMI. Retrieved August 5, 2025, from https://developer.nvidia.com/system-management-interface
Oh, J., Kim, S., & Kim, Y. (2019). Toward an adaptive fair GPU sharing scheme in container-based clusters. Proceedings - 2018 IEEE 3rd International Workshops on Foundations and Applications of Self* Systems, FAS*W 2018, 79–85. https://doi.org/10.1109/FAS-W.2018.00029
Openja, M., Majidi, F., Khomh, F., Chembakottu, B., & Li, H. (2022). Studying the Practices of Deploying Machine Learning Projects on Docker. ACM International Conference Proceeding Series, 190–200. https://doi.org/10.1145/3530019.3530039
Shea, R., & Liu, J. (n.d.). On GPU Pass-Through Performance for Cloud Gaming: Experiments and Analysis. https://doi.org/https://doi.org/10.1109/NetGames.2013.6820614
Shetty, J., Upadhaya, S., Rajarajeshwari, H. S., Shobha, G., & Chandra, J. (2017). An empirical performance evaluation of docker container, openstack virtual machine and bare metal server. Indonesian Journal of Electrical Engineering and Computer Science, 7(1), 205–213. https://doi.org/10.11591/ijeecs.v7.i1.pp205-213
Shi, S., Wang, Q., Xu, P., & Chu, X. (2017). Benchmarking State-of-the-Art Deep Learning Software Tools. http://arxiv.org/abs/1608.07249
Tiying, F., & Zhengwei, L. (2019). A GPU resource scheduling method and apparatus based on AI cloud.
Walters, J. P., Younge, A. J., Kang, D.-I., Yao, K.-T., Kang, M., Crago, S. P., & Fox, G. C. (n.d.). GPU Passthrough Performance: A Comparison of KVM, Xen, VMWare ESXi, and LXC for CUDA and OpenCL Applications.
Wang, Y. E., Wei, G.-Y., & Brooks, D. (2019). Benchmarking TPU, GPU, and CPU Platforms for Deep Learning. http://arxiv.org/abs/1907.10701
Zhao, L., Jin, Y., Hu, G., Zhou, W., Wei, H., Li, R., Zhu, X., Xu, Y., Jin, J., & Li, Q. (2025). Design and Implementation of GPU Pass-Through System Based on OpenStack. Computation, 13(2). https://doi.org/10.3390/computation13020038
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2025 Ahmad Faisal Sani, Rifa Khoirunisa, Darmawan Lahru Riatma, Yusuf Fadlila Rachman, Masbahah

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.















