Practical Machine Learning for Self-Tuning Buffer Pools in DBMS Mainframe Systems

Authors

  • Eduardo Pingarilho Mendizabal Universidade de Brasília (UnB)
  • Geraldo P. Rocha Filho State University of Southwest Bahia (UESB)
  • Marcelo A. Marotta Universidade de Brasília (UnB)
  • Marcos F. Caetano Universidade de Brasília (UnB)
  • Joao Jose C. Gondim Universidade de Brasília (UnB)
  • Lucas Bondan Rede Nacional de Ensino e Pesquisa (RNP)
  • Aleteia Araujo Universidade de Brasilia (UnB)

DOI:

https://doi.org/10.19153/cleiej.29.4.2

Abstract

This paper investigates the practical applicability of Machine Learning (ML) techniques for buffer pool optimization in Database Management Systems (DBMSs), a critical component traditionally tuned manually in mission-critical environments. While prior work proposes ML-based approaches, most evaluations are limited to simulated settings, leaving their feasibility in production systems—particularly on mainframes—largely unexplored. To address this gap, we present an empirical, production-level validation of an automated, data-driven optimization methodology applied to a relational DBMS running on a mainframe. Our approach integrates Bayesian Optimization based on Gaussian Process Regression (BO-GPR) with a three-phase pipeline: Exploratory Factor Analysis with clustering for metric reduction, LASSO regression for parameter selection, and BO-GPR for fine-grained configuration tuning. We validate the methodology using real workloads from a large-scale financial system, achieving substantial performance gains, including up to a 95\% reduction in maximum synchronous I/O wait time. These results provide reproducible evidence that ML-based optimization is both effective and operationally viable in mainframe-based DBMSs.

References

IBM, DB2 13 for z/OS: Troubleshooting for DB2. Armonk, NY: IBM, 2023.

C. J. Date, An Introduction to Database Systems, 8th ed. Pearson, 2004.

L. Liu and M. T. Ozsu, ¨ Encyclopedia of Database Systems. New York: Springer, 2009.

H. A. Ghozali, M. S. Antarressa, and S. Samidi, “Database Optimization Techniques with Logic Execution Optimization on Microservices Architecture,” CogITo Smart Journal, vol. 9, no. 1, pp. 60–72, 2023.

H. Garcia-Molina, J. D. Ullman, and J. Widom, Database Systems: The Complete Book, 2nd ed. Pearson, 2008.

Kamatkar et al., “Database Performance Tuning and Query Optimization,” in Data Mining and Big Data, LNCS, Tan et al., Eds. Berlin: Springer, 2018, pp. 3–11.

Sullivan et al., “Using probabilistic reasoning to automate software tuning,” in Proc. Joint Conf. Measurement and Modeling of Comput. Syst. NY: ACM, 2004, pp. 404–405.

Y. Zhao et al., “Automatic Configuration Tuning for Databases with Machine Learning: A Survey,” VLDB Journal, vol. 32, pp. 835–862, 2023.

X. Zhou et al., “Database Meets Artificial Intelligence: A Survey,” IEEE Trans. Knowl. Data Eng., vol. 34, no. 3, pp. 1096–1116, 2022.

S. Huang et al., “Survey on Performance Optimization for Database Systems,” Sci. China Inf. Sci., vol. 66, no. 2, p. 121102, 2023.

J. Lu et al., “Speedup Your Analytics: Automatic Parameter Tuning for Databases and Big Data Systems,” Proc. VLDB Endow., vol. 12, no. 12, pp. 1970–1973, 2019.

A. Storm et al., “Adaptive Self-Tuning Memory in DB2,” Proc. VLDB Endow., 2006.

D. V. Aken et al., “An Inquiry into Machine Learning-based Automatic Configuration Tuning on Realworld Database Management Systems,” Proc. VLDB Endow., vol. 14, no. 7, pp. 1241–1253, 2021.

H. R. Vyawahare, “Machine Learning: A Solution Approach for Complex Problems,” Int. J. Sci. Res. Eng. Manag., vol. 6, no. 4, pp. 1–6, 2022.

K. K. Joshi et al., “Machine Learning - Learning Techniques, CNN, Languages and APIs,” Int. J. Sci. Res. Comput. Sci. Eng. Inf. Technol., vol. 5, no. 3, pp. 23–30, 2020.

W.-M. Lee, Python Machine Learning, 1st ed. Wiley, 2019.

J. Tan et al., “iBTune: Individualized Buffer Tuning as a Service,” Proc. VLDB Endow., vol. 12, no. 10,

pp. 1221–1234, 2019.

D. V. Aken et al., “Automatic DBMS Tuning Through Large-scale ML,” in Proc. ACM SIGMOD, 2017, pp. 1009–1024.

A. Vasyliev, “PGTune: PostgreSQL Configuration Tuning,” https://pgtune.leopard.in.ua/, 2024.

Y. Zhu et al., “BestConfig: Tapping the Performance Potential of Systems via Automatic Configuration

Tuning,” in Proc. ACM SoCC, 2017.

K. Kanellis et al., “LlamaTune: Hierarchical Configuration Tuning for DBMSs,” Proc. VLDB Endow.,vol. 15, no. 11, pp. 2973–2984, 2022.

A. Fekry et al., “Tuneful: An Online Significance-Aware Configuration Tuner for Big Data Analytics,” in Proc. ICDE, 2020.

J. Ge et al., “WATuning: A Workload-Aware Tuning System with Attention-Based Deep Reinforcement Learning,” in J. Syst. Softw., 2021.

I. Trummer, “DB-BERT: a Database Tuning Tool that ’Reads the Manual’,” in Proc. SIGMOD, 2021, pp. 1085–1104.

X. Ding et al., “A General Approach to Scalable Buffer Pool Management,” IEEE Trans. Parallel Distrib. Syst., vol. 27, no. 8, pp. 2182–2195, 2016.

W. Effelsberg and T. Haerder, “Principles of Database Buffer Management,” ACM Trans. Database Syst., vol. 9, no. 4, pp. 560–595, 1984.

P. Martin et al., “Automated Configuration of Multiple Buffer Pools,” Comput. J., vol. 49, no. 4, pp.

–499, 2006.

Intellimagic, “Db2 for z/OS Monitoring and Performance Management,” 2024.

E. Mendizabal, G. P. Rocha Filho, M. A. Marotta, M. F. Caetano, J. a. J. Gondim, L. Bondan, and A. Araujo, “Self-Tuning DBMS: A Data-Driven Approach to Buffer Pool Optimization in Enterprise Systems,” in LI Conferencia Latinoamericana de Inform´atica (CLEI). Valpara´?so, Chile: CLEI, 2025.

E. P. Mendizabal, G. P. Rocha Filho, and A. Araujo, “Otimiza¸c˜ao de parˆametros de buffer pool com aprendizado de m´aquina em ambientes n˜ao transacionais,” in Anais do XXVI Simp´osio em Sistemas Computacionais de Alto Desempenho (SSCAD). Porto Alegre, RS, Brasil: SBC, 2025, pp. 73–84.

G. Casella and R. L. Berger, Statistical Inference, 2nd ed. Duxbury, 2002.

R. Tibshirani, “Regression Shrinkage and Selection via the Lasso,” J. R. Stat. Soc. Ser. B Methodol., vol. 58, no. 1, pp. 267–288, 1996.

X. Geng et al., “Research on Database Parameters Tuning Method Based on Embedded Device,” J. Phys.: Conf. Ser., vol. 1873, no. 1, p. 012059, 2021.

N. Trendafilov and K. Hirose, “Exploratory Factor Analysis,” in Int. Encycl. Educ., 4th ed. Amsterdam: Elsevier, 2023, pp. 600–606.

A. Yong and J. Pearce, “A Beginner’s Guide to Factor Analysis: Focusing on EFA,” Tutorials Quant. Methods Psychol., vol. 9, no. 2, pp. 79–94, 2013.

H. F. Kaiser, “An Index of Factorial Simplicity,” Psychometrika, vol. 39, no. 1, pp. 31–36, 1974.

R. B. Cattell, “The Scree Test for the Number of Factors,” Multivariate Behavioral Research, vol. 1,

no. 2, pp. 245–276, 1966.

J. F. Hair et al., Multivariate Data Analysis, 7th ed. Pearson, 2010.

B. Shahriari et al., “Taking the Human Out of the Loop: A Review of Bayesian Optimization,” Proc.IEEE, vol. 104, no. 1, pp. 148–175, 2016.

C. E. Rasmussen and C. K. I. Williams, Gaussian Processes for Machine Learning. MIT Press, 2005.

G. Li et al., “Machine Learning for Databases,” Proc. VLDB Endow., vol. 14, no. 12, pp. 3190–3193, 2021.

S. Duan et al., “iTuned: a Tool for Automated DBMS Parameter Tuning,” Proc. VLDB Endow., vol. 2, no. 1, pp. 1241–1252, 2009.

J. Ansel et al., “OpenTuner: An Extensible Framework for Program Autotuning,” in Proc. PACT, 2014.

X. Zhang et al., “An End-to-End Automatic Cloud Database Tuning System Using Deep Reinforcement Learning,” in Proc. SIGMOD, 2019, pp. 415–432.

M. Kunjir and S. Babu, “Black or White? Design Choices and Trade-offs for Autonomous Database Tuning,” Proc. VLDB Endow., vol. 13, no. 12, pp. 3503–3515, 2020.

X. Zhang et al., “ResTune: Resource-Oriented Tuning Boosted by Meta-Learning for Cloud Databases,”in Proc. SIGMOD, 2021, pp. 2102–2114.

S. Cereda et al., “CGPTuner: a Contextual Gaussian Process Predictor for Optimizing Database Configurations,” Proc. VLDB Endow., vol. 14, no. 11, pp. 2501–2514, 2021.

J. Wang et al., “UDO: Universal Database Optimization using Reinforcement Learning,” Proc. VLDB Endow., vol. 14, no. 11, pp. 2102–2114, 2021.

J. R. Gunasekaran et al., “Utilizing GMM and DNN for DBMS Configuration Tuning,” in Proc. ICDE,2023.

A. Giannankouris et al., “?-Tune: Leveraging Large Language Models for DBMS Parameter Tuning,”Proc. VLDB Endow., 2024.

Z. Chen et al., “Centrum: A Unified Framework for Cross-DBMS Configuration Tuning,” in Proc.

SIGMOD, 2025.

Y. Li et al., “ADWTune: Adaptive Database Workload Tuning via DBSCAN and DDPG,” in Proc.

ICDE, 2025.

B. Cai et al., “HUNTER: An Online Cloud Database Hybrid Tuning System for Personalized Requirements,” in Proc. SIGMOD, 2022.

X. Zhang et al., “Facilitating Database Tuning with Hyper-Parameter Optimization: A Comprehensive

Experimental Evaluation,” Proc. VLDB Endow., vol. 15, no. 9, pp. 1808–1821, 2022.

F. Martinez-Plumed et al., “CRISP-DM Twenty Years Later: From Data Mining Processes to Data

Science Trajectories,” IEEE Trans. Knowl. Data Eng., vol. 33, no. 8, pp. 3048–3061, 2021.

P. Chapman, J. Clinton, R. Kerber, T. Khabaza, T. Reinartz, C. Shearer, and R. Wirth, CRISP-DM

0: Step-by-step Data Mining Guide, CRISP-DM Consortium, Brussels, 2000, sPSS Inc.

R. D. Peng, “Reproducible research in computational science,” Science, vol. 334, no. 6060, pp. 1226–

, 2011.

P. J. Rousseeuw, “Silhouettes: A Graphical Aid to the Interpretation and Validation of Cluster Analysis,” J. Comput. Appl. Math., vol. 20, pp. 53–65, 1987.

Scikit-learn, “StandardScaler,” https://scikit-learn.org/stable/modules/generated/sklearn. preprocessing.StandardScaler.html, 2024.

IBM, Ed., IBM Db2 13 for z/OS and More. IBM, 2023.

R. Turner, D. Eriksson, M. McCourt, J. Kiili, E. Laaksonen, Z. Xu, and I. Guyon, “Bayesian optimization is superior to random search for machine learning hyperparameter tuning: Analysis of the black-box optimization challenge 2020,” in NeurIPS 2020 Competition and Demonstration Track, ser. Proceedings of Machine Learning Research, vol. 133. PMLR, 2021, pp. 3–26.

A. Pavlo, G. Angulo, J. Arulraj, H. Lin, J. Lin, L. Ma, P. Menon, T. Mowry, S. Perron, I. Quah, A. Ray, S. Suri, J. Wang, and S. Wu, “Self-driving database management systems,” in Proceedings of the 8th Biennial Conference on Innovative Data Systems Research (CIDR), 2017.

T. Kraska, A. Beutel, E. H. Chi, J. Dean, and N. Polyzotis, “The case for learned database systems,” in Proceedings of the 8th Biennial Conference on Innovative Data Systems Research (CIDR), 2018.

Downloads

Published

2026-08-06