Machine Learning-based Query Performance Optimization in Large-Scale IoT Databases
view PDF
view PDF

How to Cite

Karunarathne, Lakmali. 2026. “Machine Learning-Based Query Performance Optimization in Large-Scale IoT Databases”. Journal of Information Technology and Digital World 8 (4): 287-307. https://doi.org/10.36548/jitdw.2026.4.001.

Keywords

Internet of Things (IoT)
Query Optimization
Machine Learning
CatBoost
Large-Scale Databases
IoT Telemetry Data

Abstract

The rapid growth of Internet of Things (IoT) deployments generates continuous, high-volume telemetry that strains query processing in large-scale databases, and static rule-based or cost-based optimizers do not adapt to the resulting workload variability. This research proposes a machine learning-based query performance optimization framework that identifies high-priority telemetry records from predicted operating conditions and formalizes a mathematically defined priority-scoring mechanism that a database system can use to allocate indexing, caching, and execution resources adaptively. The framework was developed and validated using the publicly available Environmental Sensor Telemetry Dataset (405,184 raw IoT records collected over eight days from three Raspberry Pi-based sensor nodes; cleaned to 385,452 records), with a binary hazard label derived from carbon-monoxide (CO) concentration against an EPA-informed threshold. Seven candidate classifiers, K-Nearest Neighbors, Logistic Regression, an Artificial Neural Network, Gaussian Naive Bayes, Linear Discriminant Analysis, a Support Vector Machine, and CatBoost, were trained on a stratified-representative 10,000-record sample (8,000 train / 2,000 test, class balance within 0.5 percentage points of the full population) and compared on accuracy, precision, recall, F1-score, prediction variance, and mean absolute error. CatBoost was selected as the proposed model, achieving 99.9% accuracy (95% CI: 99.76–100.00%), compared with 92.1–95.0% for the next-best classifiers. A correlation analysis quantified strong collinearity among CO, LPG, and smoke channels (r ≈ 1.00) and a moderate negative association with humidity (r ≈ −0.70 to −0.71), which is shown to reduce the effective dimensionality of the sensor feature space and motivates a composite-index recommendation for the underlying time-series store. The trained classifier's output probability is used to define a formal query/index priority score and a resource-allocation objective for adaptive optimization. Results demonstrate that the proposed framework provides a statistically validated, computationally efficient basis for adaptive IoT query optimization; full deployment-scale latency benchmarking against a live database engine is identified as future work.

References

  1. Zou, Benyuan, Jinguo You, Quankun Wang, Xinxian Wen, and Lianyin Jia. ”Survey on Learnable Databases: A Machine Learning Perspective.” Big Data Research 2022, vol. 27: 100304. https://doi.org/10.1016/j.bdr.2021.100304
  2. Milicevic, Bogdan, and Zoran Babovic. ”A Systematic Review of Deep Learning Applications in Database Query Execution.” Journal of Big Data 2024, vol. 11, no. 1: 173. https://doi.org/10.1186/s40537-024-01025-1
  3. Marcus, Ryan, and Olga Papaemmanouil. ”Deep Reinforcement Learning for Join Order Enumeration.” In Proceedings of the First International Workshop on Exploiting Artificial Intelligence Techniques for Data Management 2018, 1-4. https://doi.org/10.1145/3211954. 3211957
  4. Qiao, Shao-Jie, Han-Lin Fan, Nan Han, Lan Du, Yu-Han Peng, Rong-Min Tang, and Xiao Qin. ”Learning Database Optimization Techniques: The State-of-the-Art and Prospects.” Frontiers of Computer Science 2025, vol. 19, no. 12: 1912612. https://doi.org/10.1007/s11704-025-41116-7
  5. Li, Shancang, Li Da Xu, and Shanshan Zhao. ”The Internet of Things: A Survey.” Information Systems Frontiers 2015, vol. 17, no. 2: 243-259. https://doi.org/10.1007/s107 96-014-9492-7
  6. Shi, Weisong, Jie Cao, Quan Zhang, Youhuizi Li, and Lanyu Xu. ”Edge Computing: Vision and Challenges.” IEEE Internet of Things Journal 2016, vol. 3, no. 5: 637-646. https://doi.org/10.1109/JIOT.2016.2579198
  7. Leis, Viktor, Bernhard Radke, Andrey Gubichev, Atanas Mirchev, Peter Boncz, Alfons Kemper, and Thomas Neumann. ”Query Optimization Through the Looking Glass, and What We Found Running the Join Order Benchmark.” The VLDB Journal 2018, vol. 27, no. 5: 643-668. https://doi.org/10.1007/s00778-017-0480-7
  8. Chen, Xingguang, Rong Zhu, Bolin Ding, Sibo Wang, and Jingren Zhou. ”Lero: Applying Learning-to-Rank in Query Optimizer: X. Chen ” The VLDB Journal 2024, vol. 33, no. 5: 1307-1331. https://doi.org/10.1007/s00778-024-00850-3
  9. Zhou, Xuanhe, Chengliang Chai, Guoliang Li, and Ji Sun. ”Database Meets Artificial Intelligence: A Survey.” IEEE Transactions on Knowledge and Data Engineering 2020, vol. 34, no. 3: 1096-1116. https://doi.org/10.1109/TKDE.2020.2994641
  10. Marcus, Ryan, Parimarjan Negi, Hongzi Mao, Chi Zhang, Mohammad Alizadeh, Tim Kraska, Olga Papaemmanouil, and Nesime Tatbul. ”Neo: A Learned Query Optimizer. ” Proceedings of the VLDB Endowment 2019, vol. 12, no. 11: 1705–1718. https://doi.org/10.14778/3342263.3342644
  11. Valavala, Mounicasri, and Wasim Alhamdani. ”Automatic Database Index Tuning Using Machine Learning.” In 2021 6th International Conference on Inventive Computation Technologies (ICICT), IEEE, 2021, 523-530. https://doi.org/10.1109/ICICT50816.2021. 9358646
  12. Bharati, Subrato, and Prajoy Podder. ”Machine and Deep Learning for IoT Security and Privacy: Applications, Challenges, and Future Directions.” Security and Communication Networks 2022, vol. 2022, no. 1: 8951961. https://doi.org/10.1155/2022/8951961
  13. Stafford, G. (2020). Environmental Sensor Telemetry Data [Data set]. Kaggle. https://www.kaggle.com/datasets/garystafford/environmental-sensor-data-132k
  14. U.S. Environmental Protection Agency. (2023). Carbon Monoxide's Impact on Indoor Air Quality. https://www.epa.gov/indoor-air-quality-iaq/carbon-monoxides-impact-indoor-air-quality
  15. Dorogush, Anna Veronika, Vasily Ershov, and Andrey Gulin. ”CatBoost: Gradient Boosting with Categorical Features Support.” arXiv Preprint arXiv:1810.11363 (2018). https://doi.org/10.48550/arXiv.1810.11363
  16. Pedregosa, Fabian, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel ”Scikit-Learn: Machine Learning in Python.” the Journal of Machine Learning Research 2011, vol. 12: 2825-2830.
  17. Abadi, Martín, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin. ”{TensorFlow}: A System for {Large-Scale} Machine Learning.” In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16), 2016, 265-283.