国际妇产科学杂志 ›› 2026, Vol. 53 ›› Issue (3): 261-266.doi: 10.12280/gjfckx.20251231

• 产科生理及产科疾病:论著 • 上一篇    下一篇

多种机器学习算法构建新生儿低血糖症预测模型

赵维丽, 胡丽燕()   

  1. 030000 山西医科大学(赵维丽);山西省儿童医院,山西省妇幼保健院妇产科(胡丽燕)
  • 收稿日期:2025-11-06 出版日期:2026-06-15 发布日期:2026-07-06
  • 通讯作者: 胡丽燕 E-mail:15935136709@139.com

Multiple Machine Learning Algorithms to Build A Prediction Model for Neonatal Hypoglycemia

ZHAO Wei-li, HU Li-yan()   

  1. Shanxi Medical University, Taiyuan 030000, China (ZHAO Wei-li); Department of Obstetrics and Gynecology, Shanxi Children's Hospital, Shanxi Women and Children Hospital, Taiyuan 030000, China (HU Li-yan)
  • Received:2025-11-06 Published:2026-06-15 Online:2026-07-06
  • Contact: HU Li-yan E-mail:15935136709@139.com

摘要:

目的:比较多种机器学习算法在构建新生儿低血糖症(neonatal hypoglycemia,NH)预测模型中的性能,以期早期识别可能分娩低血糖新生儿的高危孕妇。方法:回顾性分析2021年1月—2025年7月山西省儿童医院收治的286例孕妇,将其按照7∶3比例随机分成训练集(200例)和测试集(86例),主要结局指标为是否发生NH。特征筛选采用单因素分析、LASSO回归和Boruta算法,基于所筛选出的特征构建预测模型,应用Bootstrap交叉验证进行模型验证。结果:通过单因素分析、LASSO回归和Boruta特征筛选发现,妊娠前体质量指数、妊娠期增重、分娩方式、妊娠期糖尿病、妊娠期高血压疾病、促胎肺成熟治疗、新生儿出生体质量和早产是NH的影响因素(P<0.05)。在训练集中,随机森林、决策树、支持向量机和Logistic回归模型的AUC分别为0.967、0.943、0.800和0.955,在测试集中则分别为0.921、0.860、0.882和0.910,综合比较训练集和测试集的性能指标,随机森林模型的性能最优。结论:随机森林模型可用来预测NH发生风险,具有较高的临床应用价值。

关键词: 低血糖症, 婴儿,新生, 机器学习, 预测模型, 妊娠期管理

Abstract:

Objective: To compare the performance of multiple machine learning algorithms in developing predictive models for neonatal hypoglycemia (NH), with the aim of early identification of high-risk pregnant women likely to deliver neonates with hypoglycemia. Methods: A retrospective analysis was conducted on 286 pregnant women admitted to Shanxi Children's Hospital from January 2021 to July 2025. Participants were randomly allocated into a training set (n=200) and a test set (n=86) at a 7:3 ratio. The primary outcome measure was the occurrence of NH. Feature selection was performed using univariate analysis, LASSO regression, and the Boruta algorithm. Predictive models were constructed based on the selected features, and Bootstrap cross-validation was applied for model validation. Results: Through univariate analysis, LASSO regression, and Boruta feature selection, pre-pregnancy body mass index, gestational weight gain,mode of delivery,gestational diabetes mellitus, hypertensive disorders of pregnancy, antenatal corticosteroid therapy for fetal lung maturation, neonatal birth weight,and preterm birth were identified as influencing factors for NH (P<0.05). In the training set, the AUC values for the random forest, decision tree, support vector machine, and Logistic regression models were 0.967, 0.943, 0.800, and 0.955, respectively; the corresponding values in the test set were 0.921, 0.860, 0.882, and 0.910, respectively. A comprehensive comparison of performance metrics across both the training and test sets demonstrated that the random forest model achieved superior overall performance. Conclusions: The random forest model can be applied to predict the risk of NH occurrence and demonstrates considerable clinical utility.

Key words: Hypoglycemia, Infant, newborn, Machine learning, Predictive model, Pregnancy management