常规血液检查优于CRP用于区分儿童细菌与病毒感染
摘要
本研究表明,使用全血细胞计数(CBC)数据的机器学习模型在区分儿科患者的细菌和病毒感染方面优于单独的CRP。
arXiv:2609.21332v1 Announce Type: new
Abstract: Acute infectious diseases are among the leading causes of medical consultations and hospitalizations in children worldwide. These infections are predominantly caused by viruses or bacteria, yet differentiating between the two remains a common clinical challenge. As a result, pediatricians often default to the safer option of prescribing antibiotics contributing to the growing problem of antimicrobial resistance. The objective is to assess the additional predictive value of CBC towards determining the current infection. This retrospective study used data from 906 pediatric patients aged between 2 and 14 years who were tested positive either for viral or bacterial infection between 2022 and 2026. Inclusion criteria further required availability of CBC results and CRP level measurements. These laboratory parameters as well as age were used as input features for several supervised classification models. Model performance was evaluated using AUC, sensitivity and specificity. The best performing model is XGBoost, which included all features, achieving out of-sample performance of AUC of 81.7% and sensitivity of 70.8%, specificity of 79.2%. All trained models outperform a CRP-based only decision-rule model in terms of AUC. We suggest that the decision to prescribe antibiotics should be based on a number of factors, including but not limited to CBC, some of which are not currently incorporated into routine practice.
查看缓存全文
缓存时间: 2026/09/21 09:31
# Routine Blood Tests Outperform CRP for Distinguishing Bacterial From Viral Infection in Children
Source: [https://arxiv.org/html/2609.21332](https://arxiv.org/html/2609.21332)
Zhecho MitevAffiliation:Vector Labs, Sofia, BulgariaDjuna Chinareva\-KlimentovaAffiliation:Vector Labs, Sofia, BulgariaSvetoslav IvanovAffiliation:Vector Labs, Sofia, BulgariaGeorgi NalbantovAffiliation:Vector Labs, Sofia, BulgariaDimitar MitevAffiliation:Zdraveto Hospital, Sofia, Bulgaria
###### Abstract
Background:Acute infectious diseases are among the leading causes of medical consultations and hospitalizations in children worldwide\. These infections are predominantly caused by viruses or bacteria, yet differentiating between the two remains a common clinical challenge\. As a result, pediatricians often default to the safer option of prescribing antibiotics contributing to the growing problem of antimicrobial resistance\. Complete Blood Count \(CBC\) and C\-reactive protein \(CRP\) tests are widely available and represent a standardized source of information reflecting the host immune response to infection\. While the usage of blood tests to determine an infection is more often explored among adults, little information is available on the predictive capabilities of these tests among children\. The objective is to assess the additional predictive value of CBC towards determining the current infection\.
Methods:This retrospective study used data from 906 pediatric patients aged 2–14 years who were evaluated for suspected infection at a pediatric hospital in Bulgaria between 2022 and 2026, focusing on the post\-covid period\. These patients were tested positive either for viral or bacterial infection\. Inclusion criteria further required availability of CBC results and CRP level measurements\. These laboratory parameters as well as age were used as input features for several supervised classification models\. The commonly used method of using CRP alone to differentiate between viral and bacterial infection was compared to logistic regression and Gradient Boosting methods\. Model performance was evaluated using Area Under the Curve \(AUC\), sensitivity and specificity\.
Results:The best performing model is XGBoost, which included all features, achieving out\-of\-sample performance of AUC = 81\.7% and sensitivity = 70\.8%, specificity = 79\.2%\. On the other hand, the XGBoost model without CRP performs slightly worse with 1 percent less in AUC and specificity, but similar in sensitivity\. All trained models outperform a CRP\-based only decision\-rule model \(later mentioned as CRP baseline model\) in terms of AUC\.
Conclusions:Both logistic regression \(linear\) and XGBoost \(non\-linear\) models better distinguish between viral and bacterial infection compared to the CRP baseline model\. We suggest that the decision to prescribe antibiotics should be based on a number of factors, including but not limited to CBC, some of which are not currently incorporated into routine practice\. Factors such as white blood count \(WBC\), lymphocyte count \(LYM\) and monocyte percentage \(MON%\) appear to be more informative than the value of CRP for the distinction according to our best XGBoost model\.
## 1Introduction
Despite significant advances in medicine and technology, distinguishing between viral and bacterial infections remains a difficult task\. Symptoms such as fever, muscle pain, and generalized weakness may be caused by either type, while the required treatments differ substantially\. Early and accurate determination of the infection type is crucial for prescribing appropriate therapy—bacterial infections usually require antibiotic treatment, whereas antibiotics can be ineffective against viral infections\.[1](https://arxiv.org/html/2609.21332#bib.bib1)
Antibiotics overprescription is a well\-known problem, still researched actively\. Wrong usage of antibiotics can lead to development of bacterial resistance to antibiotics\.[2](https://arxiv.org/html/2609.21332#bib.bib2)
The doctors’ practice tends to be different among the different countries\. A pediatric study from 5 European countries[3](https://arxiv.org/html/2609.21332#bib.bib3)examines the correctness of antibiotics prescription\. The researchers found that the incorrect prescription of antibiotics reached a maximum of 64\.7% in Bratislava, Slovakia\. Moreover, according to the study, in western countries more doctors prescribed wrong antibiotics as a first line medication in case of Tonsillitis or Otitis media\. A more recent study conducted in the United States[4](https://arxiv.org/html/2609.21332#bib.bib4)indicates that at least 30% of prescribed antibiotics are unnecessary, despite the implementation of programs aimed at reducing inappropriate antibiotic use\.[5](https://arxiv.org/html/2609.21332#bib.bib5)
In case of symptoms, hinting bacterial or viral infection, doctors might use external factors to decide whether to prescribe or not antibiotics\. Zaykova et al\. \(2024\)[6](https://arxiv.org/html/2609.21332#bib.bib6)found evidence for different approaches among different types of physicians\. Counterintuitively, decisions also varied by weekday versus weekend/holiday periods\.
A common practice among doctors is to use CRP \(C\-reactive Protein\) tests as an indicator of the existence of bacterial infection and therefore as a decision rule whether to prescribe antibiotics or not\. Medical papers and laboratories define low CRP level to indicate viral infection, while high value – bacterial\. The range of 10–40 mg/L is considered critical, since the infection can be either bacterial or viral\.[7](https://arxiv.org/html/2609.21332#bib.bib7)However, there are 2 main reasons why one can not use CRP as the only identification for the type of infection\. First, when CRP is below the critical range, bacterial infection is still possible especially if the infection started less than 48 hours before the test\.[8](https://arxiv.org/html/2609.21332#bib.bib8)Second, adenovirus, influenza and SARS\-CoV\-2 could cause CRP to rise above 40 mg/L, therefore higher levels could not exclude the presence of a virus\.
Various biomarkers have been investigated for their ability to differentiate infection types in adults\. Multiple authors suggest that combining multiple biomarkers can improve the accuracy of the prediction outcome\.[9](https://arxiv.org/html/2609.21332#bib.bib9),[10](https://arxiv.org/html/2609.21332#bib.bib10)It is also believed that some key blood indicators vary in terms of gender and age\.[11](https://arxiv.org/html/2609.21332#bib.bib11),[12](https://arxiv.org/html/2609.21332#bib.bib12)Therefore such factors must also be taken into account for a more accurate classification\. More recently, machine learning approaches have demonstrated improved diagnostic accuracy by integrating multiple blood parameters\. Gunčar et al\. \(2024\)[7](https://arxiv.org/html/2609.21332#bib.bib7)developed a model using 16 routine blood test results along with CRP, achieving 82\.2% accuracy in distinguishing bacterial from viral infections in adults—outperforming CRP\-based decision rules alone\.
In pediatric populations, distinguishing infection etiology presents additional challenges due to the relatively immature immune system, different epidemiology of infections, and difficulties in specimen collection compared to adults\.[13](https://arxiv.org/html/2609.21332#bib.bib13)Several studies have focused on febrile children, particularly infants and preschool\-aged children\.[14](https://arxiv.org/html/2609.21332#bib.bib14)
Other studies examining infection differentiation in children have focused on novel biomarkers requiring specialized assays,[15](https://arxiv.org/html/2609.21332#bib.bib15),[16](https://arxiv.org/html/2609.21332#bib.bib16)while the combined diagnostic potential of routinely available tests—complete blood count and CRP—remains underexplored in pediatric populations\. In our research, we address this gap by evaluating the utility of CBC parameters combined with CRP for distinguishing viral from bacterial infections in children aged 2–14 years\. Our goal is to evaluate what is the contribution of CRP as a feature in such a classification model and how much better a classification model performs compared to the simple decision rule using CRP only\.
## 2Materials and methods
### 2\.1Patient Population
Between 2022 and 2026, a total of 5347 pediatric patients between 2 and 14 years with suspected infection have been tested for either viral, bacterial or both infections in hospital Zdraveto\. We applied the following eligibility criteria: proven bacteria or viral infection; available CBC and CRP tests\. As a result, in 3617 patients an infection was identified\. Since our focus is on differentiating viral or bacterial infections, only observations with confirmed infections from any of the both kinds of tests were kept\. Five patients have been positive for both viral and bacterial infection, thus due to ambiguity they have been removed\. Out of the remaining 3612 patients only a cohort of 956 patients have no missing values for all blood tests \(CBC, CRP\)\. 50 patients, which are the most recent, are taken out and used later for held\-out validation\. Therefore our final cohort consists of 906 patients\. Out of these, 424 are with bacterial infection and 482 are with viral infection\.
Descriptive statistics of all features are given in Table[1](https://arxiv.org/html/2609.21332#S2.T1)\.
Table 1:Descriptive statistics of model parameters in the whole dataset and split by output \(Virus versus Bacteria\)Values are mean±\\pmSD\. Abbreviations:nn, number; SD, standard deviation\.
### 2\.2Data processing and labelling
Several viral and bacterial diagnostic tests were performed\. Viral testing included assays for SARS\-CoV\-2, influenza A and B, respiratory syncytial virus \(RSV\), and adenovirus\. Due to the fact that test results were in text form, data preprocessing was required in order to define the final virus label as positive or negative\.
Some of the patients have been tested for bacteria through quick tests for streptococci and/or nasal or throat secretions\. For the latter, additional data preprocessing has been made since results indicate the presence and the type of bacteria as well as its intensity in text form\. We exclude the ambiguous cases and leave only those with no bacterial finding or with proven presence of such\.
Our outcome variable is defined as follows: “1”: presence of bacterial infection and “0”: presence of viral infection\.
### 2\.3Statistical analysis
Using the dataset described above, we formulated the prediction task as a binary classification problem\. Two main well\-known supervised machine learning approaches were considered and compared – logistic regression and XGBoost\.[17](https://arxiv.org/html/2609.21332#bib.bib17)For each of which a hyperparameter search was performed in order to optimize our chosen performance criteria – Area under the curve \(AUC\)\. For the logistic regression all features have been min\-max scaled\. Along with the four mentioned models, a univariate logistic regression analysis was also performed\.
Apart from age and CRP, the standard parameters from a CBC have been used as input variables:
> LYM– Lymphocytes \(absolute count\),LYM%– Lymphocyte percentage,MON– Monocytes \(absolute count\),MON%– Monocyte percentage,NEU– Neutrophils \(absolute count\),NEU%– Neutrophil percentage,RBC– Red Blood Cell count,HGB– Hemoglobin,HCT– Hematocrit,MCV– Mean Corpuscular Volume,MCH– Mean Corpuscular Hemoglobin,MCHC– Mean Corpuscular Hemoglobin Concentration,RDW– Red Cell Distribution Width,WBC– White Blood Cell count,PLT– Platelet count,MPV– Mean Platelet Volume,PCT– Plateletcrit,PDW– Platelet Distribution Width\.
Both logistic regression and XGBoost models have hyperparameters which have to be selected for a final model\. For logistic regression we use L2\-norm penalization whose shrinkage parameter is the only hyperparameter that must be tuned\. For XGBoost the following parameters were tuned: maximum tree depth \(max\_depth\), number of boosting rounds \(n\_estimators\), feature subsampling ratio \(colsample\_bytree\), minimum child weight \(min\_child\_weight\), minimum loss reduction required to make a split \(gamma\),fraction of observations used for each tree \(subsample\) and L2 regularization \(reg\_lambda\)\.
A standard procedure for hyperparameter selection is the 5\-fold cross\-validation with stratification selection\. For each hyperparameter combination we obtain cross\-validation out\-of\-sample AUC\. The same 5\-fold split was used for all models to ensure fair comparison\. Final models were built with the hyperparameters which produced the highest cross\-validation AUC values\.
The chosen models were also compared with a CRP Baseline Model using only CRP as a predictor\. Similar to other authors[7](https://arxiv.org/html/2609.21332#bib.bib7)we chose an optimal cutoff value for the CRP Baseline Model based on the highest accuracy from the whole dataset\. We note that this is in\-sample performance, therefore overoptimistic\. In our case the cut\-off is 22 mg/L\. According to this rule all cases below 22 mg/L were considered viral infection, and all equal to 22 and above – bacterial\.
Additionally, SHAP \(SHapley Additive exPlanations\) provides a unified method for feature importance and is widely used for model interpretability\.[18](https://arxiv.org/html/2609.21332#bib.bib18)We computed SHAP values on the final model trained on the full dataset\.
## 3Results
First, 12 variables were found to be statistically significant \(all except for RBC, HGB, HCT, MCV, RDW, MPV, MCH and PDW\) in an univariate non\-penalized logistic regression analysis \(Table[S1](https://arxiv.org/html/2609.21332#Ax1.T1), Supplementary Materials\)\. Given the large sample size, thepp\-values should be interpreted with some caution, as they may be inflated in significance\.
In the multivariate case, our results were obtained using 5\-fold cross validation with stratification on the earlier described dataset with 906 observations\. We compare 2 machine learning algorithms, each with 2 models – one with CRP as a feature and one without CRP using AUC, sensitivity and specificity as performance metrics\. To assess the statistical performance of the best models and the CRP baseline model 95% confidence intervals were estimated by bootstrapping the out\-of\-fold predictions of the final selected models using 1000 bootstrap resamples\.[19](https://arxiv.org/html/2609.21332#bib.bib19),[20](https://arxiv.org/html/2609.21332#bib.bib20)
Our best performing model is XGBoost with CRP as an input feature \(see Table[2](https://arxiv.org/html/2609.21332#S3.T2), Figure[2](https://arxiv.org/html/2609.21332#S3.F2)\)\. It achieves an area under the curve \(AUC\) of 81\.7%, with a sensitivity of 70\.8% and a specificity of 79\.2%\. The XGBoost model trained without CRP as a feature performs slightly worse in terms of AUC \(AUC – 80\.8%, sensitivity – 71%, specificity – 78%\)\. As shown in Table[2](https://arxiv.org/html/2609.21332#S3.T2)for both cases XGBoost performed better than logistic regression in sensitivity and in AUC\. However, the specificity of both logistic regression models is higher compared to the XGBoost models on the account of sensitivity of∼67%\{\\sim\}67\\%\.
As shown in Figure[1](https://arxiv.org/html/2609.21332#S3.F1), the SHAP analysis identified white blood count \(WBC\), lymphocyte count \(LYM\) and monocyte percentage \(MON%\) as the most influential features of the best mode\. Higher WBC and LYM values generally increased the probability of bacterial infection, while lower WBC and LYM values shifted predictions toward viral infection\. Other laboratory parameters had comparatively smaller effects, with most SHAP values clustered near zero \(e\.g\. MCV, HGB, MPW, RDW, NEU%, PDW\)\.
Considering AUC, all our models perform substantially better than the CRP Baseline model at all thresholds, which only has an AUC of 57\.4% \(Figure[2](https://arxiv.org/html/2609.21332#S3.F2)\)\.
Finally, we compared our XGBoost model with CRP as a feature to a CRP Baseline model with a cut\-off 22 mg/L\. The CRP Baseline model achieved a sensitivity of 30\.9% and a specificity of 85\.6%, being outperformed significantly by all four trained models\.
To validate our findings, we have tested all models on a held\-out set consisting of 50 patients \(16 viral, 34 bacterial\)\. The pattern of results remained consistent \(see Table[S2](https://arxiv.org/html/2609.21332#Ax1.T2), Supplementary Materials\)\. Even on unseen data the trained models add value compared to the CRP Baseline model, which has an AUC barely above the random guess \(52\.8%\)\. Given the small size and pronounced class imbalance of this held\-out set \(68% bacterial vs\. 46\.8% in training\), class\-specific metrics such as specificity and sensitivity should be interpreted with caution\.
Table 2:Model performance with 95% confidence intervals using bootstrap resampling of the out\-of\-fold predictions \(1000 iterations\)\.Abbreviations: w/ CRP – model including CRP as a feature; wo/ CRP – model excluding CRP as a feature; LR – Logistic Regression model\. XGBoost w/ CRP is the best performing model overall \(AUC = 81\.7% and sensitivity = 70\.8%, specificity = 79\.2%\)\.
Figure 1:SHAP values for the whole dataset on the final model trained\. Each dot represents an observation and the color and its intensity represent the effect of that feature on the dot\. Effects to the right of 0\.0 are positive, and to the left – negative\. WBC, LYM, MON%, CRP, PCT and age have the highest impact, and PDW, NEU%, RDW, MPV, HGB and MCV – the lowest\.Figure 2:Receiver operating characteristic \(ROC\) curves on all trained models using out\-of\-fold predictions from 5\-fold cross validation with stratification and CRP Baseline model on the whole dataset\. AUC\_CRP\_baseline – Area Under the Curve for baseline model with CRP alone; AUC\_XGB\_crp – Area Under the Curve for XGBoost model with CRP as variable; AUC\_XGB\_wo\_crp – Area Under the Curve for XGBoost model without CRP as variable; AUC\_LR\_crp – Area Under the Curve for logistic regression model with CRP as variable; AUC\_LR\_wo\_crp – Area Under the Curve for logistic regression model without CRP as variable\.
## 4Discussion
In this research, we have employed a statistical approach capable of differentiating between viral and bacterial infections in children\. Our results suggest that the combination of multiple blood indicators outperforms the predictions produced by the CRP baseline\.
According to our best performing model, the most important input features are WBC, LYM and MON%\. All these features were shown to be statistically significant in the univariate logistic regression analysis as well \(Table[S1](https://arxiv.org/html/2609.21332#Ax1.T1), Supplementary Materials\)\. Other researchers show that a virus such as SARS\-CoV\-2 is leading to a decrease in the absolute lymphocyte count\.[21](https://arxiv.org/html/2609.21332#bib.bib21)Our findings indicate that lymphopenia may serve as a broader indicator of viral infection beyond COVID\-19 as the period we are covering is post\-pandemic \(after 2022\)\.
Elevated WBC levels are commonly associated with bacterial infections, reflecting the immune system’s response through increased production of these cells to fight the invading pathogens\.[22](https://arxiv.org/html/2609.21332#bib.bib22)Our research confirms this observation throughout our study population\. Our model places WBC as the most important indicator for the differentiation of virus vs bacteria\. This also aligns with other statistical studies, confirming WBC to be an important indicator for discovering bacterial infections\.[9](https://arxiv.org/html/2609.21332#bib.bib9)
It is important to note that none of these biomarkers alone are sufficient to make a reliable classification of infection\. Our results indicate that a sophisticated combination of multiple factors is required to achieve optimal results\.
C\-reactive protein is widely regarded as an important indicator of bacterial infection\.[23](https://arxiv.org/html/2609.21332#bib.bib23)However, our study suggests that CRP alone has limited discriminative ability for differentiating viral from bacterial infections\. This is supported by the histograms of patients grouped by infection type \(viral vs\. bacterial\), which show substantial overlap and similar means between the two classes \(Figure[S1](https://arxiv.org/html/2609.21332#Ax1.F1), Supplementary Material\)\. Despite selecting the optimal threshold of 22 mg/L for the CRP baseline model it produces a specificity slightly above our best performing model \(85\.6% vs 79\.2%\), and a sensitivity much lower than XGBoost \(30\.9% vs 70\.8%\)\.
Even when combined with complete blood count \(CBC\) parameters, CRP contributed only marginally to model performance\. This was reflected in both machine learning algorithms \(XGBoost and logistic regression\), where inclusion of CRP resulted in only a modest improvement in AUC \(1 percentage point for XGBoost and 0\.2 percentage points for logistic regression\)\.
Other authors developing machine learning\-based models have used similar models for adults[7](https://arxiv.org/html/2609.21332#bib.bib7)and very young children \(<<6 years old\)\.[14](https://arxiv.org/html/2609.21332#bib.bib14)They have utilized similar CBC parameters and CRP as input features\. However, they note that CRP is a very good indicator of the type of infection, while in our cohort of children, CRP is not crucial for the distinction\. This discrepancy may be attributable to differences in age distribution, timing of biomarker assessment, or clinical setting, as CRP kinetics and baseline inflammatory responses can vary across populations\.[12](https://arxiv.org/html/2609.21332#bib.bib12),[24](https://arxiv.org/html/2609.21332#bib.bib24)This is a limitation of our study since we have not analyzed the differences in the age distribution due to the lack of enough data per age group\.
The findings of this study may support specialists in making more informed decisions when evaluating children with suspected infection\. Our method of combining routine blood tests, such as complete blood count and CRP, provides a way to make distinction between viral and bacterial infections\. This may help pediatricians reduce the antibiotic prescriptions or recognise early the need for one, which could lead to more appropriate treatment overall\.
While this research showed significant results, it has several limitations\. First, all models were trained on data collected from a single hospital, which may limit the generalizability of our findings\. Second, the time elapsed between illness onset of the patients and the timing of laboratory examinations were not available\. The absence of this information prevents the model from better understanding of the stages of the illness\.
Future work could involve developing a multicentric study including observations from other hospitals, capturing other relationships between CBC parameters, CRP measurement and infection type\. Furthermore, we could compare our model output with the decisions that pediatricians make based on the same variables and check if the model has a better prediction accuracy compared to practitioners\.
## 5Conclusion
A classification model was developed for distinguishing between viral and bacterial infections in children\. The models include complete blood count and CRP as input features and were all found to outperform the CRP Baseline model\. Moreover, some CBC input features add more contribution to the overall predictive performance of the model compared to CRP\. WBC, LYM and MON% appear most informative to our model\. A larger amount of data from various hospitals is required to validate our findings and improve the model performance\. Nonetheless, the results of our study show that such a model has the potential to be used as a tool facilitating general practitioners and pediatricians’ in prescription decisions\.
## References
- 1Tanday S\. Resisting the use of antibiotics for viral infections\.Lancet Respir Med\.2016;4\(3\):179\.
- 2Llor C, Bjerrum L\. Antimicrobial resistance: risk associated with antibiotic overuse and initiatives to reduce the problem\.Ther Adv Drug Saf\.2014;5\(6\):229–241\.
- 3Sanz Emilio J, Hernandez Miguel A, Ratchina S, et al\. Prescribers’ indications for drugs in childhood: a survey of five European countries \(Spain, France, Bulgaria, Slovakia and Russia\)\.Acta Paediatr\.2005;94\(12\):1784–1790\. doi:[10\.1111/j\.1651\-2227\.2005\.tb01854\.x](https://doi.org/10.1111/j.1651-2227.2005.tb01854.x)
- 4Harris AM, Hicks LA, Qaseem A; High Value Care Task Force of the American College of Physicians and for the Centers for Disease Control and Prevention\. Appropriate antibiotic use for acute respiratory tract infection in adults: advice for high\-value care from the American College of Physicians and the Centers for Disease Control and Prevention\.Ann Intern Med\.2016;164\(6\):425–434\. doi:[10\.7326/M15\-1840](https://doi.org/10.7326/M15-1840)
- 5Fiore DC, Fettic LP, Wright SD, Ferrara BR\. Antibiotic overprescribing: still a major concern\.J Fam Pract\.2017;66\(12\):730–736\.
- 6Zaykova K, Nikolova SP, Pancheva R, Serbezova A\. Antibiotic prescribing practices to children among in\- and outpatient physicians in Bulgaria\.Acta Medica Bulgaria\.2024;51\(4\)\. doi:[10\.2478/amb\-2024\-0075](https://doi.org/10.2478/amb-2024-0075)
- 7Gunčar G, Kukar M, Smole T, et al\. Differentiating viral and bacterial infections: a machine learning model based on routine blood test values\.Heliyon\.2024;10\(8\):e29372\. doi:[10\.1016/j\.heliyon\.2024\.e29372](https://doi.org/10.1016/j.heliyon.2024.e29372)
- 8Markanday A\. Acute phase reactants in infections: evidence\-based review and a guide for clinicians\.Open Forum Infect Dis\.2015;2\(3\):ofv098\. doi:[10\.1093/ofid/ofv098](https://doi.org/10.1093/ofid/ofv098)
- 9Gille\-Johnson P, Hansson KE, Gårdlund B\. Clinical and laboratory variables identifying bacterial infection and bacteraemia in the emergency department\.Scand J Infect Dis\.2012;44\(10\):745–752\. doi:[10\.3109/00365548\.2012\.689846](https://doi.org/10.3109/00365548.2012.689846)
- 10Oved K, Cohen A, Boico O, et al\. A novel host\-proteome signature for distinguishing between acute bacterial and viral infections\.PLoS One\.2015;10\(3\):e0120012\. doi:[10\.1371/journal\.pone\.0120012](https://doi.org/10.1371/journal.pone.0120012)
- 11Adeli K, Raizman JE, Chen Y, et al\. Complex biological profile of hematologic markers across pediatric, adult, and geriatric ages: establishment of robust pediatric and adult reference intervals on the basis of the Canadian Health Measures Survey\.Clin Chem\.2015;61\(8\):1075–1086\. doi:[10\.1373/clinchem\.2015\.240531](https://doi.org/10.1373/clinchem.2015.240531)
- 12Doucoure MB, Wright JK, Criss AH, et al\. Normal clinical laboratory ranges by age and sex: relationship with complete blood count and chemistry parameters\.Clin Lab Sci\.2024;37\(2\)\.
- 13Tsao YT, Tsai YH, Liao WT, et al\. Differential markers of bacterial and viral infections in children for point\-of\-care testing\.Trends Mol Med\.2020;26\(12\):1118–1132\. doi:[10\.1016/j\.molmed\.2020\.09\.004](https://doi.org/10.1016/j.molmed.2020.09.004)
- 14Lee B, Chung HJ, Kang HM, Kim DK, Kwak YH\. Development and validation of machine learning\-driven prediction model for serious bacterial infection among febrile children in emergency departments\.PLoS One\.2022;17\(3\):e0265500\. doi:[10\.1371/journal\.pone\.0265500](https://doi.org/10.1371/journal.pone.0265500)
- 15Srugo I, Klein A, Stein M, et al\. Validation of a novel assay to distinguish bacterial and viral infections\.Pediatrics\.2017;140\(4\):e20163453\.
- 16van Houten CB, de Groot JAH, Klein A, et al\. A host\-protein based assay to differentiate between bacterial and viral infections in preschool children \(OPPORTUNITY\): a double\-blind, multicentre, validation study\.Lancet Infect Dis\.2017;17\(4\):431–440\.
- 17Chen T, Guestrin C\. XGBoost: a scalable tree boosting system\. In:Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining \(KDD ’16\)\.New York, NY: Association for Computing Machinery; 2016:785–794\. doi:[10\.1145/2939672\.2939785](https://doi.org/10.1145/2939672.2939785)
- 18Lundberg SM, Lee SI\. A unified approach to interpreting model predictions\. In:Advances in Neural Information Processing Systems 30 \(NIPS 2017\)\.Red Hook, NY: Curran Associates; 2017:4765–4774\.
- 19Davison AC, Hinkley DV\.Bootstrap Methods and Their Application\.Cambridge: Cambridge University Press; 1997\.
- 20Hastie T, Tibshirani R, Friedman J\.The Elements of Statistical Learning: Data Mining, Inference, and Prediction\.2nd ed\. Springer; 2009\.
- 21Huang C, Wang Y, Li X, et al\. Clinical features of patients infected with 2019 novel coronavirus in Wuhan, China\.The Lancet\.2020;395\(10223\):497–506\.
- 22Riley LK, Rupert J\. Evaluation of patients with leukocytosis\.Am Fam Physician\.2015;92\(11\):1004–1011\.
- 23Yo CH, Hsieh PS, Lee SH, et al\. Comparison of the test characteristics of procalcitonin to C\-reactive protein and leukocytosis for the detection of serious bacterial infections in children presenting with fever without source: a systematic review and meta\-analysis\.Ann Emerg Med\.2012;60\(5\):591–600\.
- 24Sproston NR, Ashworth JJ\. Role of C\-reactive protein at sites of inflammation and infection\.Front Immunol\.2018;9:754\.
## Supplementary Material
Table S1:Univariate logistic regression analysisAbbreviations: S\.E\., standard error\.
Table S2:Held\-out set results \(n=50n=50, Virus = 16, Bacteria = 34\)Figure S1:Distribution of log values of CRP grouped by type of infection – viral or bacterial\.相似文章
基于机器学习的PCR确诊衣原体检测前风险分层:患者报告数据与尿液生物标志物的应用
本研究评估了机器学习模型在利用无创患者报告数据和尿液生物标志物对沙眼衣原体感染进行检测前风险分层中的应用,展示了适中的预测性能以及两种数据类型的互补价值。
机器学习评估炎症生物标志物对老年西班牙裔成人队列认知障碍的预测价值
本文介绍了一种可解释的机器学习方法,使用朴素贝叶斯分类器在小型临床数据集上从炎症生物标志物预测认知障碍,并识别出I-309 (CCL1)作为关键预测特征。
利用生理信号通过机器学习预测考试结果
本研究探讨了利用皮肤电活动、心率和皮肤温度等生理数据,通过机器学习模型预测考试结果,发现深度学习方法与随机森林等简单模型均能有效发挥作用。
MS-MLB:基于血液的MS分类的开放机器学习基准
本文介绍了MS-MLB,这是一个开放机器学习基准,利用公开的GSE17048队列,从全血RNA表达数据中对多发性硬化症进行分类。它提供了一个可重复、受泄漏控制的评估流程,并报告梯度提升(Gradient Boosting)为表现最佳的方法。
机器学习用于培养前ESBL风险分层以指导经验性抗生素选择:一项12家医院的肠杆菌科培养物研究
研究人员开发了一个成本敏感的XGBoost模型,使用电子健康记录数据来预测产ESBL肠杆菌科感染,旨在指导经验性抗生素选择并减少碳青霉烯类药物的过度使用,这是一项12家医院的研究。