Interpretable analysis of smartphone addiction status and its associated factors among college students
Yuanning Li, Najie Zhao, Yanyan Wang · Frontiers in Psychiatry · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3389/fpsyt.2026.1850706
Methodology & findings
Study design
Cross-sectional survey with machine learning predictive modeling.
Sample
N = 2761, 6 groups
Primary method
LASSO (Least Absolute Shrinkage and Selection Operator) regression for feature selection with L1-regularization. XGBoost (Extreme Gradient Boosting) classifier for predictive modeling with hyperparameter optimization via grid search and 5-fold cross-validation. SHAP (Shapley Additive exPlanations) values for model interpretation. Model evaluation metrics: classification accuracy, precision, specificity, F1-score, and AUC. Comparison of seven machine learning models: Logistic Regression, Elastic Network, K-Nearest Neighbors, Decision Tree, XGBoost, Support Vector Machine, and Random Forest. Data analysis conducted in Python 3.9 using Pandas, NumPy, and XGBoost libraries.
Main result
The study identified smartphone addiction prevalence at 22.24% among college students, with "the factors influencing smartphone addiction among college students, ranked by mean absolute SHAP values, are: Loneliness (0.437), Monthly household income (0.067), Age (0.056), Place of residence (0.056)." The feature levels "Loneliness (Yes), monthly household income (≥9000 RMB), age (≥22 years), and place of residence (Urban) all demonstrated strong positive SHAP values, suggesting that these factors substantially elevate the predicted risk of mobile phone addiction in the college student population."
Reports effect sizes and confidence intervals.
Research paradigm
Positivist/Quantitative empiricism with machine learning
Author conclusions
"In this study, an XGBoost model was employed to identify factors associated with smartphone addiction among 2,761 college students. The key predictors identified included loneliness, monthly household income, age and place of residence." The authors further conclude that "our findings underscore the importance of early psychological screening and assessment for college students by educators and university administrators. Furthermore, targeted supportive interventions should be implemented to address modifiable risk factors, thereby mitigating the risk of smartphone addiction within this population."
Risk of bias
Selection bias from convenience sampling (non-random recruitment from five universities); Recall bias from self-reported questionnaire measures; Class imbalance in outcome variable (22.24% with addiction vs 77.76% without); Potential information leakage acknowledged by authors regarding feature selection timing; Absence of external validation cohort limiting generalizability; Cross-sectional design preventing causal inference; Selection bias due to convenience sampling; Recall bias from self-reported questionnaires; Class imbalance in dataset affecting performance metrics; Potential information leakage during feature selection; Selection bias from convenience sampling; Recall bias from self-reported measures; Class imbalance in dataset; No external validation cohort
Limitations
- The authors state: "First, the reliance on convenience sampling and self-reported measures may introduce selection and recall biases
- Second, the cross-sectional design precludes causal inferences regarding the relationships among the examined variables
- Third, the predictive model was constrained by a limited set of covariates, and potential class imbalance within the dataset may have affected performance metrics
- Additionally, although cross-validation was implemented, the absence of an independent external validation cohort limits the generalizability of the findings
- There is also a possibility of information leakage during the feature selection process (e.g., if screening was conducted prior to data splitting), which may have contributed to an overestimation of model performance
- Finally, the study did not incorporate biological markers (e.g., genomic data) or certain behavioral covariates (e.g., physical activity), thereby limiting insights into underlying physiological mechanisms and potential confounding pathways."
Open questions raised
- Future research should "prioritize longitudinal designs, recruit larger and more diverse cohorts for external validation, and integrate multi-dimensional data (including behavioral and biological indicators) to elucidate causal mechanisms and enhance predictive robustness."
- "Future research should prioritize longitudinal designs, recruit larger and more diverse cohorts for external validation, and integrate multi-dimensional data (including behavioral and biological indicators) to elucidate causal mechanisms and enhance predictive robustness."
- Future research should prioritize: longitudinal designs, larger and more diverse cohorts for external validation, integration of multi-dimensional data including behavioral and biological indicators to elucidate causal mechanisms and enhance predictive robustness.
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations