Auditing Construct Overlap in Explainable Machine Learning: Evidence from Burnout-Depression Prediction Across Student Cohorts
2026-07-12 • Machine Learning
Machine LearningArtificial Intelligence
AI summaryⓘ
The authors show that explainable machine learning models predicting complex mental health outcomes can seem reliable across different groups, but this is mostly because of how the outcomes are defined. They used a specific model on medical and non-medical students and found that some factors like trait anxiety repeatedly appeared most important. When they adjusted for overlap between related variables, the model's accuracy and feature importance dropped a lot, showing the initial results were influenced by shared information. Their main advice is that researchers should do similar adjustments to avoid misleading conclusions when working with correlated mental health measures.
Explainable Machine LearningElasticNetTrait AnxietyResidualizationComposite OutcomesKendall TauR-squaredMental Health AssessmentCorrelated VariablesPrediction Interval
Authors
Alireza Dehghan, Negin Ashrafi
Abstract
Explainable machine learning (XML) pipelines applied to composite mental health outcomes can produce apparently-robust, cross-population-stable risk hierarchies that are largely artefacts of how the outcome was constructed. We demonstrate this using an ElasticNet pipeline applied to 886 medical students at the University of Lausanne (primary cohort, 2022), validated across 2,580 longitudinal observations at three time points and 701 non-medical students from eight faculties; all three datasets share identical instruments. The pipeline produces a hierarchy in which trait anxiety and health satisfaction dominate wherever the outcome is measured, with Kendall $τ= 1.0$ for the top-two positions across all five evaluation sets and consistent transfer performance ($R^2$: 0.41-0.49). Two residualization experiments, which isolate shared variance between correlated variables via regression, reveal the mechanism: when trait anxiety (STAI-T) is residualized against the co-included depression subscale (CES-D, $r = 0.72$), model $R^2$ drops from 0.41 to 0.16 and STAI-T falls from rank 1 to rank 6; when burnout subscales are residualized against CES-D, $R^2$ collapses to 0.016. Prediction intervals average 35.4 units on a 0-100 scale (2.4 outcome standard deviations), independently ruling out individual-level deployment. The residualization protocol is the paper's transferable contribution: any XAI study combining correlated predictor and outcome constructs should apply this check before interpreting apparent stability as a finding.