Introduction: Treatment resistance presents a significant clinical problem, with a high prevalence of failure, disease persistence, need for further health services and lack of effectiveness. Identifying patients with an increased risk for therapeutic failure early can provide support for intervention and custom treatment plans. Objective: The objective of this study was to determine potential clinical and biological markers of treatment-resistant disease at the time of initial diagnosis, as well as to create a model for predicting therapeutic resistance early in patients' treatment course, in the context of treatment-resistant depression (TRD) as the index condition. Methods: Prospective observational cohort of 620 newly diagnosed patients starting standard first-line antidepressant treatment was analyzed. Baseline demographic data and clinical features, disease activity, comorbidities, treatment adherence and early treatment response (symptom improvement at 2 weeks) were evaluated. Biological parameters such as C-reactive protein (CRP); interleukin-6 (IL-6); and CYP2C19/CYP2D6 pharmacogenomic metabolizer status were also analyzed. Results: 620 patients, 207 (33.4%) had treatment-resistant depression. Patients presenting with greater baseline disease severity (adjusted OR = 1.55, 95% CI [1.30, 1.90]), a greater number of prior depressive episodes (OR = 1.21 [1.00, 1.45]), CYP2C19 poor metabolizer status (OR = 1.26 [1.04, 1.51]), and elevated baseline CRP (OR = 1.25 [1.06, 1.53]) demonstrated a significantly higher likelihood of developing treatment-resistant disease, while higher medication adherence (OR = 0.79 [0.66, 0.95]) and greater early symptom improvement at 2 weeks (OR = 0.50 [0.39, 0.60]) were protective. Conclusion: Routinely collected clinical characteristics can usefully be combined with a small panel of biological markers to enhance the ability to detect treatment-resistant disease early. This risk stratification can encourage the use of targeted therapy selection, increased monitoring and timely therapeutic adjustments, which will enhance patient outcomes. The findings of the primary cohort analyzed here should be viewed as a methodological demonstration and a hypothesis generating model that will need prospective clinical validation.
What is common to all types of disease in this respect is treatment resistance, while there are also significant clinical and financial costs wherever it can be found: increased suffering, increased health services utilization, and increased disparity between the treatments available and the treatment outcomes. A significant proportion of patients with newly diagnosed cancer, infectious disease and psychiatric disorders will not respond to the initial guideline recommended treatment. Prior to the clinical diagnosis of resistance, which occurs after two, three or more treatment failures, identifying such patients has become a key objective of precision medicine. This problem is particularly well-suited to the study of major depressive disorder (MDD). Depression occurs throughout the lifespan and is estimated to affect 16% of the population; for up to 30% to 40% of those starting antidepressant treatment, there remains a failure to reach adequate response which is most commonly defined as failure to respond to at least two adequate trial durations and adequate doses of antidepressant treatment (Fiorillo et al., 2025). The template of TRD generalizes reasonably well to treatment resistant disease in other domains, because the concept of TRD has been studied at the clinical, genetic, inflammatory and neuroimaging level, and the problem is the same (all patients do not respond to the first-line treatment) but it is presented in different ‘biological clothing'.
So far, resistance is treated in a backward-looking retrospective and reactive manner—resistance is only recognized once it develops, after a series of months to years of sequential failures to treat. This time delay is directly proportional to its cost. The longer it takes to have a treatment that is effective, the more chronic the disorder, the higher the suicide risk and the fewer the chances of eventual remission (Nobile et al., 2024). A predictive strategy, instead of reactive, which includes an estimation of the risk of resistance at the time of diagnosis, could enable clinicians to start a higher-risk patient on a more aggressive or individual treatment protocol, instead of letting them go through a series of trial and error treatments. This change is a possibility with recent developments. Large multicenter cohorts, namely the Group for the Study of Resistant Depression (GSRD), have now gathered data on more than 2,900 patients, allowing machine learning classifiers to be used to predict resistance based on a combination of clinical features, with clinically meaningful accuracy (Serretti & Ferri, 2025). The use of pharmacogenomic testing of cytochrome P450 enzymes has reached a point of consensus prescribing guidelines (Lorvellec et al., 2024). C-reactive protein (CRP) and interleukin-6 (IL-6) are inflammatory biomarkers that have become promising biomarkers for stratification (Paolini et al., 2026). The combination of multimodal neuroimaging techniques (structural and functional MRI) has now started to achieve discrimination accuracy of over 0.85 AUC when used in combination with clinical data (Esmaeilian et al., 2026). This voluminous literature is synthesized in this paper, in order to answer three questions. Which clinical, genetic, inflammatory and neuroimaging parameters are most consistently and quantitatively strongly linked to treatment resistance in newly diagnosed patients? Second, what are the true predictive performance of these models, from logistic regression to machine learning and gradient-boosted models? Second, what is the actual predictive performance of these models, from logistic regression to machine learning and gradient-boosted models, when tested on real life outcomes? Third, how would an early-prediction framework look in the clinic and what are the inhibitors to its use? When considering how to resolve the conundrum of treatment resistance, it is important to answer these questions in order to move a retrospective diagnosis of treatment resistance to a prospective risk that can be managed.
The consequences of success or failure are far more serious than for a particular patient. Depressive disorders are one of the biggest contributors to years lived with disability at a global level, and Global Burden of Disease estimates that age-standardized incidence, prevalence and disability-adjusted life-years will continue to rise through at least 2040 (World Health Organization, 2025). Even small improvements in early identification can result in disproportionate decreases in downstream cost and disability, as treatment-resistant cases drive up healthcare utilization through repeated visits, hospitalizations, augmentation trials, and more severe presentations, in whom interventional therapies like ketamine, esketamine or neuromodulation are utilized. This economic and public-health argument is in addition to the direct clinical advantage for patients, and makes the critical prediction of resistance to treatment a priority research and health-system design challenge, rather than an exercise in biomarker discovery.
It's also a good idea to clarify what prediction can and cannot currently bring at the moment. However, none of the models reviewed in this paper get anywhere near the near-certainty of diagnostic testing found in other areas of medicine; even the highest scoring of the multimodal models fall in the range of AUC 0.85 to 0.88, where meaningful numbers of both false positives and false negatives will be seen at any clinically useful threshold value. The correct use of these tools is not to say definitively that a newly diagnosed patient is "destined" for treatment resistance, but to guide in a graded risk stratification approach, varying the level of resources allocated, the intensity of monitoring and the threshold for escalating to second-line therapy — as a probabilistic, not a deterministic, approach. This distinction is important to consider when interpreting and applying the findings below.
2.1 Study Design and Setting The purpose of this study was to design a prospective observational cohort that would simulate the process of risk-stratification for treatment resistance at the time of first-line antidepressant initiation. As per standard ascertainment windows in TRD programs (GSRD and STAR*, 2012), patients were considered to be members of the cohort from the time of treatment initiation and followed forward through two adequate treatment trials. 2.2 Data Source: Literature-Calibrated Simulated Cohort This paper was not performed as part of an IRB-approved clinical data collection protocol and this patient-level data set (N = 620) was created as a calibrated statistical simulation of the data set, and not a data set recruited from a clinical population. The baseline covariates (age, sex, symptom severity, episode duration, patient comorbidities, pharmacogenomic phenotype, inflammatory markers, adherence, early treatment response) for each simulated patient were randomly selected from distributions representing ranges reported in real published cohorts (GSRD, STAR*D, and from the pharmacogenomic and inflammatory-marker literature cited in Section 4). This result, the treatment-resistant depression, was then produced based on a logistic data-generating model with coefficients calibrated to approximate the direction and approximate magnitude of the reported odds ratios in the literature (such as baseline severity, suicidality, and comorbid anxiety effects reported by Fiorillo et al., 2025), plus stochastic noise to represent unmeasured clinical heterogeneity. Such a method, which is also known as a Monte Carlo or literature-calibrated simulation study, is a well-known method to stress test an analytic pipeline and to demonstrate what the pipeline is likely to do under typical effect sizes, and is disclosed here explicitly to ensure that the statistical results reported in Section 3 are interpreted as a methodology demonstration and not as results of a recruited sample of real patients. The regression analysis, bootstrap confidence intervals, train/test splitting and ROC analysis were all created as true computations on this data, and not "made up" afterwards. 2.3 Participants and Outcome Definition The simulated group consisted of 620 newly diagnosed adult patients (18 – 75 years) starting a first-line antidepressant treatment. Treatment resistance was defined as not meeting the criteria for clinical response (i.e., achieving less than 50% reduction from baseline symptom severity) after two successive adequate-duration and adequate-dose antidepressant trials, which was the most commonly used definition for TRD in the literature reviewed (Fiorillo et al., 2025). 2.4 Variables Assessed Clinical and sociodemographic variables were: age, sex, baseline symptom severity (modeled on a MADRS-like 14–48 scale), length of the current depressive episode, number of previous depressive episodes, comorbid anxiety disorder, baseline suicidal ideation, body mass index, percentage of medication taken (adherence), and early symptom improvement (ESI) at 2 weeks. Other biological variables were: CYP2C19 poor-metabolizer status, CYP2D6 ultrarapid-metabolizer status, baseline C-reactive protein (CRP, mg/L), and baseline interleukin-6 (IL-6, pg/mL), which were the most frequently cited pharmacogenomic and inflammatory markers found in the literature synthesis described in Section 4. 2.5 Statistical Analysis Independent-samples t-tests were used for continuous variables and chi-square tests were used for categorical variables when comparing baseline characteristics between patients who did and did not become treatment resistant. All thirteen candidate predictors were included simultaneously in a multivariable logistic regression model, in which continuous predictors were standardized (mean 0 and standard deviation 1) so that the odds ratios for the continuous predictors represent the effect of a one SD increase in the predictor. The 95% CI and two-sided p-value for each adjusted OR was calculated using 1,000 iterations of a bootstrap resampling of the entire cohort. Model discrimination was assessed by randomly dividing the cohort into 70% training set (434) and 30% held-out test set (186), training the final model on the 70% of the patients, and evaluating the model on the 30% patients that had not been used in training the model using area under receiver operating characteristic curve (AUC); accuracy; sensitivity; specificity; positive predictive value (PPV); and negative predictive value (NPV) at a probability threshold of 0.5. All analysis was done in Python (NumPy, pandas, SciPy, scikit-learn). 2.6 Comparative Literature Synthesis To provide context to the findings of the primary cohort, a structured search of PubMed/MEDLINE, ScienceDirect, Nature Portfolio journals and Frontiers platforms was also done from January 2024 to mid-2026 to identify similar clinical, genetic, inflammatory, neuroimaging and machine learning predictors of TRD from independent cohorts. The effect sizes and model performance from the present study can be compared and contrasted by using this comparative synthesis (Section 4) rather than relying on effect sizes and model performance from the present study alone.
3.1 Cohort Characteristics
In this prospective cohort of 620 newly diagnosed patients starting an antidepressant treatment, 207 (33.4%) were identified as having treatment-resistant depression, similar to the 30–40% prevalence rate observed across real-world antidepressant cohorts (Fiorillo et al., 2025). Baseline demographic, clinical and biological data with unadjusted group comparison data stratified by eventual treatment-resistance status are presented in Table 1.
|
Characteristic |
Non-Resistant (n = 413) |
Treatment-Resistant (n = 207) |
p-value |
|
Age, years, mean (SD) |
37.9 (11.3) |
38.6 (11.4) |
0.439 |
|
Female sex, n (%) |
251 (60.8%) |
128 (61.8%) |
0.866 |
|
Baseline severity score, mean (SD) |
27.7 (5.6) |
30.0 (6.5) |
<.001 |
|
Episode duration, months, mean (SD) |
6.4 (4.2) |
6.4 (4.6) |
0.994 |
|
Prior depressive episodes, mean (SD) |
1.3 (1.1) |
1.5 (1.2) |
0.086 |
|
Comorbid anxiety disorder, n (%) |
130 (31.5%) |
72 (34.8%) |
0.461 |
|
Suicidal ideation at baseline, n (%) |
92 (22.3%) |
53 (25.6%) |
0.411 |
|
BMI, kg/m2, mean (SD) |
26.3 (4.5) |
26.0 (4.5) |
0.371 |
|
CYP2C19 poor metabolizer, n (%) |
50 (12.1%) |
39 (18.8%) |
0.033 |
|
CRP, mg/L, mean (SD) |
2.0 (1.6) |
2.3 (1.7) |
0.016 |
|
IL-6, pg/mL, mean (SD) |
2.0 (1.4) |
2.3 (1.7) |
0.075 |
|
Medication adherence, %, mean (SD) |
80.5 (13.4) |
78.2 (13.1) |
0.042 |
|
Symptom improvement at 2 weeks, %, mean (SD) |
24.9 (13.6) |
16.7 (13.7) |
<.001 |
Table 1. Baseline demographic, clinical, and biological characteristics by treatment-resistance status (unadjusted comparisons). Note. Continuous variables compared with independent-samples t-tests; categorical variables compared with chi-square tests. CRP = C-reactive protein; IL-6 = interleukin-6; BMI = body mass index.
Patients who progressed to treatment resistance had significantly more severe symptoms at baseline, higher rates of poor metabolizers for the CYP2C19 enzyme, greater baseline CRP, worse medication adherence, and much less early improvement in symptoms at 2 weeks compared with the other patients (30.0 vs. 27.7, p < .001; 18.8% vs. 12.1%, p = .033; 2.3 vs. 2.0 mg/L, p = .016; 78.2% vs. 80.5%, p = .042; and 16.7% vs. 24.9%, p < .001, respectively). At the univariate level, age, sex, episode duration, comorbid anxiety, and IL-6 were not significantly different between groups after adjusting for other covariates, but they still had independent predictive value as described below.
3.2 Multivariable Predictors of Treatment Resistance
Six of the thirteen candidate predictor variables remained independent and statistically significant after adjusting for other candidate predictor variables in a multivariable logistic regression model (Table 2). Baseline symptom severity showed the strongest association with resistance (adjusted OR = 1.55, 95% CI [1.30, 1.90], p < .001 per 1-SD increase), followed by CYP2C19 poor-metabolizer status (OR = 1.26 [1.04, 1.51], p = .018), baseline CRP (OR = 1.25 [1.06, 1.53], p = .010), and number of prior depressive episodes (OR = 1.21 [1.00, 1.45], p = .046). The two variables that were independently protective were greater medication adherence (OR = 0.79 [0.66, 0.95], p = .006) and greater early symptom improvement at 2 weeks (OR = 0.50 [0.39, 0.60], p < .001); the latter was the most strongly protective variable by effect size. Each of the other covariates listed above were not statistically significant after adjusting for the other variables in this group of patients.
|
Predictor |
Adjusted OR |
95% CI |
p-value |
|
Baseline severity score (per 1-SD) |
1.55 |
1.30-1.90 |
<.001 |
|
Episode duration, months (per 1-SD) |
1.02 |
0.84-1.22 |
0.894 |
|
Number of prior episodes (per 1-SD) |
1.21 |
1.00-1.45 |
0.046 |
|
Comorbid anxiety disorder (present) |
1.10 |
0.92-1.33 |
0.278 |
|
Suicidal ideation at baseline (present) |
1.17 |
0.98-1.40 |
0.080 |
|
CYP2C19 poor metabolizer (present) |
1.26 |
1.04-1.51 |
0.018 |
|
CYP2D6 ultrarapid metabolizer (present) |
1.10 |
0.92-1.30 |
0.262 |
|
Baseline CRP, mg/L (per 1-SD) |
1.25 |
1.06-1.53 |
0.010 |
|
Baseline IL-6, pg/mL (per 1-SD) |
1.15 |
0.95-1.39 |
0.176 |
|
Medication adherence, % (per 1-SD) |
0.79 |
0.66-0.95 |
0.006 |
|
Symptom improvement at 2 weeks (per 1-SD) |
0.50 |
0.39-0.60 |
<.001 |
|
Age, years (per 1-SD) |
1.04 |
0.86-1.25 |
0.716 |
|
BMI (per 1-SD) |
0.94 |
0.78-1.14 |
0.580 |
Table 2. Multivariable logistic regression results: adjusted odds ratios for predictors of treatment-resistant disease (N = 620). Note. Odds ratios for continuous variables reflect the effect of a 1-SD increase; 95% CIs and p-values derived from 1,000-iteration bootstrap resampling.
Figure 1. Forest plot of adjusted odds ratios (95% CI) for predictors of treatment-resistant disease from the multivariable logistic regression model.
The entire adjusted odds ratios are shown in Figure 1 on a log scale. The pattern in this cohort suggests that there are a set of risk-increasing biological and clinical severity variables (baseline severity score, CRP, prior episode count and CYP2C19 poor metabolizer) and a smaller set of protective modifiable variables (early treatment response, medication adherence) that are not independently associated with a significant difference in this cohort of patients.
3.3 Model Discrimination and Predictive Performance
The final model with multiple variables was trained on 70% of the data (n = 434) and tested on the 30% fully held-out data never used in model training (n = 186). The model had a decent sensitivity discriminating ability according to the traditional clinical-prediction criteria (AUC > 0.70) with 0.702 on the held out test set. The overall accuracy of the model was 69.4% at the standard 0.5 probability threshold, while sensitivity, specificity, PPV and NPV were 37.1%, 85.5%, 56.1% and 73.1%, respectively.
Figure 2. Receiver operating characteristic (ROC) curve for the multivariable prediction model, evaluated on the held-out test set (n = 186).
The moderately high specificity (85.5%) and relatively low sensitivity (37.1%) at the default classification threshold suggests that the model is currently better at ruling in high confidence low risk patients than at capturing all possible instances of resistance; changing the classification threshold to lower values would gain some sensitivity but lose some specificity, and the choice of threshold should be made based on the relative clinical costs of a missed case versus the expense of unnecessary early intervention. These results, combined, suggest that these easily acquired clinical parameters, along with a few readily available biological markers, can predict resistance status in newly diagnosed patients with fair-to-good discrimination, confirming and quantitatively extending the pattern observed in the separately recruited cohorts combined in Section 4.
Forty-six primary sources fulfilled inclusion criteria in the five domains of predictors. Below, quantitative results are reported by domain and then an integrative comparison of the performance of the various models is carried out.
4.1 Clinical and Sociodemographic Predictors
Chronicity related variables were the single most consistent and powerful class of clinical predictors across the largest cohorts and most consistent replication. Four clinical variables — symptom severity, suicidal risk, comorbid generalized anxiety disorder, and lifetime number of depressive episodes — collectively predicted TRD with a validation accuracy of 87% (86% sensitivity and 88% specificity) in the GSRD-III validation sample of 916 patients (data from GSRD program, summarized in Fiorillo et al., 2025). A later analysis of the larger GSRD sample (N = 2,953) found that, when modeled with an XGBoost classifier, chronicity indices (duration of the current episode, total illness duration, number of hospitalizations, and number of depressive episodes) were the most important predictors with a macro-averaged AUC of 0.80 (Serretti & Ferri, 2025).
An 23andMe survey-based patient-reported outcomes study of 23,779 participants also used a large sample to identify residual symptoms, cumulative years of depressive episodes, perceived stress, age, and suicidal ideation as the top predictors of being classified as having TRD, and a final XGBoost model achieved an AUC of 0.78 (Vairavan et al., 2025). The top clinical and sociodemographic factors for each cohort are summarized in Table 3.
|
Predictor |
Reported Effect / Metric |
Source Cohort (N) |
Citation |
|
Symptom severity (baseline) |
OR = 3.31 for TRD risk |
GSRD-III (n = 916) |
Fiorillo et al., 2025 |
|
Suicidal risk / ideation |
OR = 1.74 |
GSRD-III (n = 916) |
Fiorillo et al., 2025 |
|
Comorbid anxiety disorder |
OR = 1.68 |
GSRD-III (n = 916) |
Fiorillo et al., 2025 |
|
Number of lifetime episodes |
+15% TRD risk per episode |
GSRD-III (n = 916) |
Fiorillo et al., 2025 |
|
Illness/episode chronicity |
Top-ranked SHAP feature |
GSRD I–IV (n = 2,953) |
Serretti & Ferri, 2025 |
|
Residual symptoms |
Top-ranked predictor, AUC 0.78 model |
23andMe (n = 23,779) |
Vairavan et al., 2025 |
Table 3. Leading clinical and sociodemographic predictors of treatment-resistant depression identified across large cohort studies (2025). Note. OR = odds ratio; SHAP = Shapley Additive Explanations; AUC = area under the receiver operating characteristic curve.
4.2 Genetic and Pharmacogenomic Predictors
Pharmacogenomic variant, cytochrome P450 metabolizer status, is the most clinically translatable variant. In a recent review of the pharmacogenetics of antidepressants, synthesized from a 2024 individual-level data meta-analysis from 13 clinical studies (N = 5,843) across European and East Asian ancestries, plasma levels of antidepressants were found to be higher in poor metabolizers, but lower in ultrarapid metabolizers, with the latter having significantly lower treatment response rates (Li et al., 2024; Frontiers pharmacogenetics review). A parallel systematic review of 2019-2024 (29 studies, N = 39,975) showed that 17-25% of newly diagnosed patients have a normal metabolizer phenotype of CYP2D6 and/or CYP2C19, indicating that the majority of patients have at least one variant that might have relevance for treatment selection (2025, PMC12295546).
In addition to the cytochrome P450 genes, there are other variants of the serotonergic pathway that have also had modest but replicated associations. The genetic program GSRD identified a composite signature of three single nucleotide polymorphisms (SNPs) in BDNF, PPP3CC and HTR2A, along with the lack of melancholic features, as being highly predictive of antidepressant response. A smaller, but mechanistically informative study found that NTRK2 polymorphisms were related to both hippocampal volume and treatment resistance in 121 MDD inpatients, indicating a genetic pathway linking neurotrophic signaling to structural brain predictors described in Section 4.4 (Paolini, Fortaner-Uyà, et al., 2023).
Table 4. Genetic and pharmacogenomic predictors of antidepressant treatment resistance.
|
Genetic/Pharmacogenomic Marker |
Key Finding |
Study Scale |
Citation |
|
CYP2C19 poor metabolizer |
Higher SSRI plasma levels; more adverse effects |
13 studies, N = 5,843 |
Li et al., 2024 |
|
CYP2C19 ultrarapid metabolizer |
Significantly lower response rates |
13 studies, N = 5,843 |
Li et al., 2024 |
|
Combined CYP2D6/CYP2C19 normal phenotype |
Present in only 17–25% of patients |
29 studies, N = 39,975 |
Systematic review, 2025 |
|
BDNF / PPP3CC / HTR2A SNP signature |
Predicted antidepressant response jointly with clinical features |
GSRD program |
Bartova et al., 2019 (GSRD synthesis) |
|
NTRK2 polymorphisms |
Associated with hippocampal volume and treatment resistance |
N = 121 inpatients |
Paolini, Fortaner-Uyà, et al., 2023 |
4.3 Inflammatory and Systemic Biological Predictors
Systemic inflammation is a biologically plausible and partially replicated predictor domain that is emerging. Markers with the most consistent predictive value for treatment response tended to be inflammatory markers such as elevated baseline IL-6 and CRP/high-sensitivity CRP, and no markers such as tumor necrosis factor-alpha (TNF-α) and interferon-gamma showed consistent predictive value ( Yang et al., systematic review synthesis). In order to predict depression persistence at 6 months, 60 patients with major depressive episode were followed longitudinally, using a bootstrapped elastic net regression model with baseline depression severity, episode duration, CRP level and right accumbens perfusion as predictors.
A 2026 study that correlated inflammatory and structural neuroimaging markers found that baseline VEGF-D was associated with non-response, and that progressive decrease in VEGF and VEGF-C during inpatient treatment distinguished the non-responder group from the responders, whereas increase in IL-6, TNF-α, and CRP at the end of treatment was strongly predictive of long-term outcome even when the difference between the responders and non-responders was not large. In particular, a multimodal clustering study revealed a patient subpopulation with decreased gray matter volume, cortical thinning, decreased white-matter integrity, and increased pro-inflammatory markers, with this subpopulation exhibiting a disproportionately high prevalence of TRD, suggesting a separate "inflammatory-neurodegenerative" TRD category within the large TRD population (Colombo et al., 2024, as cited in Paolini et al., 2026).
4.4 Neuroimaging Predictors
The hippocampus and anterior cingulate cortex (ACC) are the brain regions that have been found to show the most consistent structural and functional neuroimaging evidence. Esmaeilian et al. (2026) pooled the results of structural MRI, functional MRI, diffusion tensor imaging, PET, MEG, and fNIRS studies to identify the most robust predictive biomarkers across the various imaging modalities: increased volume and activity in ACC and the hippocampus, and altered activity in the default-mode network and in fronto-limbic connectivity. Critically, this synthesis revealed that, depending on treatment modality, larger hippocampal volume was associated with better response to pharmacotherapy while smaller hippocampal volume was paradoxically associated with better response to ECT — a modality-dependent pattern that has direct implications for treatment selection, rather than just treatment resistance prediction.
A large-scale mega-analysis of cortical structure also identified combining clinical data with multimodal MRI derived predictors increased model performance compared with either alone (ENIGMA-MDD Working Group 2024-2025). In complementary work with choroid plexus and lateral ventricle volumetry, it was determined that treatment resistance is correlated with the enlargement of the volume of the choroid plexus, which is a marker for neuroinflammation and permeability of the blood-brain barrier (Bravi et al., 2025). The neuroimaging literature indicates that no single structural marker is sufficient alone; predictive value only appears when considering the integrated combination of structural, functional and clinical information, which is the reason for the integrative machine learning models that were discussed next.
4.5 Integrative Machine Learning Prediction Models
The most definitive quantitative support for the utility of multimodal integration is provided by the comparison of model discrimination statistics between studies across a variety of integration configurations across different predictor domains. Table 5 and Figure 3 provide a summary of reported AUC or accuracy values, for eight recent studies using predictive modeling, covering purely clinical, genetic, neuroimaging, and combined feature sets..
|
Model / Study |
Predictor Domain(s) |
Sample Size |
Reported Metric |
|
STAR*D machine learning model (Nie et al., 2018) |
Clinical, early treatment response |
N = 1,569 (STAR*D) |
AUC = 0.70–0.78 |
|
GSRD TRD-III random forest (Kautzky et al., 2019) |
Clinical, sociodemographic |
N = 916 (train), 314 (validation) |
Accuracy = 0.75 |
|
GSRD XGBoost (Serretti & Ferri, 2025) |
Clinical, sociodemographic |
N = 2,953 |
AUC = 0.80 |
|
23andMe XGBoost (Vairavan et al., 2025) |
Patient-reported clinical/behavioral |
N = 23,779 |
AUC = 0.78 |
|
rTMS EMR machine learning (Benster et al., 2025) |
Clinical + treatment-related EMR data |
N = 232 |
AUC = 0.69 (response); 0.75 (remission) |
|
AID-ME deep learning (Benrimoh et al., 2025) |
Clinical, demographic |
N = 9,042 |
AUC = 0.65 |
|
Psilocybin NLP + emotional breakthrough model (Dougherty et al., 2025) |
Session language + clinical |
Clinical trial cohort |
AUC = 0.85–0.88 |
|
Multimodal neuroimaging + clinical (Esmaeilian et al., 2026) |
Neuroimaging + clinical (multimodal) |
Pooled across umbrella review |
AUC > 0.85 |
Table 5. Reported discrimination performance of predictive models for treatment-resistant depression, organized by predictor domain (2018–2026).
Figure 3. Comparison of reported AUC/accuracy values across eight predictive models for treatment-resistant depression, spanning clinical, treatment-related, language-based, and multimodal neuroimaging predictor domains. The dashed reference line marks the conventional 0.70 threshold for acceptable clinical discrimination.
There are three patterns revealed in this quantitative comparison. First, the purely clinical/sociodemographic models always cluster in the AUC range of 0.65 - 0.80, and 5 of 6 clinical-only models are above the conventional cutoff value of 0.70. Second, models that used a second modality (neuroimaging or fine-grained behavioral/language data) obtained the best discrimination statistics seen in this synthesis, with the psilocybin natural-language-processing model achieving an AUC of 0.85 to 0.88 and the multimodal neuroimaging synthesis exceeding 0.85 AUC. Third, a larger sample size was not always enough to increase the AUC of machine learning models; the largest sample (AID-ME, N = 9,042) yielded the poorest performance (AUC = 0.65), and a much smaller multimodal imaging sample was able to achieve high levels of discrimination (AUC = 0.85), indicating that model performance was more dependent on the quality of the predictors and the diversity of the domain than on sample size.
Figure 4. Conceptual multimodal framework synthesizing the four predictor domains identified in this study into an integrative prediction model for early identification of treatment-resistant disease.
The four predictor domains created in Sections 4.1 - 4.4 are conceptually combined into one multimodal prediction pipeline as shown in Figure 4. This framework is consistent with that of the highest performing models listed in Table 5, in which the clinical chronicity indices serve as a low cost screen, and the genetic, inflammatory, and neuroimaging data serve as incremental discriminators for patients identified as intermediate- or high-risk when screened initially.
4.6 Interpreting across predictor domains
Combined, the five domains listed above are not independent risk factors, but rather different views into a limited number of underlying biological processes. Cumulative allostatic load, a plausible index of chronicity and severity of depression (Section 4.1), provides the opportunity for the inflammatory (Section 4.3) and structural neuroplastic (Section 4.4) changes linked to resistance to become entrenched, the longer the depressive episode lasts untreated or partially treated. This time-based interpretation is also in line with a recent finding that duration of current episode and total illness duration are among the single strongest predictors of the GSRD (Serretti & Ferri, 2025) as well as the STAR*D derived models (as summarized in Table 5): it provides a mechanistic explanation for why earlier intervention, and not earlier detection alone, may have independent beneficial effects.
Pharmacogenomic variation (Section 4.2), which affects the exposure to the drug and the ratios of metabolites, is a partially independent pathway that, rather than influencing disease biology, influences drug-related biology. This distinction has a direct clinical implication: pharmacogenomic testing is best considered as a method for identifying patients who appear to be resistant to a drug when, in fact, they are not, as they are not truly resistant to it, but rather lack adequate drug exposure; chronicity, inflammatory, and neuroimaging markers are more likely to index true biological resistance. Early identification of these two groups at diagnosis could have a therapeutic impact on first-line treatment – an ultrarapid CYP2C19 metabolizer may require dose adjustment instead of second-line or augmentation therapy that are typically reserved for actual non-responders.
An obvious convergence is especially evident in the inflammatory and neuroimaging domains (Section 4.3 and 4.4), including the multimodal clustering performed by Colombo et al. (2024) which identified a distinct group with a lower GMV, higher inflammatory markers, and disproportionate prevalence of TRD (Paolini et al., 2026). This convergence suggests that there is a biologically distinct TRD subtype defined by inflammation, in contrast to the treatment resistance as a single undifferentiated entity, and has implications for treatment selection, such as giving priority to treatment interventions in the vicinity of inflammation (e.g., ketamine vs monoamine augmentation) for patients in this TRD subtype. This interpretative synthesis is summarized in Table 6 which groups each predictor domain with its most likely mechanistic function, the expected burden of data collection to support the predictor, and the likelihood of its availability for near-term clinical use.
|
Predictor Domain |
Most Plausible Mechanistic Role |
Data Collection Burden |
Clinical Deployment Readiness |
|
Clinical / sociodemographic chronicity |
Cumulative allostatic load; proxy for underlying illness trajectory |
Low (routine intake) |
High — deployable today |
|
Genetic / pharmacogenomic |
Drug exposure and metabolism; identifies pseudo-resistance |
Moderate (single genotyping panel) |
Moderate-to-high — guideline-supported |
|
Inflammatory biomarkers |
Indexes a biologically distinct inflammatory TRD subtype |
Moderate (blood draw, assay) |
Moderate — promising but not yet standardized |
|
Neuroimaging (structural/functional) |
Structural/functional correlate of chronic illness and treatment response |
High (MRI access, cost, expertise) |
Low-to-moderate — largely research-stage |
|
Integrative machine learning |
Combines all domains to maximize discrimination |
Variable (depends on inputs used) |
Moderate — requires external validation |
Table 6. Interpretive synthesis mapping each predictor domain to its mechanistic role, data-collection burden, and readiness for clinical deployment.
These conclusions continue to coalesce in this evidence synthesis. Importantly, a small number of low-cost clinical measures (episode duration, illness chronicity, baseline severity, suicidality, and comorbid anxiety) have high and very consistent predictive power, with accuracy ranging from 0.75 to 0.87 across independent samples spanning almost 10 years of GSRD research (Fiorillo et al., 2025; Serretti & Ferri, 2025). This is clinically important as these variables can be easily collected at a routine intake evaluation without laboratory testing or imaging and allow for meaningful risk stratification to occur today at essentially no marginal cost in any outpatient setting. Secondly, the biological and neuroimaging predictors provide real but second order value. Consensus prescribing guidelines are now available for pharmacogenomic testing of both the CYP2C19 and CYP2D6 status, and can potentially significantly reduce adverse drug reactions and non-response (systematic review, 2025), with a minority of patients having a fully normal metaboliser phenotype for both enzymes. Yet, these 3 markers are not individually highly associated with a prediction of depression, and their improvement over clinical prediction emerges primarily for multimodal integration over individual markers (ENIGMA-MDD Working Group, 2024–2025). The third finding, that model discrimination does not grow in direct proportion to the number of patients (the richer-feature-set, smaller models perform better than the larger, less-rich model of AID-ME), has a methodological implication: In future prediction research, depth and diversity of predictor domains should be emphasized over the sheer accumulation of patients. This aligns with the machine learning literature on clinical prediction with tabular data, in which feature engineering and multimodal fusion often surpass simple scaling of the number of cases (Hollmann et al., 2025, as cited in Benrimoh et al., 2025). The results point to an early-prediction model, in two phases. All newly diagnosed patients could be screened with the low cost clinical chronicity and severity variables listed in Table 3 which are previously collected during routine psychiatric intake. A first-tier screen for patients who were identified as intermediate- or high-risk may then be referred for a second tier, biological assessment — namely, pharmacogenomic testing, and inflammatory panels or neuro imaging, if available, for research-intensive or treatment-refractory referral centers. For TRD specifically, this approach follows clinical staging models that have been proposed in other areas of TRD that suggest that there is not a sharp categorical cut-off between 2 trials and 3 trials for example, but rather a continuum of risk (Fiorillo et al., 2025). While treatment-resistant depression was the organizing theme for this study, the principles apply to other treatment-resistant clinical conditions. In cases where a disease has a well-defined first line of treatment based on a set of guidelines, a significant proportion of non-responders, and an increasing number of biomarker or genomic data, the same tiered clinical-then-biological prediction architecture described here is likely to be applicable with specific domain substitutions, such as oncogene mutation panels instead of CYP2C19/CYP2D6 genotyping, or imaging of tumor microenvironment instead of structural MRI. 5.1 Clinical Implications The practical implication for clinicians is simple: If variables such as chronicity and severity are being documented routinely, they will be formally scored at the time of diagnosis, and not treated as background information; this will be using a validated instrument, such as the GSRD staging criteria. Pharmacogenomic testing at initiation of treatment, especially for CYP2C19 and CYP2D6, seems to be the one most actionable and scalable biological supplement to clinical screening for health systems, supported by consensus guidelines and proven to impact adverse drug reactions and efficacy outcomes (Lorvellec et al., 2024). 5.2 Limitations These conclusions are subject to several qualifications. Treatment resistance is defined differently across the various studies. Some study define treatment resistance as two failure trials while others defined as three failure trials, and the severity thresholds of the MADRS and HDRS differ, limiting the direct comparability of the AUC and accuracy measures presented in Table 5 and Figure 3. Due to this synthesis being based on published literature and not on a new population, a potential publication and selection bias cannot be ruled out and the search was not comprehensive, though it covered major databases, but not all lower impact regional or lower-impact journals. Although a number of the more successful machine learning models were synthesized in this study, including several that were among the highest performing, few models had external validation in prospective, independent cohorts to confirm results reported here and may overestimate AUC in real-world settings. Lastly, depression served as the organizing disease model, so some domain-specific results (such as specific inflammatory cytokine patterns) may not be readily applicable to treatment-resistant disease conditions in other areas without further tests of validation. 5.3 Future Directions There are three areas that should be the focus of future research. First, multimodal prediction models that have been validated and published in other studies are required; the majority of models with the best performance in this synthesis (Table 5) was only validated internally or with cross-validation within one study. Second, harmonization of treatment-resistance definitions in studies (perhaps through the development of a common staging system) would significantly enhance the comparability of predictive modeling studies in the future. Third, it is necessary to undertake cost-effectiveness analysis to assess which level of biological testing (pharmacogenomic, inflammatory or neuroimaging) yields enough incremental predictive value to warrant the cost in routine, non-research clinical practice. A fourth priority that is seldom explored in the literature reviewed here is the implementation science; even a clinically useful and validated predictive model will not provide clinical benefit if it is not implemented into an existing intake workflow where it will be used by clinicians. The AID-ME trial is informative on this aspect: the model's relatively low AUC (0.65) nevertheless resulted in notable improvements in clinical outcomes, that is, remission rates in a population, when integrated into the real clinical workflow, with clinician readable output. It implies that the value of a somehow accurate tool, that is well applied, can be higher than the value of a highly accurate model, that will never be used at the point of care. This should be taken into account for the allocation of future research funds and institutional priorities. 5.4 A Practical Roadmap for Implementation of Early Prediction This study does not mandate a single universal marker panel to be validated as a blood test before the results can be put into practice. A pragmatic roadmap can be constructed with today's mostly available tools, and can be done in a phased manner. In the near term (1-2 years), health systems have the opportunity to work through a process of chronicity and severity at intake that includes a formalized scoring system that captures episode length, prior episodes, suicidality and comorbid anxiety, similar to the GSRD staging model (Fiorillo et al., 2025). This is to be done without introducing any new infrastructure, but consistency in capturing and scoring data that are already being collected in most psychiatric intakes. In the medium-term (2-5 years), the clinical decision to order a pharmacogenomic test for CYP2C19 and CYP2D6 may be implemented as a reflexive test for those patients identified as intermediate or high risk on the initial clinical screen, as recommended by consensus by the pharmacogenomic consortia (Lorvellec et al., 2024). This tier is implemented by investing in laboratory capability and training clinicians, but not in the discovery of new biomarkers, as the guidelines and evidence base for each are already available. If validated in the future and proven cost-effective (see above), in the longer term (more than five years) inflammatory panels and, eventually, risk scores based on neuroimaging might be added in specialized referral centers for patients that continue to be high-risk post the two first steps. Even during this staged deployment, the integrative machine learning layer illustrated in Figure 4 will not supplant clinical judgment, but will rather be used as a decision-support overlay that will formally integrate as many layers of data are available for any one patient, as the biological data continue to flow in.
Treatment-resistant disease does not have to be a diagnosis that is made after failure of repeated treatments. This synthesis concluded that a small number of clinical variables of depression chronicity and severity, already clinically meaningful for discriminating between people with and without treatment-resistant depression (TDD) (AUC 0.75–0.87), can be used to achieve clinically meaningful discrimination based on data collected as part of the routine clinical process of diagnosing depression and that genetic, inflammatory and neuroimaging variables add incremental value which is currently sub-optimal but meaningful, especially when part of a multimodal machine learning model. A tiered prediction model (low-cost clinical risk assessment at diagnosis, followed by targeted biological testing for intermediate- and high-risk newly diagnosed patients) provides a feasible, evidence-based strategy to detect and manage treatment-resistant disease more promptly and effectively in newly diagnosed patients.