Regression to the mean is a phenomenon whereby, naturally, initially very high scores tend to fall, and vice versa, low ones tend to rise; in other words, extreme values tend to normalize.
For example, if we study a group of patients with anxiety twice over time, it may be that at the initial assessment some patients have anxiety levels far higher than their usual levels due to transient factors (a bad day, a one-off problem). Thus, at the second assessment, when these factors have disappeared, these patients would show lower levels of anxiety, closer to their usual levels. Conversely, some patients might present at baseline with anxiety levels lower than their usual ones due to transient factors (a calm day, a small joy), and at the second assessment show higher anxiety levels, close to their usual ones. Thus, since those who scored high at baseline tend to decrease and those who scored low tend to increase, a model might appear capable of predicting the course of anxiety from initial anxiety, when in reality it is only capturing the natural tendency of extreme values to move towards the mean. Therefore, in this case, the model would have appeared to have good predictive ability without having identified any truly useful information about the progression of symptoms.
Four researchers from the Clínic-IDIBAPS, Joaquim Raduà, Eduard Vieta, Vincenzo Oliva and Michele De Prisco, from the research groups on Imaging of mood- and anxiety-related disorders (IMARD) and Bipolar and depressive disorders, published a methodological commentary in the journal Nature Mental Health that explains this problem using conceptual examples, simulations and an illustration with real data from 942 people with repeated measures of anxiety and depression. The problem, however, is the same in any medical speciality and whether or not AI is used.
Joaquim Raduà explains that “the work does not imply an immediate change in patient care, but aims to protect the dissemination of prediction tools, such as those based on artificial intelligence, which appear more precise than they really are, and thus prevent clinical decisions being made based on predictions that seem better than they actually are.”
The article proposes a simple procedure to check whether prediction tools provide genuine predictive information beyond this statistical phenomenon. The researchers propose applying this criterion in future studies, incorporating it into methodological recommendations on creating symptom prediction models, and reviewing the extent to which some published results may be influenced by this statistical artefact.
