I think it was in one of my very first statistics lectures as a very enthusiastic 20-year old psychology student that my respect-demanding professor looked into every single student’s eyes and said: terminology is important! When you talk about statistical findings that fail to meet the criteria to reject the null-hypothesis, you must refer to them as “non-significant”, NEVER as “insignificant”. These are not the same thing! (quoted from my brain, with no claim for accuracy)
I am reminded of this, every time I read discussions of scientific findings, or even journal papers that get this terminology wrong. You may deem it nit-picky, but I am convinced that it’s important to know the difference.
The next paragraph is complicated and might make your head spin a little. Skip it if you want but don’t let it stop you from reading the rest of this post!
You see, what we do when we evaluate quantitative data (i.e. numbers that we have collected to describe a phenomenon), is we first (ideally) create a hypothesis of how two or more variables are related (on the basis of previous research), then collect the data, and finally calculate the likelihood of this data occurring under the assumption that our hypothesis is wrong (this is not intuitive, I know but it is due to the fact that we can never collect every single person’s data in the world so we cannot, unequivocally, ever, prove that our hypothesis is correct). Normally, we then accept a cut-off of 5% (p=0.05), that means, if our data is less than 5% likely to occur under the assumption that the null-hypothesis is correct, we regard the result as supporting our alternative hypothesis.
Are you confused yet?
Let’s use an example. We have read the literature and want to test whether socio-economic status (i.e. financial status) is related to depression risk. Our hypothesis is: Socioeconomic status is related to depression risk. That means our null-hypothesis is: socioeconomic status is not related to depression risk. We collect our data, run a test, and find a p-value of 0.001. That means, if the null hypothesis is true, our data is less than 5% likely to occur (0.01%, in fact). Our finding is significant. Both statistically and societally: We now know that it is very likely that depression risk is related to an individual’s ability to make ends meet, financially. This is an actionable point for politics, and health- and social care (described, e.g. in Elovainio et al, 2020).
Here comes the point about “non-significant” vs “insignificant”, however. I recently read a paper where, had I been a reviewer, I would have criticised the use of language. It summarised (among other factors) the association between older adults’ depression risk and education and found that, overall, most studies, and especially those deemed “good quality” (by objective metrics used for systematic reviews), found no association between education and depression risk, calling this result (statistically) “insignificant”. Statistically, this association was “non-significant”. However, this finding is anything but “insignificant”, right? It means that most studies evaluated here found that depression risk is likely equal across groups with different levels of education. That means, even though socioeconomic status might be related to depression risk in older adults, education might not be, those two factors are separate if this is true.
A similar example from my own research on older adults’ habitual use of digital devices to communicate: In one of my studies (Macdonald, Luo & Hülür, 2021), we examined whether there was an association between how older adults communicated with their social circle (i.e. face-to-face, on the phone, or digitally) and their well-being. We found no significant association, that means, it doesn’t matter which medium older adults use, social interaction, in general, has a positive effect on their daily well-being. So in practice, we should encourage and enable older adults to stay in touch with their social circle by whichever means possible. This is a non-significant finding which is very significant in practice (of course there are also limitations to this study that affect how we can interpret it – but you get my point).
There are also findings that are statistically significant but unlikely to be significant in practice. These are sometimes used to promote pseudoscience or wellness practices when studies show statistically significant results that might be due to inappropriate study design, scientific malpractice, or conclusions that go beyond the scope of the data. If there’s interest in more information on this part, tell me, and I might see if I can find some examples.
For now – pay attention to terminology, it might seem insignificant but it can make all the difference!
Sources:
Elovainio, M., Vahtera, J., Pentti, J., Hakulinen, C., Pulkki-Råback, L., Lipsanen, J., Virtanen, M., Keltikangas-Järvinen, L., Kivimäki, M., Kähönen, M., Viikari, J., Lehtimäki, T., & Raitakari, O. (2020). The Contribution of Neighborhood Socioeconomic Disadvantage to Depressive Symptoms Over the Course of Adult Life: A 32-Year Prospective Cohort Study. American Journal of Epidemiology, 189(7), 679–689. https://doi.org/10.1093/aje/kwaa026
Macdonald, B., Luo, M., & Hülür, G. (2021). Daily social interactions and well-being in older adults: The role of interaction modality. Journal of Social and Personal Relationships, 38(12), 3566–3589. https://doi.org/10.1177/02654075211052536
Maier, A., Riedel-Heller, S. G., Pabst, A., & Luppa, M. (2021). Risk factors and protective factors of depression in older people 65+. A systematic review. PLOS ONE, 16(5), e0251326. https://doi.org/10.1371/journal.pone.0251326
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.