Latent variable models are statistical tools used extensively throughout clinical and experimental research in medicine and the life sciences. Disciplines such as psychology and neuroscience routinely employ these models to answer complex research questions ranging from the impact of personality traits on workplace success (1) to measuring inter-correlated activity in neural populations based on neuroimaging data (2). Through latent variable modelling, dispositions, states, or processes that must be inferred rather than directly measured can be linked causally to more concrete, observable measurements.
While some latent variable approaches are exploratory, Structural Equation Modelling (SEM) and Confirmatory Factor Analysis (CFA) are primarily used to test a priori hypotheses. They are designed to evaluate the causal relationships between observable (manifest) variables and their corresponding latent variables within an inter-correlated dataset. A critical assumption of SEM is that the model is correctly specified. Even a minor misspecification can affect all parameter estimations in the model, rendering approximations inaccurate in unpredictable ways (3).
Therefore, with any postulated statistical model, it is imperative to assess and validate model fit before interpreting the results. The most rigorous and widely accepted method for evaluating exact fit across structural equation models is the chi-squared (χ2) statistic.
Interpreting the Chi-Squared (χ2) Statistic
A statistically significant χ2 statistic is indicative of the following:
- Model Misspecification: The degree of misspecification is a function of the χ2 value. The set of parameters specified in the model does not adequately fit the data, meaning the parameter estimates are likely inaccurate.
- The Need for Investigation: Because χ2 operates on the same statistical principles as parameter estimation, trusting the parameter estimates requires trusting the χ2 test. A significant result necessitates investigating where the misspecification occurred and readjusting the model to improve accuracy.
While incorrect hypotheses are a common cause, misspecification can also stem from other statistical violations. To properly diagnose a significant model fit test, researchers must evaluate the following assumptions:
- Heterogeneity: Does the causal model vary between subgroups of subjects? Are there intervening within-subject variables?
- Independence: Are the observations truly independent? Latent variable models assume that all manifest variables are independent after controlling for latent variables, and that an individual’s position on a manifest variable is solely the result of their position on the corresponding latent variable (3).
- Multivariate Normality: Is the assumption of multivariate normality satisfied?
The 2015 Meta-Analysis: A Cautionary Tale
A 2015 meta-analysis of 75 latent variable studies, drawn from 11 psychology journals, highlighted a troubling tendency among clinical researchers to ignore the χ2 exact fit statistic when reporting and interpreting their results (4).
The findings revealed that 97% of papers reported at least one “appropriate” model. However, 80% of these models did not pass the criteria for exact model fit, and the χ2 statistic was simply ignored. Only 2% of overall studies concluded that the model didn’t fit at all, and one of these interpreted the model anyway (4).
Why is
χ2 being ignored? Researchers often cite reasons for bypassing the exact fit test, including: the statistic is overly sensitive to sample size, it penalizes models with a high number of variables, or a general objection to the logic of the exact fit hypothesis. This has led to a broad consensus in favor of using Approximate Fit Indices (AFIs), such as the RMSEA or CFI, to justify models.
However, relying on AFIs often leads to questionable conclusions. In the meta-analysis, only 41% of studies reported χ2 model fit results. Furthermore, 40% of the studies that failed to report a p-value for their
χ2 did report the degrees of freedom. When researchers used these degrees of freedom to cross-check the unreported p-values, they found that all of the non-reported p-values were, in fact, significant.
Other glaring reporting issues included:
- 43% of studies failed to report which fit function (e.g., Maximum Likelihood) was used.
- 30% of studies selectively applied more lax cut-off criteria for AFIs than were conventionally acceptable.
- 53% failed to report their AFI cut-off criteria at all.
- Assumption testing for univariate normality was assessed in only 24% of studies (4).
Defending the Exact Fit Test
A common criticism of χ2 is that as datasets grow larger, the test detects increasingly trivial discrepancies as sources of misspecification. Strictly speaking, this is not a flaw it is an increase in statistical power. It means the level of certainty with which discrepancies can be considered important has increased.
Model misspecification can result from both theoretically relevant and peripheral causal factors. A significant model fit statistic is not trivial just because the underlying cause is trivial; rather, it means that a trivial cause is having a statistically significant effect that needs to be addressed. The χ2 model fit test remains the most sensitive way to detect misspecification in latent variable models and should be prioritised, even when sample sizes are high. Importantly, in the context of SEM, a rejection of model fit does not necessarily require the rejection of every individual hypothesis within the model (4).
The Problem with Approximate Fit Indices (AFIs)
AFIs provide a conceptually heterogeneous set of fit indices for each hypothesis. None of these indices are accompanied by a formal critical value or significance level, and nearly all arise from unknown distributions. While they are a function of χ2, unlike the χ2 statistic itself, AFIs do not have a verified statistical basis for testing model fit. Despite this, satisfactory AFI values are routinely used to override a significant, failing χ2 test.
Monte Carlo simulations of AFIs have concluded that it is impossible to determine universal cut-off criteria for any form of model tested. Furthermore, using AFIs, the probability of correctly rejecting a mis-specified model actually decreases as sample size increases – the exact inverse of the χ2 statistic. Additionally, as model misspecification or correlated errors become more severe, AFIs become increasingly unpredictable, whereas χ2 reliably reflects the severity of the issue (4).
Best Practices for Latent Variable Modelling
Based on the findings of the meta-analysis and the statistical principles of SEM, researchers should adhere to the following best practices, alongside rigorous testing of heterogeneity, independence, and multivariate normality:
- Pay strict attention to distributional assumptions.
- Have a strong theoretical justification for your model.
- Avoid post hoc model modifications (e.g., dropping indicators, allowing cross-loadings, or correlating error terms) just to achieve fit.
- Avoid confirmation bias.
- Use an appropriate and clearly stated estimation method.
- Recognize the existence of equivalent models.
- Justify all causal inferences.
- Use clear, transparent reporting that does not selectively omit model fit statistics.
Image: Michael Eid, Tanja Kutscher, Stability of Happiness, 2014. Chapter 13 – Statistical Models for Analyzing Stability and Change in Happiness. https://www.sciencedirect.com/science/article/pii/B9780124114784000138
References:
- Latent Variables in Psychology and the Social Sciences.
- Structural equation modelling and its application to network analysis in functional brain imaging. https://onlinelibrary.wiley.com/doi/abs/10.1002/hbm.460020104
- Chapter 7: Assumptions in Structural Equation modelling. https://psycnet.apa.org/record/2012-16551-007
- McIntosh, C. N. (2015). A cautionary note on testing latent variable models. Frontiers in Psychology. https://www.frontiersin.org/articles/10.3389/fpsyg.2015.01715/full



Really Great article, Amazing Write Up, I can agree with your point of view. Much appreciation for the information, Really interesting article, It’s well-structured and has good visual description, I would like to thank you for putting the time together to construct this article. It gave me a lot of information that I really enjoyed reading.