This is going to be a very brief conceptual discussion on hierarchical linear modelling (HLM), also known as multi-level modelling in social sciences. People sometimes call this mixed-effects modelling or mixed modelling, but be careful that not all mixed-effects models have a hierarchy.
Note: Mixed-effects models by definition are models which have both random and fixed effects (see below).
(1) When do we use it?In social sciences, the data types are usually hierarchical in nature, i.e. they have a certain hierarchy that is based on grouping or nesting. We use HLM precisely in cases where observations are nested or clustered in some ways at a different level. It has many applications in cross-sectional and longitudinal research (repeated-measures design). Research studies create data hierarchies through the way the data are sampled. For example:
- Students (S) are nested within a particular Class, which again is nested in a School.
- Time points are nested within participants in a longitudinal study.
![]() |
| Fig-1: Pictorial representation of (a) two-level hierarchy and (b) three-level hierarchy (Source: SAGE Ency. 2018) |
Suppose we wanted to examine A-level exam scores this year to predict the university entrance test in the following semester. So entrance exam is the outcome variable [DV] and the A-level score is the predictor [IV]. One way is to randomly sample students from the whole population such that each individual is guaranteed to be independent. However, what if the students recruited only belong to a particular high school? In HLM, the students are treated as the first level or the base, and the school as the second level (2-level hierarchy). In other words, we performed sampling primarily from four different local schools only. To advance further, we have to introduce 'Class' in between both levels (now, 3-level hierarchy), since each school surely has different classes.
In HLM, independence can no longer be freely assumed. Why? E.g. think of how students of Teacher John of School ABC have a different profile, intelligence, general ability than those attending the class of Teacher Mary from School XYZ. In such a nested condition, assumptions fulfilled for the OLS will be violated: observations are no longer independent of each other, and residual errors are correlated within clusters. The error variance will be different for different clusters. As a result, the standard errors of the regression coefficients will generally be underestimated and the significance level will be wrong, leading to misinterpretation. One way is to bring the students' data of Level 1 up to the next level through some averaging, but this is a poor practice. HM treats levels as something to take into account. It is basically an extension of the ordinary least square (OLS) of the linear regression.
(2) How to construct the model?
For a given A-level exam score xi in a simple linear regression, we estimate the intercept b0 and slope b1. The term ei is the residual. The equation states that for a unit increase in the A-level exam score, the score on the entrance test yi increases by the value of b1. But wait! The outcome dataset (entrance test scores yi) is clustered by the school. If this is not taken into account, then the data from all students is treated as unique observations.
Suppose we adopt a 2-level HLM comprising Students and Schools, i.e. the students are grouped or nested within a school. We introduce a group-level variation by changing the suffices, such that yij is the score on the entrance test for student i in school j, and uj is the group-level residual for that j-th school. This uj is a new random variable, referred to as the Level-2 or group residuals, and is assumed to follow a normal distribution, ~ N(0, σ2). The new HLM equation has the following terms:
- We see b0 as the grand mean of the university entrance test y.
- Level 1 contains the basic form of a linear regression, yij = b0j + b1xij + eij ; in turns, Level 2 consists of b0j = b0 + uj.
- The mean of y for the school group j is b0 + uj, where uj is the "school effect", i.e. the difference between the mean of the school group j and the grand mean. With this, each school can have different mean scores on the entrance test.
- The individual-level residual eij is the difference between the value of y for a student i and the individuals group mean b0 + uj. This reflects differences in students’ individual test scores from their respective school (or group) mean.
What can this improved model tell us? Some schools will have means that are higher than the grand mean, suggesting that their students perform better on the A-level exam on average, and some schools will have lower cluster mean values.
(3) Defining fixed and random terms
The core of hierarchical models is the assignment of fixed effects and random effects, two terms that already appeared in another blog post on group-level fMRI analysis. You have to define each term of a mixed-effects model to be either a fixed or random effect.
- A fixed effect is a parameter that is fixed across all groups and does not vary in the model. It is the term of interest in experimental manipulation. Estimating a fixed-effect of a term is like estimating a regression slope b. Differences in the slope If all terms in the model above are fixed, it behaves as the usual linear regression.
- In contrast, a random effect allows each group to have a different estimate. In HLM, random effects represent a higher level variable under which data points are grouped. This implies that random effects must be categorical (but cannot be continuous!). For example, the residual errors (uj and eij) are considered random effects. Estimating a random effect is like looking for an effect in our data to come from a large group of normally distributed datasets.
The most basic model has only an intercept without any predictors xij, which is also called the intercept-only model. Another name for it is the Unconditional Model or Null Model. Usually, the model has only participants eij as a random error.
The next basic model is the one shown in the equation above. It allows for different groups or schools j to have a varying intercept or mean b0j = b0 + uj. Thus, the model is also known as the random intercept model.
(4) The more complete model
A random intercept model shown above assumes that the relationship between the university test and the A-level exam score (the predictor) is the same for each group. This assumption can be relaxed by allowing for different slopes for the predictor in each group, making it a random slope model as shown in the right panel of Fig-2 below. Each slope will be estimated separately for each group. In this new equation, there is a new term u1j xij , where:
- The intercept for school group j is now b0 + u0j.
- The slope for school j is now b1 + uij , where b1 is the average slope across groups.
- Subscript “0” differentiates the random effect for the intercept u0j from the random effect for the slope u1j. Both random effects are assumed to be normally distributed.
The variance of the intercept and slope are assumed to be correlated; the covariance between the intercept and slope is estimated as part of the random slopes model. The random slopes model is also commonly known as the random coefficient model or a growth curve model when using repeated measures or longitudinal analysis.
(5) Final notes
How is the implementation in practice? Apart from defining fixed/random terms, we have to know which one nested under which variable.
- When we deal with a dataset we have to begin with the simplest model, which is the intercept-only model.
- Then we add additional terms, the simplest being the random intercept model, i.e. each group has its own group mean.
- If we have more than one predictor or IV, define the fixed/random effect. A mixture of fixed and random effect slopes is possible.
- The most complete model is the random slope model, where intercepts and slopes are treated as random effects. Often in a longitudinal dataset, we can add the quadratic term "Time" (Level-1).
So how do we know which model is best? We can use the usual performance metrics such as Bayesian Information Criterion (BIC) or Akaike Information Criterion (AIC). The two most common libraries for HLM analysis in R are lme4 and nlme. In the commands, you have to indicate which ones are the fixed and random effects, and which variable is nested within what. The outputs of the HLM in R is quite similar to what we expect from OLS in linear regression.
The random slope model often mimics reality, but this requires a large sample size. Parameter estimation in HLM is achieved using either the Maximum Likelihood (ML) or Restricted Maximum Likelihood (ReML). ReML works better if the sample size is small.
Some good references:







