Consider the following scenario. You want to know the distribution of hair length between adult men and women, but doing this for all world population is just impossible. You start to select four men and women. For every single person, there is bound to be differences in hair length which represents a within-subject variability. In addition, different people in our study will have different hair cut or style, a form of between-subject variability. You can sample a strain of hair of each person or multiple hairs.
- Fixed effects model: we assume that any systematic variability introduced in our measurement is due to within-subject variability only. In other words, the four men (or women) are more or less equal in terms of hairstyle.
- Random effects model: this model assumes a more realistic scenario where a man (or woman) is randomly sampled from a population. The model takes into account another factor called "subject", and any unpredictable variation in the data is contributed by an individual or between-subject differences.
In functional MRI, we employ mixed-effects modeling as accurate statistics when analyzing group data. This is done by taking sample voxel's time series per subject. Theoretically speaking, a fixed-effects model is used when a single subject has multiple runs. In order to assess multiple subjects that have multiple runs, we have to use mixed-effects modeling because every person is understood to be unique and we want to make a statement that can generalize to all people.
Fixed Effects Analysis
Suppose Yi = Xibi + εi is the fixed effects model of our neuroimaging data of N-subjects, each is acquired with n time points. The subject-level GLM consists of p parameter estimates. The within-subject residual εi has a multivariate normal distribution with a mean vector 0 and variance-covariance matrix of Σi = σw2.I. Under the null hypothesis, H0: cTbG = 0, we can then compute:
which obeys a t-distribution with # dof = N(n – p). Note how we incorporate the group-level PE from the average of individual subject's PEs, and within-subject error variance from the pooled individual error variance.
Random Effects Analysis
Our group modeling will now have both fixed and random effects. Suppose Yi = Xibi + εi is the model of our neuroimaging data of N-subjects, each is acquired with n time points. The within-subject parameter estimates and residual are bi and εi respectively. In the random-effects model, we further treat the individual subject as coming from a set of the population with its own between-subject variability. Thus, bi can be further decomposed into bG + εGi. the group parameter estimates and group error of which the subject belongs. Under the null hypothesis, H0: cTbG = 0, we can then compute:
which obeys a t-distribution with # dof = N – 1. Note that the numerator is the same as the fixed effects model, but the denominator contains more terms. The denominator in random effects is comparatively larger than that in fixed-effects model. So, if you use fixed effects for group analysis, you tend to get unnecessary activation, that is, a false positive.
Notes on Linear Mixed Model
The statistics mentioned in this post are also known by linear mixed model (LMM). With mixed modeling, the model will fit the average intercept and slope as a fixed effect, and each subject is assigned a different intercept/slope. Hence, LMM is an extension of GLM, but this is unlike a standard linear regression model which has only fixed effects. LMM approach is powerful to model data where there is a hierarchical or nested structure, or we deal with multilevel modeling. LMM can also be used in the case of repeated-measures design, or cases when there are some missing data.
Let's take another classic example of LMM: modeling distribution of performance of college students in New York City. Variation in the performance can primarily be due to a random variation at the individual (each student) level. But suppose the city has many colleges. Each college can also be seen as a contributing factor to the performance of each of the individuals at that school, say, due to the good teachers and attitude of the school. Hence those observations cannot be treated as fully independent of each other, but dependent on a higher level grouping called (which college?) - breaking the major assumptions of more traditional linear models. This example reflects different sources of randomness which are in a hierarchy, i.e. individuals-classes-college. This resembles our fMRI dataset, as we move up from an individual subject data to a group-based inference.
LMM is also common for two explanatory variables that are not independent. For example: suppose we want to predict the pitch level as a function of two independent variables: age and gender. Following GLM notation, we can write down:
Fixed Effects Analysis
Suppose Yi = Xibi + εi is the fixed effects model of our neuroimaging data of N-subjects, each is acquired with n time points. The subject-level GLM consists of p parameter estimates. The within-subject residual εi has a multivariate normal distribution with a mean vector 0 and variance-covariance matrix of Σi = σw2.I. Under the null hypothesis, H0: cTbG = 0, we can then compute:
Random Effects Analysis
Our group modeling will now have both fixed and random effects. Suppose Yi = Xibi + εi is the model of our neuroimaging data of N-subjects, each is acquired with n time points. The within-subject parameter estimates and residual are bi and εi respectively. In the random-effects model, we further treat the individual subject as coming from a set of the population with its own between-subject variability. Thus, bi can be further decomposed into bG + εGi. the group parameter estimates and group error of which the subject belongs. Under the null hypothesis, H0: cTbG = 0, we can then compute:
Notes on Linear Mixed Model
The statistics mentioned in this post are also known by linear mixed model (LMM). With mixed modeling, the model will fit the average intercept and slope as a fixed effect, and each subject is assigned a different intercept/slope. Hence, LMM is an extension of GLM, but this is unlike a standard linear regression model which has only fixed effects. LMM approach is powerful to model data where there is a hierarchical or nested structure, or we deal with multilevel modeling. LMM can also be used in the case of repeated-measures design, or cases when there are some missing data.
Let's take another classic example of LMM: modeling distribution of performance of college students in New York City. Variation in the performance can primarily be due to a random variation at the individual (each student) level. But suppose the city has many colleges. Each college can also be seen as a contributing factor to the performance of each of the individuals at that school, say, due to the good teachers and attitude of the school. Hence those observations cannot be treated as fully independent of each other, but dependent on a higher level grouping called (which college?) - breaking the major assumptions of more traditional linear models. This example reflects different sources of randomness which are in a hierarchy, i.e. individuals-classes-college. This resembles our fMRI dataset, as we move up from an individual subject data to a group-based inference.
LMM is also common for two explanatory variables that are not independent. For example: suppose we want to predict the pitch level as a function of two independent variables: age and gender. Following GLM notation, we can write down:
pitch ~ nationality + gender + errorNotice how age and gender are dependent as coming from the same person or subject. Taking into consideration fixed and random effects with LMM, we rewrite the equation to be:
pitch ~ nationality + gender + (1 | subject) + error