Tuesday, November 8, 2016

Experimental Design - Statistical Tests

Basic inferential statistics.   I decided to separate this from the earlier entry. We shouldn't forget that what we observe is based on a sample, not a population. Inferential statistics has its goal in making a conclusion about a population from just this sample. In doing so, there bound to be some errors as our measure is just an estimate. Suppose we have a population of interest and that we draw all possible samples of size n. And suppose we compute a statistic e.g., a mean, proportion, standard deviation for each sample. The probability distribution of this statistic for all possible samples is called a sampling distribution; the standard deviation of which is called standard error which roughly tells us the spread around the mean of the sampling distribution. Mathematically, statisticians have to make some assumptions in order to model the problem accurately. This will be discussed later on.

How can we infer about a population parameter by using our sample? We use the concept of the sampling distribution (of means) to see where the true population mean lies w.r.t our sample mean. Having obtained the sample mean and standard deviation, we can find a confidence interval within which the true population parameter lies. This interval can be found with a certain confidence level, say 95% or 99%, with the help of a transformation. This data transformation converts our sample into a known sampling distribution, e.g. using a Student t-distribution, by first converting our scores into a t-value. The conclusion of our experiment is obtained by making use of a known sampling distribution and putting that distribution in a certain hypothesis test. The type of statistical test used (e.g. one-sample t-test, ANOVA, or chi-square test) depends on the nature of our variables.

In constructing a hypothesis, there are by definition the null and alternative hypothesis, indicated as H0 and H1. In the language of experimental design, the H0 states that the IV or treatment (marijuana) has no effect on DV (exam score). Conversely, the H1 states that the treatment has an effect on DV. To do, we first create a sampling distribution assuming the null, H0 is true. We decide an alpha level α or significant threshold to represent a concept of "very unlikely-ness" or chance that we observe a treatment effect when H0 is true. We then compare the likelihood of getting our sample statistics p to the alpha level that is usually fixed at 0.05 or 0.01. The final outcome of the test is whether we disapprove or retain the null hypothesis. For example,
  • The probability of our sample statistics of 0.01 means there is only 1 out of 100 chance you find no treatment effect in your data when the null H0 is true  ..... We reject H0.
  • The probability of our sample statistics of 0.80 means there is 8 out of 10 chance you find no treatment effect in your data when the null H0 is true  ..... We retain H0.
  • Rejecting the null while in reality it's true is a Type I error with a chance α. Accepting the null while in reality it's false is a Type II error with a chance β. Note that these errors are not measurement errors made by mistakes but rather a measure of chance or probability.
Fig 1: Summary of hypothesis test with its associated Type I/II errors. 

Behavioral science is unique in the sense that the test must first assume the null hypothesis to be correct. Also, the test result simply makes a case without even saying which hypothesis, H0 or H1 is indeed true or factual. Even if H0 is indeed true, we don't know how much the difference is with the untreated sample group. For such reason, the effect size d comes into play. The effect size is defined as the difference in means between the experimental group and control group, divided by the standard deviation of the control group.

Lastly, the ability to correctly reject the H0 or find something significant in your data while H0 is indeed wrong is called power, and it is simply 1 – β. A powerful design is able to detect a small effect size < 0.20 unit. Power increases or β is reduced when the alpha level increases, the mean difference between two conditions under the null hypothesis increases, and sample size n increases. The concept of power is crucial because it determines the replicability of our study. According to Cohen (1966), most of the behavioral studies have low power, i.e. the chance of replicating the same experiment yielding the same result is 50-50.

A technique called power analysis is used to determine the appropriate sample size for the study. Generally, power = 0.80 is an acceptable value. Tools such as "pwr" in R is able to help us in power analysis, e.g. find the minimum sample size to achieve certain p-value and power of 0.80.

Okay, we move on to talk about assumptions in statistics, a crucial component to make our problems mathematically solvable. Most statistical procedures such as t-test, linear regression, and ANOVA are based on assumptions. For example, a linear regression is valid when the DV and IVs are linearly related to begin with (for categorical variables such as "Yes/No", one uses logistic regression instead). Regression has other important assumptions in general, e.g. the residuals are independent, normally distributed, and their variance is homogenous. The type of test that requires certain assumptions to be met is called parametric statistics. Violating these assumptions causes incorrect inference or decision. The assumptions for parametric tests are presented below.

[Reference #1]: Among some relevant textbooks, I chose the materials from the classic JL Myers' "Fundamentals of Experimental Design" (1979).

The normality assumption.   As mentioned, statisticians make assumptions on their methodology to simplify the maths. Both t-test and ANOVA have to satisfy the following main assumptions:
(a) Samples are independent of each other; the score of Tom is independent of Ni;
(b) The normality assumptions (for ANOVA, it means the distribution within groups);
(c) The homogeneity of variance.
The normality assumption means that the sampling distribution comes from a population that is normally distributed or Gaussian. This is hard to know as we don’t have the luxury of knowing all possible samples related to our population of interest. Out of faith, we assume that if the data we take is normally distributed then the sampling distribution will also be normally distributed. Every stats guru also tells us about the central limit theorem that is, the distribution of the sample means approaches a normal distribution as the sample size N increases, regardless of the shape of the population of interest. As a rule of thumb, N > 30 is generally accepted for the assumption to be valid. The average of our sample means is itself the population mean, and the standard deviation is the standard error.

Fig 2: Comparison between normal data and its QQ-plot (left) and "not-really" normal data (right)

This can be tricky if N is as few as 15 or 20. Now, how do we test for normality?
(1) Visual comparison. Normal data has a bell-shaped distribution. It has a specific kurtosis and skewness characteristic. It is recommended to transform the kurtosis and skew values to z-scores. Another useful visual test is through using Q-Q plot that compares theoretical values against the actual or observed values. An ideal case will be a straight line with 45-deg w.r.t the horizontal. The following figure shows an example how the data that is less normal has the tendency to have Q-Q plot that is not straight, but rather curved away from the perfect straight line. Most of the times, the presence of outliers distort the normality, the skewness or kurtosis of your data. It is quite common to treat outliers by removing the particular data point or by transforming the data.

(2) The Shapiro-Wilk test. In R, you can simply write shapiro.test(variable name); or in Matlab as swtest(x, alpha). This test is super cool as it directly assesses whether a certain distribution is significantly different from the normal distribution with the same mean and standard deviation. If the p is significant (p < 0.01, say), then the data is different from the normal distribution.

Homogeneity of variance.   Suppose you have three groups in your experiment. Homogeneity of variance says that the variability of the data in all groups should be equal, because strictly speaking, they come from the same population. The same notion applies, for example, in the case of repeated measures where you test the same group in two different experimental manipulations. If the homogeneity principle holds true, the variance of the data in either case remains rightfully the same. In most cases, we usually ignore this principle and immediately perform t-test or ANOVA, etc. Refer to the following illustration.

There are two tests that you can use to assess this: Levine’s Test and Hartley’s F-test (good for large sample). Both methods test the null hypothesis that the variance of the two datasets or groups are equal. The Hartley's F-test is also called the equality of variance test, or sometimes variance ratio test since you're actually taking the ratio of. The computed F-statistics will be compared against the standard F-table according to their respective degree of freedom. In general, people have shown that ANOVA is robust against a mild non-normality and heterogenous variance (only for equal sample size among groups!).
Fig 3: Comparison between dataset with homogenous (left) and non-homogeneous variance (right)

Non-parametric statistics.    Non-parametric statistics means the methods don't rely on any assumption discussed above! In fact, the information on the distribution is obtained directly from the dataset itself. This is the most fundamental idea of non-parametric statistics! Bootstrapping, for example, is one popular non-parametric test that relies on "random sampling with replacement". This test can be used to assess stability of your statistics and it is pretty straightforward. Another test called permutation test builds, rather than assumes, the sampling distribution by 'shuffling' the observed data. Unlike bootstrapping, we use only the existing data set and the shuffles are without replacement.

I took this from Matlab website. Suppose you have a set of LSAT scores and  GPA. You can easily plot the data and compute the correlation. It is found that there is a positive relationship between LSAT and GPA, r = +0.78. Now although it may seem large, it is unknown whether this observation is statistically significant or just happens by chance. Using the bootstrp function you can resample the two vectors as many times as you like and consider the variation in the resulting correlation coefficients. If the finding is true, you would expect that bootstrapping process yields more or less the same correlation value. Refer to the following histogram. Bootstrapping gives a distribution of correlation coefficients that is centered near r = +0.78. Using permutation strategy, we can also show how the correlation found is not by chance. We permute the LSAT-GPA pairs for 1000 times, each time is followed by computing the correlation. The correlation values of the permuted data sets should collapse since the relationship has been "broken" following shuffling.
Fig 4: Distribution of correlation coefficients with bootstrapping and the location of the original r = +0.78 (red line).

During the period when I first learnt statistics, we didn't care much about such assumptions and treated them as dogma. Now, I'm learning to accept the presence of non-parametric statistical tests used when the assumptions are not met, that is, to replace the usual t-test etc. The principle behind those tests is assigning ranking. High scores will be represented by large ranks, low scores by small ranks. The analysis is then performed on the ranks rather than the original raw data. In other words, non-parametric tests are useful when dealing with ordinal data. Example of ordinal dataset includes: the satisfaction rating from 1 - 5, the order of athletes based on speed, and the school ranking.

When we are interested in comparing two independent datasets, we need to use the Mann-Whitney test or the Wilcoxon's rank-sum test. This is the counterpart of the independent t-test. In R, you can do this by:  wilcox.test(x,y,____), some parameters inside the brackets have to be defined. In Matlab, use the following instead ranksum(x,y). When we deal with repeated measures or paired sample, then we have to use the Wilcoxon signed-rank test. For example: testing the performance before and after drug consumption. You can use the same command as in (1) in R but by defining "paired = TRUE". In Matlab, however, the command will be different now, signrank(x,y). If we are interested in testing more independent groups, we have to do non-parametric version of the regular ANOVA, e.g. Kruskal-Wallis test and Friedman's ANOVA for repeated-measure designs. Other non-parametric analysis such as the Spearman's correlation coefficient (the counterpart of Pearson's r) and Kendall tau coefficient.



Final note: ANOVA can be understood from the perspective of Linear Modeling and in itself is a very huge topic, covering multiple regression, mixed-effects or hierarchical modeling, and generalized linear models. The mathematical discussions on this theme will be beyond this blog post.....

[Reference #2]: Andy Field et al, "Discovering Statistics using R" (2012).

Wednesday, November 2, 2016

Experimental Design - Basic Concepts

Terminology.    The main goal of any scientific research is to describe behavior, predict the properties, determine the main cause, and explain it as part of a scientific question. To answer a research question, we traditionally follow the following steps: begin with a hypothesis, identify what factors whose effects are to be studied, variable(s) that are manipulated in the study, carefully select a measure of effects, the observed variable; as much minimize the unwanted or irrelevant effects or confounds, design an appropriate experiment and do sampling. Lastly, we perform statistical analyses and draw a conclusion.

There are 4 types of variables according to their nature: nominal (non-numeric, category), ordinal (rank, or ordering), interval (no true zero, e.g. temperature), and ratio scale (zero means absence, e.g. weight). These types exist also according to the way we measure them, or their level of measurement. The first two are qualitative, i.e. kinds, types, categories. The last two are quantitative variables or quantity, amount.

From the perspective of an experiment, there are 3 main types of variables. We have come across the dependent and independent variables a few times, including when talking about linear regression. Independent variables are under the control of the experimenter, independent of the participants. Dependent variable(s) is usually measured. A design that has only one DV is called a univariate design, more than 1 DV simultaneously analyzed is called a multivariate design. The following table nicely summarizes the differences. The third type, the confounding variable, is undesirable because it seems to explain our research question, but the actual fact is irrelevant.

Independent variable (IV): a variable that you are manipulating, a.k.a
- Explanatory variable
- Regressor
- Controlled or manipulated variable
- Predictor variable
- Input variable
Dependent variable (DV): a variable that is affected in the experiment, a.k.a
- Explained variable
- Regressand
- Measured variable
- Response variable
- Outcome or output variable

We then design an appropriate experiment to test our hypothesis, determine how many groups of subjects, and wisely sample our subjects from a population of interest. There are a few ways to do sampling, e.g. random sampling, stratified sampling. The independent variables are manipulated to provide a different experimental procedure to samples or subjects. These different procedures are called treatments, and hence independent variables are also called treatment variables. In designing our experiment, there are two main ways we can assign the participant into. Refer to the table below.

Within-subject design, repeated-measure design = the same participant experiences multiple treatment or protocols in the study.
- Lower error variance, 
- Higher statistical power,
- Risk of carryover effect or order effect,
Counterbalancing and/or Latin square method is able to overcome this risk.
Between-subject design, independent-group design = participant is randomly assigned to only one treatment or one level of the independent variable of the study.
- Higher error variance (inter-subject var.)
- May require a larger sample, time & cost.

A simple example is here. Suppose you want to test the effect of marijuana on the final exam score. The independent variable here is the consumption amount. The dependent variable is the final exam score. Here, it is good to have 2 independent or treatment groups of subjects, i.e. no single subject goes through the experiment more than once. The obvious statistical comparison is between the group of students who take marijuana and another group who doesn't. In an ideal experiment, the only difference between the two groups is the treatment (marijuana consumption). In real life, this is hard to achieve. Irrelevant differences must then be controlled. A controlled experiment is taken exactly to minimize those differences.

What about confound variables? An obvious confound may be the student IQ. If this variable is able to explain the exam score well, this result is not directly useful to our study. In other words, our design should only reflect the influence of our independent variable on our dependent variable. Sometimes, a control group is used to minimize the confounding effects in our analyses. What we can do now is to compare control vs marijuana and control vs no marijuana. The presence of confounding effects and other subjective influences can also be minimized by performing randomization as a form of control during the experiment. For example, we can assign each participant to a group randomly regardless of their IQ level as the focus here is on marijuana consumption.

Another aspect to understand is the term fixed and random variable. A fixed variable whose different levels are have been arbitrarily, but rationally, chosen. Inferences taken about a certain fixed variable is limited to the selected members. A random variable has levels that are subsets of a larger population, bearing in mind that each level has an equal opportunity to be selected. It helps us to link our data to a population in general. Subjects are typically a random variable. There is always a risk that our experiment introduces a bias in our data. This may occur simply because of not selecting the subjects wisely, error in measurement, and finally, the confounding variables not accounted for. It is important that the dependent variable is sensitive and reliable to describe our hypothesis.

Cause and effect are important to explain certain behavior. Saying the average exam scores between these two groups are significantly different from each other does not directly say that drug consumption is the main cause of the effect. In addition, correlation and regression analyses do not imply causation. Traditionally, one has to prove that marijuana consumption is the necessary and sufficient factor for the effect to occur. But in reality, this is difficult to prove simply because there are many confounding factors that may explain an outcome or result. One way is to construct a chain of reaction and apply mediation analyses.

One critical factor in behavioural research is the validity of the variables being investigated. Validity means truth, how accurate a variable is to represent information. There are a few kinds of validity:

(1) Construct validity: adequacy of the operational definition of a variable. In social sciences, a construct is a measure we are interested in. Unlike in physics, constructs are often abstract or subjective. For example: if a social anxiety measure for teens is the one we are researching on, then that is our construct. Construct validity says that the measure has to describe or quantify social anxiety, not other things such as mood or temperament. 

(2) External validity: describes how a conclusion of a study can be generalized to other populations or similar studies. In our example using marijuana consumption on campus, how true does this happen across different schools in general? Hypothesis testing is performed to achieve this purpose. 

(3) Concurrent validity: when validity is tested against an established measure of the same or a related underlying construct assessed within a similar time frame, then we are looking for this type of validity. 

(4) Internal validity: refers to how well the effect in our experiment is attributed to our independent variable. A good experimental design must ensure this is the case. 

(5) Ecological validity: not only does it apply in the laboratory context, but it also has to accurately reflect what we want in the real-world setting. This is important as some measures, e.g. in sports or healthcare, require subjects to be tested outside of the laboratory. 

Finally, the results of our trials or experiments have to be reliable when we perform measurements or observations. A reliable observation is one that is consistent and repeatable. Consistency can be measured in terms of inter-rater reliability (using a few judges, measure once) or test-retest reliability (using a single judge, measure a few times).

Note: For more discussion, see Cozby (2011).

Factorial Design.     So far we only see one simple treatment, e.g. drug consumption, in our study. Let's introduce a new term: an independent variable is also called a factor, often a categorical variable. For example, when we talk about drug abuse, the drug type is a factor. Each substance: alcohol, marijuana, nicotine, and amphetamine is called the level of that factor.  In our experiment, if in addition to drug consumption, we introduce another independent variable in our example, e.g. criminal history and study hours we have two factors in the design. Proper design allows us to explain behavior in a multi-dimensional view.

Often, an experimental design can be a little complicated. Suppose we have 2 factors. If every level of factor-A occurs with every level of factor-B, it is called crossing. A study that has many levels or multiple crossing factors is termed factorial design. We can construct a factor A x B table. In general, a factorial experiment is typically the case. However, often factor-A may contain different level of factor-B. This is called nesting, factor-B is nested in factor-A.

Which statistical test should we use? Comparing two means in your data can be tested using t-test, be it paired or independent samples t-test. In a factorial design, when dealing with one factor with two levels only, the t-test is sufficient. When more levels are involved, however, the Analysis of Variance or ANOVA is used as an extension of t-test. ANOVA is based on F-test, which is a ratio of two variances (or mean-squares, MS). Refer to the equation below. When the null hypothesis is true that there is no difference in means, the F-value should be 1.0. A large F-value indicates some differences may exist among the means due to a systematic treatment or manipulation effect. Note that it doesn't tell which means are different! To find which means, we should do post-hoc tests.
There are two main types of ANOVA: (1) a 1-way analysis compares levels of a single factor based on a single continuous response variable; and (2) a 2-way analysis compares levels of two factors on a single continuous response variable; and so on.... Now, it seems that study hours, a continuous independent variable, is correlated with the exam score. To treat this in ANOVA, we use a variant analysis called Analysis of Covariance or ANCOVA.  ANCOVA compares the dependent variable by both a factor (categorical) and a continuous independent variable. ANOVA characterises between-group variations due to a treatment effect. In contrast, ANCOVA divides between-group variations to treatment and a continuous covariate.

Knowing what statistical test to use is paramount in research!! Several other hypothetical situations and the type of possible statistical tests used will be presented here:
To see the difference in exam score between Harvard and UPenn students in their "Intro to Statistics" class, I recruit a total of 30 first-year undergraduate students from each university. Then, I perform unpaired t-test (nondirectional or 2-tailed, with α = 0.05) as the samples come from two independent groups or origins.
To see whether there is a difference in exam score among Harvard, MIT, and Stanford students in their "Intro to Statistics" class, I recruit a total of 30 first-year undergraduate students from each university. Then, I perform 1-way ANOVA to test the main effect of university (2-tailed, with α = 0.05).
To see whether there is a difference in exam score among Harvard, MIT, and Stanford students in their "Intro to Statistics" class, I recruit a total of 30 first-year undergraduate students from each university. I introduce a new factor called gender (male and female). Then, I perform 2-way ANOVA to test the main effect of university and of gender. If the score of different university is also influenced by gender, I'm looking at the interaction between the two factors. Interpretation of the main effects can be different in the presence of an interaction!
To see whether there is a difference in exam score among Harvard, MIT, and Stanford students in their "Intro to Statistics" class, I recruit a total of 30 first-year undergraduate students from each university. It seems that the score is correlated with IQ level. If I treat the IQ level as a continuous independent variable or covariate, then, I perform 1-way ANCOVA to test the main effect of university (2-tailed, with α = 0.05) having controlled the IQ level as the covariate.

[Reference #1]: Paul Cozby, "Methods in Behavioral Research, 10 ed" (2011). 
[Reference #2]: JL Myers' "Fundamentals of Experimental Design" (1979).

Friday, July 29, 2016

The Reward Circuit - a brief overview

My current research project makes use of positive feedback as a form of reward or incentive in learning. This is actually exciting as it combines motor learning processes and positive reinforcement. Reinforcement and reward-based learning implicate the limbic network, in which the basal ganglia, BG, is one of the most important members. Specifically, it is the ventral portion of the striatum called the nucleus accumbens (NAcc) and underlies a series of behavior e.g., motivation, and emotion. Initially thought as purely motoric, BG evolved to engage more diverse behavior such as executive, and then the limbic functions. The inclusion of the limbic region as a part of the BG has also been proven anatomically (Nauta, et al., 1978; Mogenson et al., 1980; Haimer et al., 1986). More importantly, the BG limbic circuit does not work in isolation. This summary will be based on an excellent review by Haber and Knutson (2010) and a few other relevant good stuffs will be provided as references at the end.

Neuroanatomy of Reward
Prefrontal Reward Regions
The involvement of prefrontal cortex in reward comes naturally. Scientists have been interested in studying the role of the frontal lobe in cognition in which reward processing is one of the main features. The main reward-based prefrontal regions traditionally include the anterior cingulate cortex (ACC) that includes BA 24, 25, and 32 and the orbitofrontal cortex (OFC) that includes BA 11, 12, 13, and 14. Unlike sensorimotor cortex, the prefrontal cortex is diverse in terms of cytoarhitectonic features and functions. The recent paper in PNAS (Neubert FX, et al, 2015) scrupulously describes similarities and differences between primate and human prefrontal cortex associated with reward. In general, the human prefrontal cortex is sub-divided into a few areas:
  (a) Sensory region: the orbital part of the brain linked to the olfactory bulb and the insula.
  (b) Ventromedial prefrontal cortex (vmPFC) that includes BA 10, 11, and 32.
  (c) Rostral OFC that covers BA 11, 12, and 13.
  (d) Dorsal ACC or dACC that is BA 24.

Reinforcement-based motor learning presumably implicates basal ganglia (Schultz et al., 1997; Graybiel, 2005). In fact, reward-based action may involve more complex neuronal processes beyond the traditional basal ganglia and sensorimotor loops. For example, it is possible that reward-based decision-making is also involved (Rushworth et al., 2004), such that during learning rewards may influence the production of subsequent movements. Prior studies show that regions in the prefrontal cortex are involved in this type of activity (Shima and Tanji, 1998). Using fMRI in gambling tasks (Daw et al., 2006), it has been shown that the intraparietal sulcus and frontopolar cortex are preferentially active during exploration. In contrast, regions of the striatum and ventromedial prefrontal cortex are involved in exploitative decision making to accumulate more rewards. The vmPFC is a region in which activity is associated with stimulus-reward value, selecting actions that are more rewarding (O'Doherty et al., 2003; Rushworth et al., 2004; Daw et al., 2006) and encoding the value of performed decisions (Knutson et al., 2001; Smith et al., 2010).

The Ventral Striatum
In 1954, Olds and Milner discovered how a tiny structure called the NAcc that lies ventral to the sensorimotor striatum is linked to reward behavior in rats. In 1978, Heimer described the link between NAcc and olfactory tubercle in rats. Ventral striatum was later regarded as the reward center and thought to be the interface between the limbic and motor systems (Mogenson et al., 1980). In recent years, the traditional boundaries of ventral striatum have expanded beyond NAcc which includes the ventral portion of the striatum and the ventral caudate. Thus, the name ventral striatum (VS) is legit to contrast this with the more dorsal, sensorimotor-related striatum. The NAcc structure has an outer shell and a core, each contains different neuronal cell types and functions.

(1) Afferent projections
Like the dorsal sensorimotor striatum, VS receives massive topographic glutamatergic inputs from the cerebral cortex, from the thalamus, and the brain stem. The word topographic deserves an emphasis, refer to Fig-2. From the prefrontal reward regions, vmPFC sends projections mainly to the NAcc. From NAcc, we go more dorsally and laterally to regions that cover ventral caudate nucleus and ventral putamen. Principally, these two areas receive projections from the OFC. The dACC also projects to a more central and lateral portion of caudate and putamen. Lastly, the DPFC (dorsal prefrontal cortex) terminates more diffusely along the rostro-caudal striatum, particularly the head of the caudate. A critical difference with the dorsal striatum, VS also receives projections from the amygdala and hippocampus. The afferent axons are concentrated within the NAcc. Amygdala is known to play a role in processing reward, e.g. an emotional aspect of reward, associating stimuls and reward or punishment. Thalamus is the final link connecting VS with the rest of the brain regions. In the thalamus, reward processing is managed by one of the largest nuclei there called the medial-dorsal nucleus (MD) that carries bidirectional wiring between the frontal lobe and VS, amygdala, and hippocampus.

(2) Efferent projections
Like the dorsal sensorimotor striatum, VS sends efferent axons to two main targets. First, the more specific efferent outputs to the ventral pallidum, not GPe or GPi. Second, more broad midbrain regions such as the ventral tegmental area (VTA) and middle substantia nigra (SNpc) which are rich of the dopaminergic neurons. The pathway between midbrain and VS is known as the mesolimbic pathway and is known to mediate reward processing. Other efferent projections are to the pedunculopontine nucleus (PPT) and nucleus basalis (NB, the principal source of cholinergic fibers to the cortex and amygdala). In particular, PPT carries broad functions such as arousal, attention, motivation, and voluntary limb movements.

Midbrain dopamine neurons
The involvement of dopamine in reward is perhaps first shown elegantly by the work of W. Schultz since 1970s. The VTA and SNpc, the substantia nigra pars compacta, are two important components and their neurons can be identified using a certain phenotypic marker, e.g. calbindin for dorsal SNpc and VTA (the ventral SNpc, however, is calbindin-negative).

(1) Afferent projections
The principal inputs to VTA/SNpc come from the striatum through both the GPe and VP. Other inputs come from other parts of the brain stem, mainly PPT. Anatomically, the largest source of projection comes from the ventral striatum.

(2) Efferent projections
Midbrain dopamine neurons send massive projections out back to the striatum through the mesolimbic pathway mentioned above. It has been found that there is a medio-lateral and an inverse dorso-ventral topography arrangement. Thus, the ventral portion of SNpc projects to the dorsal striatum, the dorsal portion projects to the ventral striatum. It has also been observed that there is differential efferent projections back to the striatum. In other words, VS receives the least number of projections while the sensorimotor striatum receives the most. This striato-nigro-striatal network is consistent with the idea that the limbic system is able to influence sensorimotor behavior through the interface situated in the striatum. Apart from the striatum, dopamine neurons from the midbrain send rather diffuse projections to the frontal lobe through the mesocortical pathway, so it is wrong to say that BG is the only target of dopamine in the brain.


Further Readings
[1]  SN Haber, Knutson B. (2010). "The reward circuit: linking primate anatomy and human imaging". Neuropsychopharmacology.
[2]  FX Neubert, et al. (2015). "Connectivity reveals relationship of brain areas for reward-guided learning and decision making in human and monkey frontal cortex". Proc. Nat Acad. Science.

Thursday, June 30, 2016

Linear Regression - The Basics and More

Sigh, I should have studied this even before learning about the fMRI statistical pipeline. In itself, the scope of regression and GLM is super wide. In this blog post, the emphasis would be on concepts, not calculation, since statistical software such as R/SPSS/Matlab is able to help us.

Simple Linear Regression
Linear regression is one of the most basic tools to model data. In its simplest form, the model can be represented by Equation (1). This is because the equation can be used to find or “predict” the observed variable Y based on X, and so independent variable X is also known as a predictor. The point where the line crosses the vertical Y-axis is called intercept β0 of the model. The parameter beta or β is also called the gradient or slope of the equation. The slope represents how much the Y changes resulting from a unit change in X. Both coefficient β and β0 are called regression coefficients. If there is more than one predictor, the term multiple regression is used instead, see Equation (2). Here, we attempt to predict Y based on several predictors such as in the case of functional MRI statistical analysis. Using multiple regression performed at every brain voxel altogether, I have discussed how a task-activation map can be produced.
(1) How can we estimate the coefficients?
Performing a linear regression will also mean finding the best fit model or Ŷ (the ^ hat sign indicates that the variable is an estimated or predicted value). We position or "fit" the line where the deviation from each data point is the least. The regression line used to model the actual data set does not obviously pass every data point. Whereas variance is like the deviation from the data center (mean), deviations from the regression line are called residuals. A method. called the method of least squares, which aims to get a numerical estimate of β that minimizes these deviations or residuals. The Gauss-Markov Theorem has been presented when talking about fMRI statistics. The estimated beta is sometimes named parameter estimate, PE. Briefly, the formula is shown in Equation (3). Assumptions are made for this formula to be valid. In fact, for an accurate model, several assumptions have to be held:
  • The outcome variable has to be quantitative, continuous, and unbounded.
  • All predictors must be quantitative or categorical (at least two categories). They should not have zero variance. 
  • The relationship between predictors and the outcome variable is linear.
  • Residuals Σe = σ2.I are uncorrelated and independent of each other. They are normally distributed with zero means.
  • Residuals maintain homogeneity of variance.
(2) How well does the model fit my data?
The model performance can be quantified by using three different sums of squares defined as follow:
  1. Residual sum of squares, SSR = deviation of the observed data from the fitted line. This is actually an error or a measure for the inaccuracy of the model. Hence, it represents a portion of the data that cannot be well predicted by the line. This portion is minimized through the method of least squares. Refer to the red solid line in the figure above.
  2. Model deviation sum of squares, SSM = deviation of the predicted Ŷ based on the model from the mean of the data set. Refer to the green solid line in the figure above.
  3. Total deviation sum of squares, SST = deviation of the observed data Y from their mean. This is simply the variability of the data set around its mean. 
Variance partitioning says that SST = SSM + SSR . The total variance in our observed data can be decomposed into two parts: the portion explained by the model and by the residuals. Hence, SSM signifies the merits of using the regression line. The ratio of SSM  and SST is called the R2, or r-squared. It is a statistical measure of how close the data are to the fitted regression line. Generally speaking, between 0.16 ‒ 0.25 is considered a good model fit. There is also another measure called the F-ratio which is defined as follow:
F-ratio is a ratio between how good the model is with how bad it is (residual). Mathematically, it is the ratio between mean-squared error of the model MSM, and mean-squared error from the residual, MSR. F-ratio is used a lot in analysis of variance (ANOVA) and the significance of the value can be assessed using a F-distribution table.
Actually, a mean-squared error is the sum of square divided by a degree of freedom (dof). For a simple linear regression with just one predictor, the model has # dof equals 1 and the whole data set has N ‒ 1. The # dof of the residual is, therefore, N ‒ 2.

Residuals are important to identify poor model fit. Mathematical software such as R is able to give us a summary of the linear regression, lm. Residual can be quickly computed using resid. The sum of all residuals and the expected value E[res] equals zero. Lastly, residuals can be thought of as outcome variable y having removed the linear relationship with the predictor x.
(3) How well does the predictor do the job?
To evaluate further, we can also test whether X is a good predictor of Y. Because the slope β represents how much Y changes for every unit of predictor X, a lousy predictor means a flat or very small slope. Using one-sample t-test on the slope β (technically on the parameter estimate of that), we test whether it is significantly different from zero. Refer to Equation (4) in the context of simple linear regression. The denominator is the standard error that tells us the variability β's across different samples and σR2 is the residual variance. Commonly, a significant p-value of the t-test suggests that the predictor is reliable enough to predict the outcome variable.

Multiple Linear Regression
When there is more than one predictor, multiple regression kicks in. For example, we want to model or predict electricity bills during winter as a function of outdoor temperature and oil price. The basic principle remains, R2 represents the proportion of the variance explained by the model with the total variance in the observed data set. Note that more predictors means higher variance explained by the model, artificially inflating the R2. Some adjusted measures based on AIC, BIC, etc. are used as parsimony adjusted measures that penalize you for adding more independent variables in the model unnecessarily. As a rule of thumb, higher AIC/BIC means worse fit. Adjusted R2 tells you the percentage of variation explained by only the predictors that actually affect the outcome, dependent variable. It is computed using the formula 1 – ((1 – R2)((N – 1) /( N – k – 1)) where k is the number of predictors. The more junk variables, the lesser the value will be.

Fig-1: Two R outputs from a linear model operation where you have one predictor (IV2) and three (IV1, IV2, IV3) to predict the same DV. Click to enlarge! SPSS, like R, is able to produce similar detailed regression results.


In Fig-1, I want to predict DV using one (left) and 3 predictors (right panel). The beta coefficients are shown together as "Estimate" with their standard error, t-statistics, and p-value results. In R, these p-values are two-tailed. On the left panel, a positive slope means that there is ~0.27 increase in IV1 for 1 unit increase in DV. A very small p-value indicates that we cannot retain the null hypothesis that the β equals zero. Also, the summary tells us about the residual component, the multiple and adjusted r-squared. Look at how adding 2 extra predictors reduces our degree of freedom to 19, but does not give many benefits to the model. First, the residual is not dramatically reduced. Second, the left panel has a greater reduction in adjusted r-squared. The last line tells us about the F-statistics together with their p-value. A good model should have a high F-value with a low p-value against the null hypothesis. For a simple linear regression on the left, the p-value of the F-statistics equals the p-value from the t-statistics. This is unlike the regression with 3 predictors on the right.
In the context of multiple regression, we can use F-statistics similar to that used in ANOVA. The null hypothesis is that all βi equals zero. The alternate hypothesis is that at least one βi > zero. The F-ratio formula doesn't change, but the degree of freedom of the model and residual would now be k and N ‒ k ‒ 1 respectively, where k is the # predictors.
In multiple regression, the statistical test on the betas can be further analyzed according to different contrast. Such term may be familiar in ANOVA. For example, if there are 3 predictors, the contrast of [~1 0 +1] attempts to test whether the difference in the first and last β values is significant.

(1) Parameter estimates in multiple regression
How to find β-parameters in multiple regression? Exactly what we learnt in the fMRI statistics. We can derive betas by minimizing ∑(Y ‒ X1β1 ‒ X2β2 ‒ ... ‒ Xk βk )2  just like before. The formula for R2 and F-ratio are similar but DOF of the model and residual are different. Conceptually, the estimate for β1 is the regression through the origin estimate having [X2, X3, ... , Xk βk ] regressed out of both Y and the X1. Similarly, the estimate for β2 is the regression through the origin estimate having [X1, X3, ... , Xk βk ] regressed out of both Y and the X2. Regression through the origin happens when Y = 0 for X = 0.

The left graph below depicts the first model with IV2 only, the slope is +0.2712. The second model has IV1, IV2, and IV3 to predict DV. Here, I'm plotting IV2 and DV but with the influence of IV1 and IV3 removed. It can be seen that the slope becomes negative, ‒0.382776. Hence, the meaning of coefficient in multiple regression is the expected change in the outcome or dependent variable per unit change in one predictor, holding all other predictors fixed. By holding those predictors constant, it is said that we have controlled for other predictors in the model. Each regression coefficient is also called partial regression coefficient. Note: italicized words in brown carry similar context!
Fig-2: The slope between IV2 and DV is positive (left). When two other predictors are introduced, the same slope changes to negative (right). The result in the right panel is when we partial out IV1 and IV3 from the model. In this way, we say we have "adjusted" the model that only IV2 contributes to DV.


If we are to make a model with several predictors, there are several methods available. In hierarchical regression, the most important predictor is entered first followed by other predictors. But in what sequence? Based on past literature or if we have a strong hypothesis in mind. After the known predictor is included, we can add new predictors either in one go or stepwise manner. Then, the R2 is checked each time to see whether adding predictors reliably improves the model. If you want to include all predictors at one go, this is called forced-entry method. The most exploratory method that receives less scientific respect is actually stepwise regression where you start with the whole bunch of predictors, and predictors that contribute the least will be removed step by step. Here, what eventually remains is a set of "good" predictors.

Lastly, there is a predictor that helps to increase the variance explained in the model when placed together with other predictors, although in itself it isn't correlated with the dependent variable. This is called suppressor variable, since it suppresses the residuals.

(2) Adjustment and interaction
Adjustment, is the general idea of putting predictors into a linear model to investigate the role of a third variable on the relationship between another two. Example, suppose you recorded blood pressure (continuous var) based on two types of participants, those who do and don't seek treatment (binary, 0/1). A third, continuous variable x, is bodyweight that is correlated with blood pressure. We may model our data with  Y = β0 + β1X + β2T + e. Refer to Fig-3. Here, the treatment status T and their fitted line are depicted as a different color. The horizontal dotted line represents the mean of each group when we disregard X. Try to regress X out of Y in R, and see whether the distance of two dotted lines is similar to the difference in intercept!
Fig-3: You want to predict blood pressure (Y) from treatment effect (T). A third variable X (body weight) carries a linear relationship with blood pressure as well. Now, each treatment status has the same relationship (slope) with Y. The black line is what you would get if you just fit X and ignored the group.


The predictor X is unrelated to treatment status (color). The groups have different intercept that depends on treatment status. So, β1 represents the difference in intercept, and β2 is the slope common to both treatments. Notice that the estimated relationship between the group variable and the outcome doesn’t change much, regardless of whether X is accounted for or not. This is seen by comparing the distance between the horizontal dotted lines and the distance between the intercepts of the fitted lines: blue and red. That the relationship doesn’t change much is ultimately a statement about balance. The nuisance variable (X) is well balanced between levels of the group variable. So, whether you account for X or not, you get about the same answer. One way to try to achieve such balance with high probability is to randomize the group variable. This is especially useful, of course, when one doesn’t get to observe the nuisance covariate.

(3) Issues: outliers, confounds, and multicollinearity
Outliers mislead people in finding parameter estimates and eventually drawing conclusions. That's why, scatter-plot is our good friend in correlation/regression analysis. There are a few diagnostic tools for quantifying an outlier. We can find the distance from an outlier to the fitted line, e.g. Cook's distance, dfbetas() in R, etc. A confound is defined as a variable that is correlated to both your outcome and predictor variables in a way able to explain the relationship. Not accounting for this confound may introduce bias to the estimates.

Multicollinearity arises because there is a strong correlation between two or more predictors in the model. Look at the case of Fig-4 where IV1 and IV2 has a very strong correlation. Each predictor alone is able to reliably predict DV with an almost identical slope. When together, IV1 and IV2 have different slope estimates, but very low p-values > 0.10 because the variance of the unexplained component (residual) is too big. In multiple regression, the problem of multicollinearity makes it difficult to see the isolated effects of each predictor on the dependent variable. In practice, this can be diagnosed by first computing the correlation matrix and see which predictors are correlated one another. It can also be diagnosed systematically using variance inflation factor (VIF), where VIF >> 10 usually means high multicollinearity. Refer to Fig-5 to visualize multicollinearity using a Venn's diagram.
Fig-4: What happens when > 2 predictors have a strong correlation. This concept becomes useful when dealing with fMRI data!

Visualization using Venn's Diagram
I like this concept; learning multiple regression becomes easier using Venn's diagram. Each variable in the regression model is represented by a different circle: a blue circle normally for the dependent variable, while red circles for independent variables/predictors. The radius or size represents the variance. The manner in which the circles overlap illustrates how they behave in the same manner or covary (covariance). E.g., area A in the left panel = cov(X1, Y). It tells us the variance of Y explained by X1. Note how r-squared is computed visually and how it becomes tricky when two predictors are present.

Computing parameter estimates is our next topic. The shaded area in the left panel shown as green is relevant when we want to explain Y using X1 alone. When predicting Y from both predictors altogether, the green shaded area is the only relevant information for estimating the slope of X1. Why? As mentioned before, β1 is the regression through the origin estimate having X2 removed from both Y and the X1. To visualize this situation, we remove (regress out, partial out) areas A + A" + D" from both Y and X1 (look at the inset on the right!!).
Fig-5: With Venn's diagram, we can try to explain multiple regression. Example, the blue shaded area on the right panel shows variance of the observed data (y) explained only by x2, not x1. Note that this diagram does not provide the sign, meaning there's no way to see negative covariance and parameter estimates, or suppressor effect. The lower left panel shows what happens when two predictors are related. The residual part increases, but the variance of each predictor alone decreases (B >>, C <<). This causes inflation in standard error of each estimates, and a big drop in the t-statistics.




Suppose we start with a predictor X1 and an outcome variable. The bias in slope estimate only occurs when we fail to take into account another predictor, say X2, that is correlated with both the dependent variable and one of the predictors (also called the confounding variable). Look at the lower right panel of Fig-5. Technically, bias will not occur if there is no correlation between X1 and X2. However, it will give a lower t-statistics. Why? As with multicollinearity, the value of the denominator that's proportional to the variance of the residual σR2 would go up.


Main references
(1) Andy Field, et al, "Discovering Statistics using R" (2012).
(2) Brian Caffo, notes from online Coursera, Regression.
(3) This link for visualizing multiple regression using Venn diagram.

Thursday, June 16, 2016

Version control: Backing up your data

In my first job as an engineer, we already practiced data backup. The main goal obviously to keep a copy of your dearest work so that you can get it back following an unfortunate circumstance. We used a simple USB stick or portable harddisk. When I worked in the company just before starting my PhD, we used an online repository with version control. All these were done in Windows machines. Now, version control is particularly important when you work in a team with a certain tight schedule to achieve. With version control, multiple users or programmers can work on their piece at the same time and commit ("apply changes") the updates accordingly.

About four years later, I'm doing the same thing. There are many things we need to backup. A large part of my graduate work involves coding: Matlab script, bash script for neuroimaging analysis, and InMotion robot script in Tcl/Tk. A set of codes we are working on is called a working copy. Currently, I'm backing up my so-called working copy from a Linux computer using a Subversion or SVN, a type of version control platform. Another platform called Git has a similar purpose but different set of commands. For example: git clone is similar to svn checkout;  git pull is similar to svn update;  git commit -a .... then git push (push to a server) are the same as svn commit. In git, one has to ensure that a new file/change is added before doing any commit. A "commit" represents a collection of changes that we are ready to apply to the repository.

This website is kinda helpful for beginners;  http://www.tutorialspoint.com/svn/

I found the following schematics from this website very useful, I'm a visual person maybe.
I'm also introduced to three different online repositories: Assembla (for SVN), GitHub, and BitBucket. These three are free for personal use. I put my important items such as research notebook, lecture notes, forms, scientific papers in Dropbox (with 7 GB size). This online space doesn't require any version control. It's quite sad to see Copy closing down, it has a bigger drive space for my data.

How does git work? What is a git branching and merging? The following website sums up everything: https://git-scm.com/book/en/v1/Git-Branching-What-a-Branch-Is
If one is working in a team, he/she has to work on one particular function while others on their functions too. How do they code in parallel? They have branches that are derived from a Master copy. Once they are confident the code works, they usually run a test suite on their code first before merging back into the Master copy.

Another useful info is the cheat-sheet here.

Thursday, April 14, 2016

More on Basal Ganglia & Cerebellum

The themes concerning basal ganglia and cerebellum have been summarized in my previous blog post. Both structures have been traditionally known to play dominant roles in voluntary movements. However, as time goes by, such roles have developed into a wider perspective related to non-motor functions such as general learning, executive functions, and emotion.

Functional topography of Basal Ganglia

There is a convergence of cortical information in the striatum from the cortex. This means that axons of cortical neurons terminate on a far smaller number of striatal neurons, Similarly, the number of neurons in the pallidum and substantia nigra is smaller than that in the striatum, allowing further convergence along the direct and indirect pathways. There also seems to be a somatotopic organization with the cortico-striatal-thalamic pathways. Evidence mainly comes from animal studies using tracers systematically injected into the monkey brain. This includes projections from GPi to the M1, SMA, and premotor cortex (Hoover & Strick, 1993; Akkal, Dum, & Strick, 2007), to Area 7b in the parietal lobe (Clower, Dum, and Strick, 2007).

Important evidence showing how BG is involved in cognitive processes has been shown anatomically through the work of Middleton and Strick (2002). The authors injected a few different prefrontal areas such as Area 9m /9l, Area 46v/46d, and Area 12l ['m' and 'l', 'v' and 'd', 'l' stands for medial/lateral, ventral/dorsal, and lateral respectively]. They found that labels associated with these areas comprised almost 1/3 of the GPi, indicating the importance of BG for executive tasks.

Non-motor Aspects of Basal Ganglia
Although in the early days, scientists thought that BG is important only for sensorimotor functions, there has been considerable acceptance of the four parallel divisions or systems in BG. They include circuits responsible for the skeletal/sensorimotor system, oculomotor system, executive functions, and limbic functions (Alexander et al, 1990). The oculomotor system engages the frontal eye field (FEF) and the supplementary eye field (SEF) anterior to the dorsal premotor cortex (PMd). From these two areas, projections go into the caudate area of the striatum and in turn controlling the SNpr/GPi via the direct pathway before going out to the superior colliculus (brain stem). The oculomotor system is very important for saccadic eye movements and memory-guided saccades.  Refer to the figure below.
Fig-1:  Four different basal ganglia functional divisions (adapted from Kandel, 5e)

The prefrontal cortex subserves higher-order behavior such as cognitive control, reasoning, problem-solving, general attention, working memory, which are collectively known as the executive functions. Two areas, the dorsolateral prefrontal cortex (DLPFC) and lateral orbitofrontal cortex (LOFC), project to the dorsal caudate of the striatum. DLPFC is important for organizing behavioral responses to complex problems and using verbal skills in problem-solving. LOFC, on the other hand, controls empathy and socially-appropriate behavior. The limbic circuit [Nauta, 1986 for review] begins with projections from the anterior cingulate cortex (ACC) and ventromedial prefrontal cortex (MPFC) to the ventral striatum, which also receives input from the memory structures: hippocampus, amygdala, and entorhinal cortices. The ACC/MPFC complex is important for motivating behavior and reinforcement learning.

Patients with cerebellar damage
When a person performs skilled and goal-directed movements, the cerebellum maintains accurate and timely limb movements. One influential idea says that the cerebellum is crucial for error-based motor adaptation. Prior evidence that point to this idea comes from studies involving visuomotor and force-field adaptation (e.g. the works from Amy Bastian or Reza Shadmehr). In 1970s, Marr and Albus independently suggested that the cerebellum may be involved in motor learning. This is associated with the concept of plasticity between the Purkinje cells and parallel fibers inside the cerebellar cortex. Ito later proposed the role of complex spikes from the climbing fibers as the learning signal. Experiments using motor adaptation have shown that people with cerebellar damage impacting climbing fibers aren't able to adapt to both force-field and visuomotor perturbation.

The output gates of the cerebellum are located in the deep cerebellar nuclei in the white matter. These nuclei send projections back to the cortex via the thalamus, and they are strongly excitatory necessary to maintain muscle tone. This, however, isn't related to strength but more on the timing. Lesions of the interposed nucleus reduce the accuracy of reaching movements because of errors in timing, in direction and extent, in straightness due to poor joint coordination. The hand oscillates irregularly around the target. Neurological examinations show that damage to cerebrocerebellar path delays movement timing (Holmes G., 1939). A complex movement can be decomposed into a sequence of movement components. In healthy people, this decomposition is not obvious as the movements are performed smoothly and timely.

Experiments with primates help to elucidate the role in the motor disruption. When a monkey is trying to keep its arm in a fixed position, the application of a force to extend the elbow triggers a stretch reflex in the biceps. This will pull the arm rapidly and precisely to its initial location. The contraction of the extensor triceps plays a dominant role in keeping this precision, preventing the elbow from overshooting after the biceps contract. This appears as an anticipatory mechanism, i.e. feed-forward control. When the two deep cerebellar nuclei, dentate and interposed nuclei are deactivated, the arm oscillates instead of firmly going back to the original position. There seems to be an excessive, yet inaccurate, feedback correction (Vilis, Hore, Flament, 1984 & 1986).

Different Cerebellar Lobules
Like the cerebral cortex, the cerebellar cortex can be divided into four functional divisions: the vermis, intermediate zone or paravermis, lateral hemispheres, and flocculonodular lobe. Animal studies, in particular, in rats and cats have been used as models for anatomical studies. Perhaps, the most popular and classic human cerebellar anatomy comes from the works by Larsell & Jansen (1970). Refer to the Fig-2 below. The lateral hemisphere is especially important because of its massive connections with the cerebral cortex. Both vermis and the lateral hemisphere ('H') are further divided into nine subdivisions or lobules, each assigned a Roman number I to IX.

The earlier attempts to localize cerebellar functional organization were done as a result of lesion studies in both higher-order mammals and humans. Like the sensorimotor cortex, the existence localization and cerebellar somatotopic map do not represent the bodily extent, but rather, the functional demand. A series of the different somatotopic maps were released by Bolk (1904), Adrian (1943), and Snider & Stowell (1944). The progress continued with the help of tracers and, more recently, neuroimaging. Using fMRI, Grodd et al. (2001) did an elegant study producing sensorimotor topography of the cerebellum. The study was based on 46 human subjects performing a series of motor tasks such as opening/closing of right/left fist, extending right/left arm and leg, and moving the lips. I also recommend another article by Manni & Petrosini (2004) in Nature Rev. Neurosci. for a more elaborated summary with historical contexts. The following diagrams are taken from their paper.
Fig-2: Gross anatomy of the cerebellum (right, after Larsell & Jansen) and functional map (left, after Grodd et al.). Both vermis and the lateral hemisphere ('H') are subdivided into lobules. Of utmost interest is the left somatotopic map. The arms are represented by Lobule V-VI. In particular, fine motor control of the hand and fingers are by Lobule VI and VIII of the lateral hemisphere. Note that the connections to bodily parts are ipsilateral.



The involvement of non-motor functions are shown by the neuroimaging work by Stoodley et al (2012), They found right-handed finger-tapping activated right cerebellar lobules IV–V and VIII. Verb generation engaged right cerebellar lobules VI and Crus I and a second cluster in lobules VIIB–VIIIA. Furthermore, mental rotation activated medial left cerebellar lobule VII (Crus II). Lastly, 2-back working memory task activated bilateral regions of lobules VI–VII.

Important Anatomical Connections
Studies elucidating specific connections between the cortex and cerebellum have been done using anatomical tracers in animals. For example, retrograde transneuronal transport with herpes simplex virus, HSV1, was injected into the cerebral cortex of Cebus monkeys to label neurons in the dentate nucleus (Dum & Strick, 2003, Akkal, Dum, & Strick, 2007). With sufficient survival time, it was enough to reveal a set of first-order neurons in the thalamus, and second-order neurons in the dentate nucleus. Dentate nucleus was targeted by researchers at the time as the shape is the largest and easily recognized. The study showed different dentate output channels to the motor areas (M1 and SMA) and non-motor regions in pre-SMA, Area 7b, Area 46, and Area 9L.

Another important study adopting both retrograde and anterograde tracers is by Kelly & Strick (2003) that reveals the existence of cortico-cerebellar loop in both motor and non-motor domains. In the motor loop, neurons in M1 project to cerebellar lobules V, VI, and HVIIB and HVIII, and project back to the same regions of cortex via dorsal parts of the dentate nucleus and the motor thalamus. In the prefrontal loop, however, Area 46 projects to lobule HVIIA/B (mainly to Crus II) via the pontine nuclei, and returns back to the same areas of the prefrontal cortex via ventral parts of the cerebellar dentate nucleus and prefrontal thalamus. This segregated domain also confirms the neuroimaging works mentioned above, separating motor and non-motor areas in the cerebellum. Lobules that are connected to prefrontal regions are considered non-motoric. Refer to Fig-3.
Fig-3:  Anatomical projections from the cerebral cortex and cerebellum forms loop (from Kelly & Strick, 2003)

Basal ganglia and cerebellum receive inputs from cortical motor areas and send projections back to the same areas. These multisynaptic links suggest the extent of influence of both structures to movement production. Most anatomical connectivity has been revealed by tracing studies done by Strick and his group. For example: injections of retrograde tracing to a specific area in M1 or SMA (e.g. arm) with right survival time have revealed how neurons in the dentate nucleus and internal segment of the globus pallidus are labeled. In fact, there seems to be clear functional segregation (e.g. arm, face, digits) in both dentate nucleus and GPi. This suggests that neurons in both the cerebellum and basal ganglia do project to the motor cortex. However, there is no direct connectivity between basal ganglia and cerebellum.