Tuesday, November 8, 2016

Experimental Design - Statistical Tests

Basic inferential statistics.   I decided to separate this from the earlier entry. We shouldn't forget that what we observe is based on a sample, not a population. Inferential statistics has its goal in making a conclusion about a population from just this sample. In doing so, there bound to be some errors as our measure is just an estimate. Suppose we have a population of interest and that we draw all possible samples of size n. And suppose we compute a statistic e.g., a mean, proportion, standard deviation for each sample. The probability distribution of this statistic for all possible samples is called a sampling distribution; the standard deviation of which is called standard error which roughly tells us the spread around the mean of the sampling distribution. Mathematically, statisticians have to make some assumptions in order to model the problem accurately. This will be discussed later on.

How can we infer about a population parameter by using our sample? We use the concept of the sampling distribution (of means) to see where the true population mean lies w.r.t our sample mean. Having obtained the sample mean and standard deviation, we can find a confidence interval within which the true population parameter lies. This interval can be found with a certain confidence level, say 95% or 99%, with the help of a transformation. This data transformation converts our sample into a known sampling distribution, e.g. using a Student t-distribution, by first converting our scores into a t-value. The conclusion of our experiment is obtained by making use of a known sampling distribution and putting that distribution in a certain hypothesis test. The type of statistical test used (e.g. one-sample t-test, ANOVA, or chi-square test) depends on the nature of our variables.

In constructing a hypothesis, there are by definition the null and alternative hypothesis, indicated as H0 and H1. In the language of experimental design, the H0 states that the IV or treatment (marijuana) has no effect on DV (exam score). Conversely, the H1 states that the treatment has an effect on DV. To do, we first create a sampling distribution assuming the null, H0 is true. We decide an alpha level α or significant threshold to represent a concept of "very unlikely-ness" or chance that we observe a treatment effect when H0 is true. We then compare the likelihood of getting our sample statistics p to the alpha level that is usually fixed at 0.05 or 0.01. The final outcome of the test is whether we disapprove or retain the null hypothesis. For example,
  • The probability of our sample statistics of 0.01 means there is only 1 out of 100 chance you find no treatment effect in your data when the null H0 is true  ..... We reject H0.
  • The probability of our sample statistics of 0.80 means there is 8 out of 10 chance you find no treatment effect in your data when the null H0 is true  ..... We retain H0.
  • Rejecting the null while in reality it's true is a Type I error with a chance α. Accepting the null while in reality it's false is a Type II error with a chance β. Note that these errors are not measurement errors made by mistakes but rather a measure of chance or probability.
Fig 1: Summary of hypothesis test with its associated Type I/II errors. 

Behavioral science is unique in the sense that the test must first assume the null hypothesis to be correct. Also, the test result simply makes a case without even saying which hypothesis, H0 or H1 is indeed true or factual. Even if H0 is indeed true, we don't know how much the difference is with the untreated sample group. For such reason, the effect size d comes into play. The effect size is defined as the difference in means between the experimental group and control group, divided by the standard deviation of the control group.

Lastly, the ability to correctly reject the H0 or find something significant in your data while H0 is indeed wrong is called power, and it is simply 1 – β. A powerful design is able to detect a small effect size < 0.20 unit. Power increases or β is reduced when the alpha level increases, the mean difference between two conditions under the null hypothesis increases, and sample size n increases. The concept of power is crucial because it determines the replicability of our study. According to Cohen (1966), most of the behavioral studies have low power, i.e. the chance of replicating the same experiment yielding the same result is 50-50.

A technique called power analysis is used to determine the appropriate sample size for the study. Generally, power = 0.80 is an acceptable value. Tools such as "pwr" in R is able to help us in power analysis, e.g. find the minimum sample size to achieve certain p-value and power of 0.80.

Okay, we move on to talk about assumptions in statistics, a crucial component to make our problems mathematically solvable. Most statistical procedures such as t-test, linear regression, and ANOVA are based on assumptions. For example, a linear regression is valid when the DV and IVs are linearly related to begin with (for categorical variables such as "Yes/No", one uses logistic regression instead). Regression has other important assumptions in general, e.g. the residuals are independent, normally distributed, and their variance is homogenous. The type of test that requires certain assumptions to be met is called parametric statistics. Violating these assumptions causes incorrect inference or decision. The assumptions for parametric tests are presented below.

[Reference #1]: Among some relevant textbooks, I chose the materials from the classic JL Myers' "Fundamentals of Experimental Design" (1979).

The normality assumption.   As mentioned, statisticians make assumptions on their methodology to simplify the maths. Both t-test and ANOVA have to satisfy the following main assumptions:
(a) Samples are independent of each other; the score of Tom is independent of Ni;
(b) The normality assumptions (for ANOVA, it means the distribution within groups);
(c) The homogeneity of variance.
The normality assumption means that the sampling distribution comes from a population that is normally distributed or Gaussian. This is hard to know as we don’t have the luxury of knowing all possible samples related to our population of interest. Out of faith, we assume that if the data we take is normally distributed then the sampling distribution will also be normally distributed. Every stats guru also tells us about the central limit theorem that is, the distribution of the sample means approaches a normal distribution as the sample size N increases, regardless of the shape of the population of interest. As a rule of thumb, N > 30 is generally accepted for the assumption to be valid. The average of our sample means is itself the population mean, and the standard deviation is the standard error.

Fig 2: Comparison between normal data and its QQ-plot (left) and "not-really" normal data (right)

This can be tricky if N is as few as 15 or 20. Now, how do we test for normality?
(1) Visual comparison. Normal data has a bell-shaped distribution. It has a specific kurtosis and skewness characteristic. It is recommended to transform the kurtosis and skew values to z-scores. Another useful visual test is through using Q-Q plot that compares theoretical values against the actual or observed values. An ideal case will be a straight line with 45-deg w.r.t the horizontal. The following figure shows an example how the data that is less normal has the tendency to have Q-Q plot that is not straight, but rather curved away from the perfect straight line. Most of the times, the presence of outliers distort the normality, the skewness or kurtosis of your data. It is quite common to treat outliers by removing the particular data point or by transforming the data.

(2) The Shapiro-Wilk test. In R, you can simply write shapiro.test(variable name); or in Matlab as swtest(x, alpha). This test is super cool as it directly assesses whether a certain distribution is significantly different from the normal distribution with the same mean and standard deviation. If the p is significant (p < 0.01, say), then the data is different from the normal distribution.

Homogeneity of variance.   Suppose you have three groups in your experiment. Homogeneity of variance says that the variability of the data in all groups should be equal, because strictly speaking, they come from the same population. The same notion applies, for example, in the case of repeated measures where you test the same group in two different experimental manipulations. If the homogeneity principle holds true, the variance of the data in either case remains rightfully the same. In most cases, we usually ignore this principle and immediately perform t-test or ANOVA, etc. Refer to the following illustration.

There are two tests that you can use to assess this: Levine’s Test and Hartley’s F-test (good for large sample). Both methods test the null hypothesis that the variance of the two datasets or groups are equal. The Hartley's F-test is also called the equality of variance test, or sometimes variance ratio test since you're actually taking the ratio of. The computed F-statistics will be compared against the standard F-table according to their respective degree of freedom. In general, people have shown that ANOVA is robust against a mild non-normality and heterogenous variance (only for equal sample size among groups!).
Fig 3: Comparison between dataset with homogenous (left) and non-homogeneous variance (right)

Non-parametric statistics.    Non-parametric statistics means the methods don't rely on any assumption discussed above! In fact, the information on the distribution is obtained directly from the dataset itself. This is the most fundamental idea of non-parametric statistics! Bootstrapping, for example, is one popular non-parametric test that relies on "random sampling with replacement". This test can be used to assess stability of your statistics and it is pretty straightforward. Another test called permutation test builds, rather than assumes, the sampling distribution by 'shuffling' the observed data. Unlike bootstrapping, we use only the existing data set and the shuffles are without replacement.

I took this from Matlab website. Suppose you have a set of LSAT scores and  GPA. You can easily plot the data and compute the correlation. It is found that there is a positive relationship between LSAT and GPA, r = +0.78. Now although it may seem large, it is unknown whether this observation is statistically significant or just happens by chance. Using the bootstrp function you can resample the two vectors as many times as you like and consider the variation in the resulting correlation coefficients. If the finding is true, you would expect that bootstrapping process yields more or less the same correlation value. Refer to the following histogram. Bootstrapping gives a distribution of correlation coefficients that is centered near r = +0.78. Using permutation strategy, we can also show how the correlation found is not by chance. We permute the LSAT-GPA pairs for 1000 times, each time is followed by computing the correlation. The correlation values of the permuted data sets should collapse since the relationship has been "broken" following shuffling.
Fig 4: Distribution of correlation coefficients with bootstrapping and the location of the original r = +0.78 (red line).

During the period when I first learnt statistics, we didn't care much about such assumptions and treated them as dogma. Now, I'm learning to accept the presence of non-parametric statistical tests used when the assumptions are not met, that is, to replace the usual t-test etc. The principle behind those tests is assigning ranking. High scores will be represented by large ranks, low scores by small ranks. The analysis is then performed on the ranks rather than the original raw data. In other words, non-parametric tests are useful when dealing with ordinal data. Example of ordinal dataset includes: the satisfaction rating from 1 - 5, the order of athletes based on speed, and the school ranking.

When we are interested in comparing two independent datasets, we need to use the Mann-Whitney test or the Wilcoxon's rank-sum test. This is the counterpart of the independent t-test. In R, you can do this by:  wilcox.test(x,y,____), some parameters inside the brackets have to be defined. In Matlab, use the following instead ranksum(x,y). When we deal with repeated measures or paired sample, then we have to use the Wilcoxon signed-rank test. For example: testing the performance before and after drug consumption. You can use the same command as in (1) in R but by defining "paired = TRUE". In Matlab, however, the command will be different now, signrank(x,y). If we are interested in testing more independent groups, we have to do non-parametric version of the regular ANOVA, e.g. Kruskal-Wallis test and Friedman's ANOVA for repeated-measure designs. Other non-parametric analysis such as the Spearman's correlation coefficient (the counterpart of Pearson's r) and Kendall tau coefficient.



Final note: ANOVA can be understood from the perspective of Linear Modeling and in itself is a very huge topic, covering multiple regression, mixed-effects or hierarchical modeling, and generalized linear models. The mathematical discussions on this theme will be beyond this blog post.....

[Reference #2]: Andy Field et al, "Discovering Statistics using R" (2012).

Wednesday, November 2, 2016

Experimental Design - Basic Concepts

Terminology.    The main goal of any scientific research is to describe behavior, predict the properties, determine the main cause, and explain it as part of a scientific question. To answer a research question, we traditionally follow the following steps: begin with a hypothesis, identify what factors whose effects are to be studied, variable(s) that are manipulated in the study, carefully select a measure of effects, the observed variable; as much minimize the unwanted or irrelevant effects or confounds, design an appropriate experiment and do sampling. Lastly, we perform statistical analyses and draw a conclusion.

There are 4 types of variables according to their nature: nominal (non-numeric, category), ordinal (rank, or ordering), interval (no true zero, e.g. temperature), and ratio scale (zero means absence, e.g. weight). These types exist also according to the way we measure them, or their level of measurement. The first two are qualitative, i.e. kinds, types, categories. The last two are quantitative variables or quantity, amount.

From the perspective of an experiment, there are 3 main types of variables. We have come across the dependent and independent variables a few times, including when talking about linear regression. Independent variables are under the control of the experimenter, independent of the participants. Dependent variable(s) is usually measured. A design that has only one DV is called a univariate design, more than 1 DV simultaneously analyzed is called a multivariate design. The following table nicely summarizes the differences. The third type, the confounding variable, is undesirable because it seems to explain our research question, but the actual fact is irrelevant.

Independent variable (IV): a variable that you are manipulating, a.k.a
- Explanatory variable
- Regressor
- Controlled or manipulated variable
- Predictor variable
- Input variable
Dependent variable (DV): a variable that is affected in the experiment, a.k.a
- Explained variable
- Regressand
- Measured variable
- Response variable
- Outcome or output variable

We then design an appropriate experiment to test our hypothesis, determine how many groups of subjects, and wisely sample our subjects from a population of interest. There are a few ways to do sampling, e.g. random sampling, stratified sampling. The independent variables are manipulated to provide a different experimental procedure to samples or subjects. These different procedures are called treatments, and hence independent variables are also called treatment variables. In designing our experiment, there are two main ways we can assign the participant into. Refer to the table below.

Within-subject design, repeated-measure design = the same participant experiences multiple treatment or protocols in the study.
- Lower error variance, 
- Higher statistical power,
- Risk of carryover effect or order effect,
Counterbalancing and/or Latin square method is able to overcome this risk.
Between-subject design, independent-group design = participant is randomly assigned to only one treatment or one level of the independent variable of the study.
- Higher error variance (inter-subject var.)
- May require a larger sample, time & cost.

A simple example is here. Suppose you want to test the effect of marijuana on the final exam score. The independent variable here is the consumption amount. The dependent variable is the final exam score. Here, it is good to have 2 independent or treatment groups of subjects, i.e. no single subject goes through the experiment more than once. The obvious statistical comparison is between the group of students who take marijuana and another group who doesn't. In an ideal experiment, the only difference between the two groups is the treatment (marijuana consumption). In real life, this is hard to achieve. Irrelevant differences must then be controlled. A controlled experiment is taken exactly to minimize those differences.

What about confound variables? An obvious confound may be the student IQ. If this variable is able to explain the exam score well, this result is not directly useful to our study. In other words, our design should only reflect the influence of our independent variable on our dependent variable. Sometimes, a control group is used to minimize the confounding effects in our analyses. What we can do now is to compare control vs marijuana and control vs no marijuana. The presence of confounding effects and other subjective influences can also be minimized by performing randomization as a form of control during the experiment. For example, we can assign each participant to a group randomly regardless of their IQ level as the focus here is on marijuana consumption.

Another aspect to understand is the term fixed and random variable. A fixed variable whose different levels are have been arbitrarily, but rationally, chosen. Inferences taken about a certain fixed variable is limited to the selected members. A random variable has levels that are subsets of a larger population, bearing in mind that each level has an equal opportunity to be selected. It helps us to link our data to a population in general. Subjects are typically a random variable. There is always a risk that our experiment introduces a bias in our data. This may occur simply because of not selecting the subjects wisely, error in measurement, and finally, the confounding variables not accounted for. It is important that the dependent variable is sensitive and reliable to describe our hypothesis.

Cause and effect are important to explain certain behavior. Saying the average exam scores between these two groups are significantly different from each other does not directly say that drug consumption is the main cause of the effect. In addition, correlation and regression analyses do not imply causation. Traditionally, one has to prove that marijuana consumption is the necessary and sufficient factor for the effect to occur. But in reality, this is difficult to prove simply because there are many confounding factors that may explain an outcome or result. One way is to construct a chain of reaction and apply mediation analyses.

One critical factor in behavioural research is the validity of the variables being investigated. Validity means truth, how accurate a variable is to represent information. There are a few kinds of validity:

(1) Construct validity: adequacy of the operational definition of a variable. In social sciences, a construct is a measure we are interested in. Unlike in physics, constructs are often abstract or subjective. For example: if a social anxiety measure for teens is the one we are researching on, then that is our construct. Construct validity says that the measure has to describe or quantify social anxiety, not other things such as mood or temperament. 

(2) External validity: describes how a conclusion of a study can be generalized to other populations or similar studies. In our example using marijuana consumption on campus, how true does this happen across different schools in general? Hypothesis testing is performed to achieve this purpose. 

(3) Concurrent validity: when validity is tested against an established measure of the same or a related underlying construct assessed within a similar time frame, then we are looking for this type of validity. 

(4) Internal validity: refers to how well the effect in our experiment is attributed to our independent variable. A good experimental design must ensure this is the case. 

(5) Ecological validity: not only does it apply in the laboratory context, but it also has to accurately reflect what we want in the real-world setting. This is important as some measures, e.g. in sports or healthcare, require subjects to be tested outside of the laboratory. 

Finally, the results of our trials or experiments have to be reliable when we perform measurements or observations. A reliable observation is one that is consistent and repeatable. Consistency can be measured in terms of inter-rater reliability (using a few judges, measure once) or test-retest reliability (using a single judge, measure a few times).

Note: For more discussion, see Cozby (2011).

Factorial Design.     So far we only see one simple treatment, e.g. drug consumption, in our study. Let's introduce a new term: an independent variable is also called a factor, often a categorical variable. For example, when we talk about drug abuse, the drug type is a factor. Each substance: alcohol, marijuana, nicotine, and amphetamine is called the level of that factor.  In our experiment, if in addition to drug consumption, we introduce another independent variable in our example, e.g. criminal history and study hours we have two factors in the design. Proper design allows us to explain behavior in a multi-dimensional view.

Often, an experimental design can be a little complicated. Suppose we have 2 factors. If every level of factor-A occurs with every level of factor-B, it is called crossing. A study that has many levels or multiple crossing factors is termed factorial design. We can construct a factor A x B table. In general, a factorial experiment is typically the case. However, often factor-A may contain different level of factor-B. This is called nesting, factor-B is nested in factor-A.

Which statistical test should we use? Comparing two means in your data can be tested using t-test, be it paired or independent samples t-test. In a factorial design, when dealing with one factor with two levels only, the t-test is sufficient. When more levels are involved, however, the Analysis of Variance or ANOVA is used as an extension of t-test. ANOVA is based on F-test, which is a ratio of two variances (or mean-squares, MS). Refer to the equation below. When the null hypothesis is true that there is no difference in means, the F-value should be 1.0. A large F-value indicates some differences may exist among the means due to a systematic treatment or manipulation effect. Note that it doesn't tell which means are different! To find which means, we should do post-hoc tests.
There are two main types of ANOVA: (1) a 1-way analysis compares levels of a single factor based on a single continuous response variable; and (2) a 2-way analysis compares levels of two factors on a single continuous response variable; and so on.... Now, it seems that study hours, a continuous independent variable, is correlated with the exam score. To treat this in ANOVA, we use a variant analysis called Analysis of Covariance or ANCOVA.  ANCOVA compares the dependent variable by both a factor (categorical) and a continuous independent variable. ANOVA characterises between-group variations due to a treatment effect. In contrast, ANCOVA divides between-group variations to treatment and a continuous covariate.

Knowing what statistical test to use is paramount in research!! Several other hypothetical situations and the type of possible statistical tests used will be presented here:
To see the difference in exam score between Harvard and UPenn students in their "Intro to Statistics" class, I recruit a total of 30 first-year undergraduate students from each university. Then, I perform unpaired t-test (nondirectional or 2-tailed, with α = 0.05) as the samples come from two independent groups or origins.
To see whether there is a difference in exam score among Harvard, MIT, and Stanford students in their "Intro to Statistics" class, I recruit a total of 30 first-year undergraduate students from each university. Then, I perform 1-way ANOVA to test the main effect of university (2-tailed, with α = 0.05).
To see whether there is a difference in exam score among Harvard, MIT, and Stanford students in their "Intro to Statistics" class, I recruit a total of 30 first-year undergraduate students from each university. I introduce a new factor called gender (male and female). Then, I perform 2-way ANOVA to test the main effect of university and of gender. If the score of different university is also influenced by gender, I'm looking at the interaction between the two factors. Interpretation of the main effects can be different in the presence of an interaction!
To see whether there is a difference in exam score among Harvard, MIT, and Stanford students in their "Intro to Statistics" class, I recruit a total of 30 first-year undergraduate students from each university. It seems that the score is correlated with IQ level. If I treat the IQ level as a continuous independent variable or covariate, then, I perform 1-way ANCOVA to test the main effect of university (2-tailed, with α = 0.05) having controlled the IQ level as the covariate.

[Reference #1]: Paul Cozby, "Methods in Behavioral Research, 10 ed" (2011). 
[Reference #2]: JL Myers' "Fundamentals of Experimental Design" (1979).