Showing posts with label coding. Show all posts
Showing posts with label coding. Show all posts

Wednesday, March 30, 2022

Linear models: Getting the syntax right!

When conducting statistical analysis and mixed-effects modelling, I often either got confused or forgot how to write the syntax properly. I thought this entry will be useful. 

Simple linear models
In simple linear models in R using the lm library, we can write the formula as DV ~ IV (DV is a dependent variable). The command employs a few arithmetic symbols. A '+ sign' indicates more than one main effect or predictor (independent variable, IV) with no interaction. A '* sign' provides a short form of main effects with interaction. Number '1' with a '+ sign' refers to an intercept. There is no need to explicitly write this, as it is always implied and estimated by default. Sometimes, you want to omit this by simply writing '0'.


When we want to include fixed/random factors, we then use, e.g. the lmer library, and the syntax is slightly different.

DV ~ 1 + IV1 * IV2        DV ~ IV | grouping                 
  1. The dependent variable (DV) is the response variable to be predicted. The independent variable (IV) is the one whose effects we will assess. By default, IV is treated as a fixed factor.
  2. A vertical bar denotes a grouping factor or random factor. It separates expressions for design matrices from the grouping factor. For example: to fit a predictor for each random factor, you can write 1 + A|S, which means the same thing as (A|S).
  3. A '/ sign' indicates nesting. So (school/class) means classes are nested within school.
  4. Fixed factors can be included without any grouping. You can have additional random factors without any fixed factor (an intercept-only model). In each random/fixed factor, ask yourself whether you allow the intercept to change, the slope, or both!


    Hierarchical Linear Modeling
    Begin with the first and simplest model which is the intercept-only model. It has no predictor or independent variable (so, no slope!). Suppose we have one random factor, usually the participant factor. The model is also called the unconditional model.

    Y ~ (1|S)            ; intercept-only model 

    The next model is the random intercept model. We assign a new variable or predictor X as Level 1, which acts as a fixed factor. Here, the intercepts for different subjects will vary but not the slopes. Note that the intercept for X can be omitted as the function understands it to be 1+X. 

    Y ~ X + (1|S)        ; random intercept model

    Now, we can add complexity by allowing different subjects to have different intercepts and slopes. This is the random intercept + slope model. There are two possibilities: the variations of intercept and slope can be independent or correlated!

    Y ~ X + (1|S) + (0 + X|S)    ; independent
    Y ~ X + (X|S)                ; correlated

     Source: here

    Thursday, July 26, 2018

    Some coding stuffs

    In this lab, I learn new things on how to code in python and R. These two languages are pretty popular among data scientists. In particular, I don't know why there is a growing popularity of big data and AI recently and Montreal becomes one of the hubs after Facebook pumped in money to our institutes here.

    Now after coding your analyses, you want to present them in a nicely fashion. One way is to use iPython notebook which is now popularly known as Jupyter (Julia-Python-R). While the installation is straightforward, the additional installation of libraries can be tricky. Yeah this is the downside of using Linux/Mac system, you gotta know what you are doing. For example, I remember I couldn't get ezANOVA to work in Jupyter. It is very handy to also learn how to customize this Jupyter notebook. The following link is useful to integrate multiple languages.

    https://vatlab.github.io/sos-docs/

    There are some drawbacks of using Linux operating system from a Windows/Mac user perspective (at least to me!) One of the most dangerous commands to use is "rm" which basically remove files/folders without putting them to the recycle bin. Others, such as "cp" and "mv" are equally dangerous if not used carefully.

    Thursday, June 16, 2016

    Version control: Backing up your data

    In my first job as an engineer, we already practiced data backup. The main goal obviously to keep a copy of your dearest work so that you can get it back following an unfortunate circumstance. We used a simple USB stick or portable harddisk. When I worked in the company just before starting my PhD, we used an online repository with version control. All these were done in Windows machines. Now, version control is particularly important when you work in a team with a certain tight schedule to achieve. With version control, multiple users or programmers can work on their piece at the same time and commit ("apply changes") the updates accordingly.

    About four years later, I'm doing the same thing. There are many things we need to backup. A large part of my graduate work involves coding: Matlab script, bash script for neuroimaging analysis, and InMotion robot script in Tcl/Tk. A set of codes we are working on is called a working copy. Currently, I'm backing up my so-called working copy from a Linux computer using a Subversion or SVN, a type of version control platform. Another platform called Git has a similar purpose but different set of commands. For example: git clone is similar to svn checkout;  git pull is similar to svn update;  git commit -a .... then git push (push to a server) are the same as svn commit. In git, one has to ensure that a new file/change is added before doing any commit. A "commit" represents a collection of changes that we are ready to apply to the repository.

    This website is kinda helpful for beginners;  http://www.tutorialspoint.com/svn/

    I found the following schematics from this website very useful, I'm a visual person maybe.
    I'm also introduced to three different online repositories: Assembla (for SVN), GitHub, and BitBucket. These three are free for personal use. I put my important items such as research notebook, lecture notes, forms, scientific papers in Dropbox (with 7 GB size). This online space doesn't require any version control. It's quite sad to see Copy closing down, it has a bigger drive space for my data.

    How does git work? What is a git branching and merging? The following website sums up everything: https://git-scm.com/book/en/v1/Git-Branching-What-a-Branch-Is
    If one is working in a team, he/she has to work on one particular function while others on their functions too. How do they code in parallel? They have branches that are derived from a Master copy. Once they are confident the code works, they usually run a test suite on their code first before merging back into the Master copy.

    Another useful info is the cheat-sheet here.