Friday, August 29, 2014

Sensorimotor Control & Learning (Part I)

"Motor learning", the main theme of my current lab, lies at the intersection between motor behavior and neuropsychology of learning. It helps us to understand how a person acquires and learns new movements, e.g. to dance, play golf, or adopt a new language. The theme is studied depending on the organ where the voluntary movements are produced: the arm, legs, jaw, and eyes.

The current summary is based on modern studies from the 1990s by a group of engineers and computational scientists, outside the domain of neuropsychology and kinesiology. Emphasis will be made on the upper limb (arm reaching) as its core manipulation.

Components of Sensorimotor Control
First, gathering sensory information associated with the task. When one wants to move or perform a task with one's limb, visual information is gathered through saccades. This gaze behavior is also task-specific. Our brain appears to be able to filter irrelevant sensory inputs. For example: the study in inattentional blindness when one fails to notice prominent visual stimuli unrelated to the task one is attending. It is worth noting that sensory streams are temporally delayed and noisy.

Second, motor tasks involve a sequence of decision-making processes in the presence of delay and noise. Why does the noise come into play? Because both sensory and motor system are inherently noisy, arising naturally at the molecular, synaptic, and system levels. The final product is that trial-by-trial movement production to the same target in space is bound to exhibit some variability. These underlying risks can be viewed with the context of reward, especially during a period of motor learning. Even when a person faces the same motor learning task, one can be a risk-averse (exploitation) or risk-seeking (exploration).

Third, our nervous system can be modeled as a controller. Traditionally, there are two basic controllers for motor control and learning: the feedback and feedforward controller. Feedback control, as the name implies, refers to the control of voluntary movements using sensory feedback. In contrast, feedforward control doesn't depend on any feedback mechanisms. Given that our sensory inflow has an inherent delay (150 - 200 msec), it becomes unreliable to rely on for an accurate movement control. To meet the demand of the motor tasks, we often rely on feedforward control which is a predictive control. As we produce certain movement, we make also make prediction about the sensory consequences of that movement. This is done using of efference copy of the motor command and the difference between the predicted sensory consequence and the actual sensory feedback will be used in state estimation.

The last controller is related to the biomechanical properties of the body and the tools used. It is within the topic of arm impedance. Impedance control depends on a few factors such as arm stiffness, i.e. how springy the musculatures are. Like the internal model, impedance control is also inspired by concepts in engineering and biomechanics. Example: we can modify the way we grip a tool (hand stiffness, arm stiffness) produced through co-contraction of the opposing muscles. Although co-contraction can be a solution to the motor task, it is inherently unstable as the sensorimotor system is noisy.

Recent Techniques or Methodologies
We can study motor learning in various ways, It can be studied through behavioral studies and quantitative movement analysis. Modern motor learning literature typically includes three well-known behavioral paradigms:
  1. Sequence learning: a type of motor skill learning where it employs serial reaction time tasks (SRTT). Here, a participant is asked to make a sequence of button pressing, key tap, or finger flexion. Motor performance is measured by the reaction or response time. It does not directly deal with the kinematic and dynamic features of motor learning. In more specific ways, learning is measured by the difference in reaction time between the random sequence and learned sequence.
  2. Visuomotor rotation: this can be achieved by e.g., providing a set of prism worn by the participant or by a certain mechanism to distort the association between the visual feedback and the actual arm movement. The performance is measured by the movement deviation. This method introduces a mismatch between two related sensory inputs: visual and proprioception.
  3. Force field paradigm: a participant performs reaching movement with a robotic manipulandum, capable of producing a velocity-dependent force that perturbs the movement trajectory. The presence of the force changes the dynamic of the motor task. The process of reaching motor performance signifies adaptation. The sudden removal of the force causes the trajectory to deflect to the opposite direction know as an after-effect. A more novel idea probes trial-by-trial performance in terms of the magnitude of the lateral force the participant produces. This is achieved by introducing catch trials in the form of force channels. This is the method used by Shadmehr and colleagues.
Lesion studies and non-invasive stimulation (TMS and tDCS) are able to complement the methods. Recently, neuroimaging methods are employed to learn regions of the brain associated with the behavioral tasks involved. Scientists employ engineering and computational modeling to represent the brain as a system or controller. In error-based learning, for example, motor adaptation can be captured by a linear time-invariant model (LTI). With this framework, in each trial, we learn new movements by employing an optimization algorithm (e.g. Kalman Filter).
Types of Motor Learning Processes
The processes of motor learning can be classified by the type of information the motor system learns.
  1. Error-based learning: motor learning through the presence of an error, i.e. the discrepancy between the desired trajectory and the actual movement outcome. The term error also means the mismatch between the predicted sensory consequences and the observed sensory feedback. Three important behavioral paradigms to study this type of learning include visuomotor rotation, prism goggle, and force-field adaptation. Error reduction happens reasonably quick and this type of learning is known as adaptation. Improvements in adaptation reach plateau after 8-10 trials. We will focus more on this type of motor learning processes as it has been widely studied for the past decade.
  2. Reinforcement learning: normally observed in a redundant system. This type of learning can happen even when there is no error involved. Learning is achieved through exploration by finding the best solution in the solution manifold. It is highly dependent on the available rewards, e.g. points, punishment, currency. A study by Izawa & Shadmehr (2011) shows that reinforcement learning, in some circumstances, can substitute for adaptation when there is uncertainty about, or no, sensory prediction error.
  3. Use-dependent learning: not strictly a learning process, but rather, adaptation through repetitive movements. Movements to a certain direction are able to reduce variability in that direction and induce a bias towards this trained direction when reaching to other directions.
  4. Learning by observation: typically it involves watching others doing the movements. This type stems from the findings of mirror neurons. Observational learning may include learning from predicting error by observing the action of others.
  5. Structural learning: learning to extract common features of different task variants. When we know the underlying structure of the task, learning can be faster, e.g. learning to swing a tennis racket bears similarity with learning using a squash racket.
What is Internal Model?
Modern research of human motor behavior has been marked by the incorporation of control engineering theories, in particular, the concepts of the internal model (see: Jordan, 1995; Kawato et al. 1987). A controller Gc(s) is used to control the process Gp(s). A good controller should be able to represent the process to be controlled. It is said that the motor system is composed of the limbs (i.e. the plant) and the controller in the nervous system, the internal model.

The internal model is an approximation of the inverse dynamics of the system being controlled. It is a model that mimics the behavior of the natural process being controlled, which refers to our motor system. This model can be adapted at any time to a novel environment, making it a suitable computational model for motor learning (refer to studies by Shadmehr's group). There are 2 variants of the internal model: forward model and inverse model.
  1. Forward model predicts sensory consequences from the efference copy generated during movement. Forward model is likened to a motor-to-sensory mapping. The efference copy is issued in conjunction with the motor command from the CNS. The model tries to anticipate the next state so that the movement goal is achieved and the error is minimized. 
  2. Inverse model tries to approximate motor commands through an inverse transformation from the incoming sensory streams. This model is reactive rather than predictive. There is a close relationship between model (1) and (2).
The internal model is used to explain motor adaptation. As mentioned, it involves a decrease in sensory prediction error through trial-by-trial adjustments in the forward model. Accordingly, the update of the forward model is translated into an update of motor commands. Mathematically, the internal model is well captured by linear time-invariant (LTI) state-space models, which have sensory errors or perturbations as inputs, sensorimotor mappings as hidden variables, and the learned or adapted motor commands as the output.
Why can't we tickle ourselves? When we tickle our body, the central nervous system predicts the sensory consequence using the efference copy of "tickling". At the same time, there is this somatic sensation generated by the "tickling". As this sensation matches the predicted sensory consequence through the forward model, the comparator circuit in the CNS doesn't detect any mismatch.

Can Motor Learning Generalize?
After going through training of a task in one context or situation, a person is able to perform as well to a similar task but in a different context or situation. This concept is called generalization. When generalization is beneficial, it is usually termed transfer. Conversely, when it is detrimental, it is termed interference. Traditionally, the studies of generalization made use of dynamic or force field paradigm. The principles derived from those studies are associated with the concept of the internal model. On top of that, generalization is related to another concept called "motor memory". If interference occurs, the transfer of learning fails.

Using force field paradigm, it is thought that motor adaptation is able to generalize in the intrinsic coordinate system, i.e. based on internal muscular patterns of activity. A salient example of intrinsic transfer is when writing "9" by right and left hand. On the other hand, using the visuomotor paradigm (Krakauer et al., 2000), motor learning generalizes in the extrinsic coordinate system, i.e. based on the external spatial coordinate frame. Transfer in motor learning has also been studied in relevant to transfer across different movement direction. There is a limited transfer of dynamic (Gandolfo et al., 1996; Sainburg et al., 1999).

Transfer occurs in different configurations of the same arm (Shadmehr & Mussa-Ivaldi, 1994; Ghez et al., 2000; Malfait et al., 2002; Shadmehr & Moussavi, 2000). How about the interlimb transfer? Tranfer occurs from the dominant arm to the non-dominant arm and this happens in the extrinsic coordinate system (Criscimagna-Hemminger et al., 2003). The opposite is not true. Further, the interlimb transfer from the dominant to non-dominant hand occurs only when the force field is introduced abruptly (Malfait & Ostry, 2004). The gradual force field, on the other hand, does not cause the apparent interlimb transfer. It seems that interlimb transfer is regarded as a cognitive process.

The Concept of Motor Memory
The initial part of learning involves more cognitive processes, where one makes use of one's memory buffer to carry out and finish the task. This readily available, temporary buffer or space is known as working memory. A popular example of working memory is when you solve mathematical problems. The later part of learning is the period when motor performance stabilizes and involves consolidation, a term related to long-term storage of motor skills. This is why after a year of not playing the piano (or skiing), we are still able to play it as well. However, what is stored inside the memory (e.g. motor commands, task dynamic, somatic experience, etc.) is still debatable and the nature of consolidation is also conflicting, e.g. Caithness et al (2006).

Memories that can be consciously recalled are named declarative memories, e.g. memory of events, words, or facts. Conversely, memories on skills and knowledge to perform some particular actions are called procedural memories, e.g. the ability to walk or ski. Typically, the domain of motor learning deals with procedural memory, the nature of which is interesting to characterize. Smith et al. introduced a two-rate state-space model of the force-field adaptation. The model says that adaptation consists of fast learning with poor retention and slow learning with more stable retention. Krakauer & Shadmehr discuss whether the formation of such memory progresses over time from a labile state, which is susceptible to interference to a stable state, which is resistant to such interference.

The memory from experiences obtained from a rapidly changing environment leads to faster decaying and unstable motor memory. As opposed, more stable memory is achieved when exposed to a gradually changing environment. What happens when, after learning A, a person immediately learns B? Called retrograde interference, task B is able to disrupt the consolidation of A. This phenomenon does not appear after a longer period of training. There is even evidence suggesting sleeps enhance/improve consolidation. Sometimes, although we forget to do a certain task, a quick relearning is sufficient to meet the expected performance. Such a phenomenon is called saving, that is, faster relearning.

Lastly, the concepts of implicit and explicit processes have a place in the context of motor learning. Explicit processes require declarative knowledge of something. When a person learns to make a golf swing, voluntary explicit processes include, for example, adjustment to the weight of the golf stick or the knowledge on the target location. In contrast, proficiency in skill performance itself or the correct timing (or speed) involves implicit processes. In some conditions, implicit planning may override explicit strategies during a visuomotor adaptation task (Mazzoni & Krakauer, 2006). Scientists are still debating which of the two are dominant during motor learning, at different stages of learning.

The famous case of HM who couldn't recall practicing a mirror writing task but performed well in the task several days later prompts scientists to think that motor learning is purely implicit. Adaptation is seen as an implicit process, but it does not rule out the involvement of explicit processes. Such "cognitive" explicit processes are thought to serve as a form of learning strategy. Taylor and Ivry (2011) found two competing processes: explicit knowledge of target error and implicit knowledge of sensory prediction error (a la the usual adaptation mechanism). Keisler and Shadmehr used an interesting approach to examine declarative memory contribution to force-field adaptation. Subjects were adapted to force-A and then a brief exposure to force-B. After a 3-min interval, they experienced channel trials. At the same time, they had to memorize words in between. This memorization interfered with the memory of the second task B.

References  
[1]  Wolpert D.M., Diedrichsen J & Flanagan J.R. (2011). "Principles of sensorimotor learning". Nature Rev. Neurosci. 12: 739-751.
[2]  Krakauer, J. W. and P. Mazzoni (2011). "Human sensorimotor learning: adaptation, skill, and beyond." Curr  Opin Neurobiol, 21(4): 636-644. 

Saturday, August 23, 2014

Additional MRI Vocabulary

1. On 2D and 3D MR image
The earlier posts on MRI/fMRI are not exhaustive and now I would like to add some more points. Let's begin by refreshing our memory on how an image is reconstructed (a 2D image, I mean). We have RF excitation pulses and we assign a gradient to localize the position in 2D space. Once we get acquire the data on the whole k-space trajectory, we have to inverse them through Inverse Fourier Transform to obtain the actual image.

Fig-1: An illustration showing a T1-weighted image transformed from its k-space equivalent data. Adapted from [1].

Okay, more basic vocabulary in MRI image formation: a slice, field of view, image matrix, slice thickness, and in-plane image resolution. Refer to the following diagram showing an anatomical or structural T1 image. If we take just the 2D image of a particular slice, it is basically a matrix of 32 ´ 32 data points, each point is acquired consecutively for every row. With the in-plane resolution of 3 ´ 3 mm, the slice field of view 32 x 3 = 96 mm. In reality, the slice is captured with a certain thickness as well. The smallest entity of an MR image is called a voxel and its size represents the image resolution. To construct a three-dimensional image, we repeatedly acquire and combine 2D slices. Now, an acquired 3D image is also called a volume or simply an MR image. A voxel that has the same size on all three sides is called isotropic, e.g. 3 ´ 3 ´ 3 mm. One volume is acquired within one TR (repetition time).

Take note that although one volume is acquired within 1 TR, all slices are not captured at once. The way we acquire or record slices follows what we called slice order that is defined by the scanner operator. Essentially, the slice order can be ascending, descending, or interleaved. The ascending slice order means that the first slice will be captured first and the difference in time point between the first and the last would be almost 1 TR. The effect of this difference on subsequent statistical stages can be detrimental! I have mentioned the slice time correction as one of the preprocessing steps.

Fig-2: A slice is made up of rows of voxels. A volume is a collection of different slices acquired in one TR.

2. More On Functional Image
The structural image consists of only one 3D image or volume and it is a T1-based image. On the other hand, a functional MRI image is a T2*-weighted image in 4D, the fourth being the time information. What it means is that what we get in fMRI is a collection of volumes, each acquired one after the other at different time points. Note that a time point corresponds to one repetition time (TR), and thus, one volume per TR. To summarize: functional MRI gives us 3D space × time dataset. In practice, the first few points ~ 2 sec are removed because the hardware (magnet) isn't yet stable.

Refer to the following diagram. Suppose we select just a slice of every volume acquired in every 2 seconds (TR = 2 sec) for 5 minutes. There are 2 experimental conditions involved. And we select a region of interest or ROI as our apriori assumption to look for brain activation.

Fig-3: In functional MRI, volumes are recorded to obtain 4-dimensional data to represent the brain's dynamic. From [3].

3. What is Spatial Normalization?
As mentioned earlier, EPI is the most common fMRI method. Although the acquisition is quick, the resulting image quality with EPI is poor. Coregistration attempts to register these poor images to the T1-weighted structural image of the same subject or patient. Why? Because the structural image carries a much higher spatial resolution. To do this, we simply perform a rigid body transformation with 6 or 7 DOF, consisting of 3 translation and 3 rotation. The goal of this transformation is to maximize the intensity-based correlation between two images, that is, the target and the reference image. This is achieved through an iterative process by minimizing some cost function.

A detailed example of 2D coregistration is as follows. We have two images, together with the information on the intensity values of every voxel. We then plot that in the histogram and assign a bin number. After that, we create another plot based on the bin number of images 1 and 2. Suppose a voxel at location [x, y] has a bin number 23 for the first image and a bin number 8 for the second image. We plot this particular voxel at [23, 8]. Once finished, we can see how this plot resembles a correlation map. If two images are aligned, the plot will be more or less a straight line.

If we then want to know the common activation happening across different subjects, we have to use a template or atlas. This is because individual brain differences can be huge. The most popular standard templates: the Talairach space and the MNI-152. I personally only use the MNI-152 template with a 2 mm resolution. Registering a structural scan of a person to the structural scan of a standard template is called normalization. Normalization is indeed crucial for group or higher-level analysis where different we want to combine different subjects. All these can be done with the currently available software. Klein et al. (NeuroImage, 2009) wrote a neat article on different methods of image registration available, e.g. SyN, ART, SPM, FNIRT, etc.

In general, spatial normalization essentially requires two steps.
  1. The first step is to align the subject image to the standard template using affine transformation with 12 DOF. This is a linear process of translation, rotation, scaling and skewing, each for the three axes. Because the resolution of the image isn't always the same with the reference image, we need to perform interpolation somehow. 
  2. The next step is to perform  non-linear transformation or geometric warping. This is a lengthy process that may require loads of iteration by minimizing some cost function (or maximizing certain 'similarity' measure). Ideally, this non-linear transformation makes use of a 'function' that is diffeomorphic. What it means is that the function maps a point in space U to exactly one and only one point in space V and the relationship is invertible. After normalization is done for all subjects, our position in space follows the standard coordinate.
  3. Note that the MNI template (by Collins et al at the MNI) has different revisions. FSL uses the MNI nonlinear template obtained from 152 healthy adults. You can choose to use the 1mm or 2mm resolution.  

Fig-4: An illustration on how we divide an anatomical slice into 6 bins of different intensity. We need to 'match' bin-1 of the source image to bin-1 of the target image, etc.


4. Orientation Convention in MRI
I find this website is useful. Note that neurological convention (RAS) is the opposite of the traditional radiological convention (LAS). The summary is shown below. In other words, the axial or transverse plane is R-L x A-P plane, coronal is R-L x S-I plane and sagittal is A-P x S-I plane.
Fig-5: Orientation conventions in brain imaging. MRI uses the LAS convention i.e. the left brain is on the right-hand side.

Wait! There is another rule. Suppose you have conducted your experiments and gotten the raw DICOM images from the machine. These images are acquired following a certain convention as well. The following is useful, especially when you work with both AFNI and FSL. This convention dictates us how a row is constructed to a slice, then a volume.
Let's say we have an AFNI image with "orientation" RAI.  This means that:
a) Voxels ordered from: Right to Left to store a row.
b) Rows ordered from: Anterior to Posterior to store a slice.
c) Slices stored from: Inferior to Superior to store a volume.
Importing a DICOM volume using to3d will create an AFNI file with RAI orientation by default. Importing a DICOM volume using dcm2nii will create an NIFTI file with RPI orientation by default.   [from: source]

5. What is AC-PC alignment?
We have just mentioned the orientation convention in the MRI. Having known the jargon, how then, should I place a patient's head in the scanner? By default, the axial or transverse plane has to cut through two important anatomical landmarks of the brain, the anterior commissure, and posterior commissure. They are usually called as AC-PC. Once the head positioning and orientation are determined prior to the acquisition, the slice orientation follows will follow the AC-PC alignment, acting as a reference or anchor. Take a look at Fig-2 above. If the head is tilted in such a way that the AC-PC is not horizontal, it is said that the head is oblique.

Why is this orientation important in the experiment? Imagine that an MRI scanner can only acquire anywhere within a certain "box". Ideally, we would want to cover the full brain as much as possible but the head size may vary. In order to fully cover important brain regions that we want, we may want to tilt the head slightly, e.g. -30 degree w.r.t. AC-PC.

Fig-6: The sagittal view of the head with the location of AC-PC.


6. What is Localizer? What is Auto-align?
Initially, when we put a subject into the bore, the scanner uses its geometric center of the bore (magnet) as its reference. This location may not be the same as the head center, which is calibrated as with the laser before we bring the subject in, Once the subject is inside, we acquire a low-resolution T1 image so that we roughly know the head position with respect to the magnet. This image, called a localizerallows the scanner software to ‘know’ where your head is. Throughout the whole session, the scanner will reference all subsequent images relative to that first reference image. This allows you to prescribe slices on each subsequent image however you like, and the scanner will track where you are in space.

Auto-align is an algorithm (Siemens scanner?) that allows you to prescribe slices on each image throughout the scan, even with different sessions. It basically does some kind of geometric transformations, taking the localizer as the reference. In the Siemens machine, this feature is called AAHScout. AAHScout plays a major role when we are scanning with multiple sessions of the same subject or patient, or when we bring the subject in/out of the bore.

7. Some more jargons to remember
Lastly, it is good to talk about the same vocabulary when discussing the design of fMRI experiments. Some useful jargons may include:
(a)   A session: all MRI/fMRI scans collected for a subject or patient in one day.
(b)   A run or scan: a continuous fMRI scanning, may last up to 7-10 minutes.
(c)   A condition: a task, a type of stimulus presented to the subject or patient.
(d)   An experiment: a set of conditions designed to answer a scientific question.
(e)   A paradigm: a set of manipulation comprising different experimental conditions.
(f)   A pulse sequence: a set of parameters, e.g. TR/TE, flip angle, slice thickness, in-plane resolution. A more recent technology called the multiband sequence becomes widely used.

A session usually consists of one or more experiments. Each experiment consists of several runs. In one run, the subject may be presented with an orderly set of two or more conditions, that is, types of behavioral manipulation or stimulus presentation. More volumes and/or runs are commonly needed when signal-to-noise (SNR) is low or the effect size is small. Thus each session consists of numerous runs, e.g. 5-10 runs, that may last to hours.

8. What is parallel imaging?
More recently, parallel imaging becomes popular to shorten acquisition time and improve SNR, while at the same time obtaining more data points. Broadly speaking, parallel imaging is a generic name denoting multiple acquisitions of images at the same time, thus parallel. In different MRI machines, the algorithm naming can be different but essentially this feature falls into 2 categories:
     (a)  Image space (encoding) following reconstruction, e.g. SENSE.
     (b)  k-Space or frequency domain, e.g. GRAPPA.

I mentioned k-space earlier. This k-space often refers to a temporary image space in matrix format, in which data from digitized MR signals are stored during acquisition. The final image is produced by some mathematical operations when the storage is full, a process sometimes known as image reconstruction. 

Parallel imaging induces more noise and the SNR is reduced by g.√R where g is the noise amplification or g-factor, and R is the acceleration factor. For example, if R = 2 then it reduces imaging time by half, but decreases the original SNR by 1/√2 ≈ 0.71.

Now, which method of parallel imaging should we choose? Well, it really depends on various factors: total time taken, SNR, coverage, etc. Parallel imaging promotes the birth of multiband acquisition in the Human Connectome Project by NIH.


References
[1]  Jezzard, P., Mathhews, P. M., and Smith, S. M. (2001). Functional MRI: An Introduction to Methods. Oxford University Press.
[2]  Lindquist M. (2008). The statistical analysis of fMRI data. Statistical Science 23: 439–464.
[3]  Lecture01 in - culhamlab.ssc.uwo.ca/fmri4newbies/Tutorials.html
[4]  FAQs from http://bic.berkeley.edu/scanning.
[5]  http://mriquestions.com/two-types-of-pi.html

Nice neuroimaging data processing, that covers some basics of the 3 most popular software packages: SPM, FSL, and AFNI.