Wednesday, September 27, 2017

Summary on Working Memory

Working memory is part of the cognitive domain called the executive function. The executive function is the mental processes that enable us to plan, focus attention, remember instructions, and juggle multiple tasks successfully. Scientists more or less agree that the triad of executive function is: working memory, flexibility, and inhibitory control.

Some important properties
The concepts of short-term memory (STM), or short-term store (STS), working memory (WM), etc. are fundamental to the field of cognitive science. The ideas date back to the time when George Miller, Brown, and Peterson & Peterson did various studies on selective attention during the ‘60s. For example, one has to remember a set of 3 alphabets. These sets are forgotten easily after 15 seconds when the person’s attention is diverted to another distracting stimulus. Hence, the concept of STM is closely related to selective attention and cognition in general.

In general, they are different from the so-called memory or the long-term memory in which the content has to be consciously recalled. William James mentioned that long-term memory has to be “brought back”, whereas STM is currently available. It is also related to conscious awareness. As a more special term, WM bears a connotation that it is a mental workbench where mental effort is applied. A popular example of WM is when you try to solve an arithmetic problem (4 + 30/2 = ?) or when you want to understand a foreign language. 

If STM is a type of memory, what is the capacity? It turns out that STM has a limited capacity. Miller originally proposed it has a magical container that can take up to 7 ± 2 items at once. Later research shows that a lesser capacity of around 4 ± 1 items at once is more acceptable (e.g. Cowan, 2010). Hence, STM is the bottleneck in our information processing system. Such capacity can be overcome by chunking or recoding strategy.

Related to the STM capacity is the term “forgetting” or losing the strength in the consciousness. In studies by Brown (1958) and a husband-wife Peterson & Peterson (1959), subjects were presented with a 3-letter (trigram) that did not make any linguistic sense (to prevent chunking), e.g. XMH. Afterward, subjects were given a number, say, 450. Then they had to count backward by three/four for a given interval, after which they have to recall what trigram they just saw. The performance is no better than chance after 3-4 seconds counting. A decay curve of “forgetting” can be fitted to predict the recall accuracy as a function of distractor interval. It seems that Brown-Peterson task points to the idea that WM storage decays as a function of time. However, Waugh and Norman (1965) proposed that it is the interference between the counting and the trigram that causes the forgetting, not time-dependent property per se. Indeed the number of interfering items in between the presentation of an item and the recall period contributes to the severity of forgetting. This temporal decay versus interference theory is still unresolved.

Once placed in the STM, how can the items be retrieved? Serial position curves reveal two properties of STM: the recency and primacy effect. In other words, an item held last or most recent in the queue and the earliest (deliberate attempt to rehearse, or put in the long-term storage) can be retrieved very accurately than the items in between (e.g. Glanzer et al.). A more developed model of STM retrieval comes from S. Sternberg (1969) which includes the way we search a desired item in the STM container, brings an item of interest to consciousness, and do an appropriate response. He was interested in whether people scan items in the short-term memory one at a time (serial processing) or all at once (parallel processing). 

In one study, Sternberg presented a display of numbers from 1 to 6 different digits to subjects. These items are so-called the memory set. Then he showed the subjects a test digit and they had to decide whether the test digit had been included in the display just shown. If parallel processing occurred, the length of the presentation should not influence the reaction time, but if serial processing occurred, the reaction time should increase as a function of presentation time. An interesting finding was that the reaction time was about the same whether the response was a "yes" or a "no." In other words, participants did not stop responding when they found a match but continued searching the entire display in their memory. This process is called an exhaustive search. Thus, Sternberg concluded that people perform an exhaustive search when retrieving information from STM. 

Distinct Components of WM
In the 60's, Atkinson and Shiffrin talked about STM model. The more elaborate model, originally proposed by Baddeley and Hitch (1974), was based on Shallice & Warrington's and other clinical cases. According to Baddeley's model, WM by no means has a central controller and two different components or "slaves" that work independently.
1)  Phonological loop: maintains, manipulates acoustical, verbal information, e.g. rehearsing words. 
2)  Visuospatial sketchpad: maintains, manipulates visual and spatial information, e.g. playing chess.
3)  Central executive: acts as a manager that controls and oversees the use of different WMs.
4)  A more recent development is a component called episodic buffer. 

Phonological loop is related to observed phenomena e.g. phonological similarity, articulatory suppression, and clinical studies of aphasic patients with dyspraxia. The primary goals of the central executive controller is binding information from a number of sources into coherent episodes, coordination of the 3 components, shifting between tasks/retrieval strategies, and selective attention & inhibition. The prefrontal cortex is important for such purposes.

The visuospatial sketchpad has also been expanded quite recently. For e.g., Logie expanded the Baddeley's visuospatial sketchpad by dividing the system into optical (visual) and spatial (mental imagery, movement information). Smyth et al. (1988) shows how subjects performed a recall of a series of movement sequence and this is thought to involve the visuospatial sketchpad. Other studies have shown that there is less interference between visual and spatial tasks than between two visual tasks or two spatial tasks. This suggests that the two entities may be more independent than initially thought.

There are tons of neuropsychological test batteries to investigate WM in humans and patients. Digit span test, where a person is presented with a series of digits one after the other and has to repeat the digits in the same order, is arguably the most common test to access verbal working memory. A variant of this, alphabets, may be used. Corsi Block Test is the spatial equivalent of the test that taps into the spatial working memory. Mental rotation test is another visuospatial test that may also involved mental imagery. Because of the online nature and the link between WM, cognition, and problem solving, WM is sometimes thought to describe some kind of intelligence.

Neural Substrates of WM
The dorsolateral prefrontal cortex (DLPFC), equivalent to BA 9 and BA 46, (or Area 9/46, Petrides & Pandya in primates), is thought to be the most important part for WM both in humans and monkeys. Much of the earlier works in non-human primates were done by Goldman-Rakic in the '80-'90s using a paradigm called delayed-response task. She found that lesion to DLPFC impaired the task performance. A series of neuroimaging studies have corroborated this finding, confirming the involvement of DLPFC in working memory tasks. According to d'Esposito and colleagues, there is a lateralization of functions, in the sense, verbal working memory is more left-lateralized versus visuospatial which is right-lateralized.

Thursday, April 20, 2017

Transcranial Magnetic Stimulation - a brief overview

TMS Principles
Studies involving electrical stimulation over muscles and nerves date back since Galvani and Volta. There were different attempts incorporating invasive stimulation on the cortex during 60's-80's in humans, e.g. Penfield, Graziano, etc. During the 80s, Merton and Morton showed that directly stimulation is able to activate the muscle. This technique, known as transcranial electrical stimulation or TES, then was applied to motor cortex through the intact scalp and able to elicit motor evoked potential (MEP). For TES, current flows from anode (+) to cathode (–) placed on the scalp. However, the technique obviously is painful. In 1985, Barker et al. from Sheffield showed that it was possible to apply Faraday's Law to excite neurons using electromagnetic coils. This is undoubtedly the start of TMS in the field of neurophysiology.

How does TMS work? A brief, high-amplitude pulse of current, lasting for approximately 100 to 200 msec, is discharged into a TMS coil. The current induces a magnetic field perpendicular to the current flow following the "right-hand rule". In tissue, this magnetic field, in turn, induces an electric field perpendicular to itself. The strength of the induced electric field mainly depends on the rate of change of the magnetic field, which, in turn, depends on the rate of change of the electrical current in the coil. In a homogeneous medium, spatial change of the electric field will cause current to flow in loops parallel to the plane of the coil, which will be predominantly tangential in the brain.

The loops with the strongest current will be near the circumference of the coil itself, anywhere on it. Conversely, the current loops become weak near the center of the coil, and there is no current at the center itself. A more focal stimulation can be achieved by a more modern figure-8-shaped coil, producing a maximal current at the intersection of the two round components. Refer to Fig-1.

Fig-1: Illustration of magnetic & electrical field generated by a TMS coil and the difference between 2 most common coils.

Neuronal elements are activated by the induced electric field by two mechanisms. If the field is parallel to the neuronal element, then the field will be most effective where the intensity changes as a function of distance. If the field is not completely parallel, activation will occur at bends in the neural element. Axons, terminals, and branches have the lowest threshold (stimulated easily), while the cell body has the highest. Also, it's important to note that the more superficial the brain tissue, the stronger the stimulation effect.

As mentioned, the sinusoidal electric current is delivered to the TMS stimulator but the way it is designed has two main types. The first, monophasic: active only during the first peak of the sine wave. It is easier to characterize. On the other hand, biphasic: functional during both positive/negative peaks and neuronal effects are thought to be quicker, more spread out. Monophasic TMS has a stronger short-term effect during repetitive stimulation than biphasic TMS, because monophasic pulses preferentially activate one population of neurons oriented in the same direction so that their effects readily summate. Biphasic pulses, in contrast, may activate several different populations of neurons (both facilitatory and inhibitory) so that summation of the effects is not so clear as with monophasic pulses. When single stimuli are applied, however, biphasic TMS is thought to be more powerful than monophasic TMS because the peak-to-peak amplitude of stimulus pulse is higher and its duration is longer when the same intensity of stimulation (the same amount of current is stored by the stimulator) is used.

Motor Evoked Potential (MEP)
The non-invasive brain stimulation by either electrical or magnetic source is able to generate observable behavior responses such as muscle twitches (for M1 stimulation) and phosphenes (for V1 stimulation). Originally, however, such stimulations were done to evoke observable movements. The electrical signal resulting from a TMS stimulation on the motor cortex is called motor evoked potential and is typically observed by two different methods. The first one is through observing the descending volley using microelectrodes, i.e. the activity of motoneurones in the corticospinal tract. This method shows us two components:
     a)  D-waves (direct), if you hit pyramidal cells residing in the motor cortex directly.
     b)  I-waves (indirect), which originate from the indirect corticospinal neurons and interneurons.
There is usually one D-wave resulting from a single stimulation but multiple I-waves that appear later than the D-waves, depending on how many possible synapses exist. D-wave measurement can be used in a clinical setting for intraoperative monitoring.

Another way, perhaps the easiest, is through electromyography (EMG) measurement from target muscles, the muscle we want to twitch or contract. EMG measures myogenic or muscle activity, or compound muscle action potential (CMAP). In many TMS experiments, EMG is usually the preferred observation method of the motor evoked potential. Refer to Figure 2 showing typical MEP curves for biceps and FDI respectively. Note that a curve consists of negative and positive peaks. The delivery of the TMS pulse is aligned with t = 0. Note that the latency of FDI activity is slower than the biceps MEP.

Fig-2: Motor evoked potential of biceps and FDI muscle respectively after delivering a single strong TMS pulse to the arm area of the primary motor cortex (M1). The plots are produced by the BrainSight navigation system.


Sometimes, such activity is also called M-waves, "M" for muscles, that have usually a larger peak-to-peak. The excitatory postsynaptic potentials in the spinal anterior horn cells summate to bring them to a threshold and fire them. M-waves are also used in the context of reflex. E.g. M1 is the earliest or short-latency onset following a sudden muscle stretch, followed by a transcortical reflex and a voluntary component. TES predominantly generates D-waves under the stimulating anode, and therefore predominantly generates M-waves in muscles contralateral to the stimulating anode

On Finding a Hotspot
Traditionally, the easiest way to induce movements is by stimulating the "hand" area of the motor cortex and observe the hand twitching. There are two most popular target muscles used in TMS studies.
     a)  FDI, first dorsal interossei, a muscle to flex the index finger.
     b)  APB, abductor pollicis brevis, a muscle to abduct the thumb.
EMG electrodes are placed in these target muscles and a TMS pulse is delivered to the "hand" area of M1. The most crucial job comes: localizing the correct area or the hotspot. One has to patiently shift and adjust the orientation of the coil from one area to the next. Higher intensity is usually used until one is able to see a good response. Once twitching occurs, it is said that we have found the hotspot.

There are two terms associated with the TMS stimulator intensity to evoke twitching.
  1. Resting Motor Threshold (RMT) is defined as the minimum stimulus intensity that evokes a minimum motor evoked response when the muscle is at rest. As a rule of thumb, the observed MEP should be 50 µV in at least 5 of 10 trials at rest. 
  2. Active Motor Threshold: the minimum stimulus intensity that produces a minimum motor evoked response (in at least 5 of 10 trials) during an isometric contraction of the tested muscle at about 10% of the maximum force.
Coil orientation influences the amplitude of MEP (replicated by Pascual-Leone, 1992). See the figure below. The actual reason why this happens is unclear, but tDCS electrode placement has a similar characteristic.
Fig-3: Different coil orientation and MEPs of lateral-medial (LM) and posterior-anterior (PA) in different intensities.


Does TMS Cause Excitation or Inhibition?
When we talk about reversible plastic changes, TMS has been shown to excite or inhibit certain neural circuits. But in what circumstance does either occur? It seems that it depends on the pulse frequency parameter, not the intensity. A series of rapid pulses of TMS over a short period of time is known as repetitive TMS, rTMS. High-frequency rTMS with pulses at about 5–10 Hz, has been used as a more powerful stimulus to produce a brief period of excitation (Pascual-Leone et al., 1994). Conversely, slow varying pulses between ~0.2 - 1 Hz can be used to inhibit neural activity (Chen et al., 1997). Such plastic changes are thought to be mediated by LTP/LTD like mechanism, that is, a persistence change in synaptic strength (Huang et al., 2007).

However, it has been shown that the plastic effect is highly variable and lasts < 1 hr. A variant of rTMS is called theta burst stimulation (TBS), where pulses are applied in bursts of three, delivered 50 times per second (50 Hz) and an inter-burst interval of 200 ms (5 Hz) (Di Lazzaro et al., 2005, Huang et al., 2005). Based on recent animal studies, TBS is also shown to be based on LTP/LTD-like mechanisms. For example, NMDAR antagonists and Ca2+ channel blockers are shown to interfere with TBS. What is interesting is that one can induce either an excitatory or inhibitory mode using TBS depending on the timing protocol, see Figure 4.

An application of TBS twice separated by a break was found to have differential effects on MEP. For example, giving 2 x cTBS separated by a 10-minute break showed that the effect of continuous TBS can last for around 1 hour (Ridding et al.). Recent research has found the efficacy of TBS has a high variability that depends on genetic factors and muscle states; and that some people do not respond well to it (see: Suppa et al., 2016).



Fig-4: Different types of TBS (cTBS and iTBS) result in a differential effect on normalized MEPs (Suppa et al. 2016)

Heterosynaptic plasticity can be realized in humans with a peripheral stimulus paired with a TMS brain stimulus. A nice set of experimental paradigms has been developed by Classen and collaborators which is called paired associative stimulation (PAS) (Stefan et al., 2000; Wolters et al., 2003). If a median nerve stimulation at the wrist is paired with a single TMS pulse to the sensorimotor cortex at 25 ms, then the two stimuli arrive at about the same time, and the MEPs will be facilitated. If the interval is about 10 ms, however, the TMS comes about 15 ms before the median nerve volley arrives, and the MEP will be depressed. The former behaves like LTP and the latter like LTD (McDonnell et al., 2007). As a simple motor learning task and PAS interact with each other, it does appear that PAS is a highly relevant model for brain plasticity (Ziemann et al., 2004).

Some major contributions of TMS
TMS can be used to localize brain function or study the neural substrate of a particular behavior. It was originally employed to study motor behavior (being the easiest to observe), but has later been used for other sensory (e.g. visual system) and cognitive functions (e.g. working memory). For example, Wasserman et al used TMS to perform MEP mapping. The authors stimulate the scalp and systematically shifted the coil to see which body parts got impacted. Similar to studies in monkeys, they found some overlapping regions responsible for movements of different body parts. One example in the motor system is the study of the role of SMA in the production of sequential finger movements. Stimulation over the SMA induced accuracy errors in complex, but not simple, sequences. Patterns of muscle activity provoked by TMS have some physiological relevance, as these can be recognized as principal components of natural movement (Gentner and Classen, 2006).

Consolidation of a simple motor skill such as phasic pinch force was disrupted by stimulation selectively over M1, without disruption of other aspects of motor function (Muellbacher et al., 2002). Another study failed to find a similar disruption of learning of motor adaptation in a force field, suggesting that only some types of motor consolidation occur in M1 (Baraduc et al., 2004). On the other hand, rTMS of M1 prior to learning of force field dynamics did interfere with consolidation without interfering with the learning itself (Richardson et al., 2006). More research has to be done on this theme.

TMS is also beneficial to understand other sensory behavior. For example, studying the visual cortex with TMS helps to understand how inhibiting the region impacts object recognition, disrupts motion perception (on V5), or reading ability. Others use TMS to understand working memory. For example, stimulating left DLPFC impairs working memory of alphabets but not of faces. Low-frequency rTMS over either the right or left prefrontal cortex (but not the parieto-occipital cortex) impaired behavior on a task involving visuospatial planning. TMS has also been used in conjunction with functional MRI to see if such stimulation changes the time course of the BOLD activity. Lastly, rTMS has been a popular non-invasive method in a clinical setting to treat major depression and, more recently, in stroke rehabilitation.

[Main source = a primer article by M. Hallet, 2007, in Neuron journal].

Sunday, April 2, 2017

Motor Learning: Behavioral Emphasis (Part II)

Introduction
Motor learning is a branch of the study of motor behavior. It cannot be separated from motor control such as muscle coordination among different body parts, e.g. eye-head-hand coordination, and the concept of executive control. There are various ways of defining motor learning:
Motor learning is defined as changes in internal processes or states, associated with repeated practice or experience, that determine a person's capability for producing a skillful movement. 
The definition above can be expanded into four properties. First, motor learning is a set of processes acquiring capability. Second, such processes are internal and not directly observable, so learning has to be probed systematically. If internal (psychological) states produce a set of motor behaviors, then behavioral changes are expected as a result of learning. Third, it is a result of repeated practice or experience (W. James called this a habit). Finally, these processes are relatively permanent, for example: a child who learns to play tennis is still able to do after a long period of break. Successful motor learning usually involves goal setting such that learners know what and how to perform a particular task.

Increased capability for moving skillfully in a particular situation defines learning. For example: the goal of learning tennis is to serve properly, make the ball enter the correct region, and direct the ball to a position difficult to reach by the opponent. The "quality" of the internal states that produce the movements is maximized as a result of motor learning. This definition is more specific, as opposed to a more general and cognitive definition of learning, i.e. a process that results in a change in behavior.

Strictly speaking, there is a distinction between motor performance and motor learning. Performance is talking about motor execution. Change in motor performance can also show changes that are not learning-related because it is only temporary, e.g. it decreases due to fatigue or is enhanced by dopping. People also make a distinction between ability vs. learning. Whereas ability can be due to growth or maturity and reflects personal traits, motor learning is especially due to repeated practices. The distinction between performance and motor learning is what makes us require a certain set of behavioral paradigms. Such paradigms allow us to measure an increase in performance even after a long pause or break of practice. See the next part: retention and transfer.

According to Fitts and Posner (1967), there are essentially 3 stages of motor learning:
  1. Cognitive stage: for a naive learner, the problem to be solved in the cognitive stage is understanding what to do. This stage is also known as the verbal-motor stage (Adams, 1971) as it involves the conveyance (verbal) and 'thinking' (cognition) of new information. There is a large gain, but inconsistent, the profile of performance.
  2. Associative stage: it is a stage of dwelling deeper into how to perform the skill; characterized as much less verbal information, smaller gains and conscious performance, a lot of corrective and adjustments. This stage is also called the motor stage proper (Adams, 1971). From the cognitive perspective, the novice is attempting to translate declarative knowledge into procedural knowledge. It's about transforming what to do into how to do.
  3. Autonomous stage: the final stage of motor acquisition where performance becomes largely automatic, where cognitive processing demands are minimal (no 'thinking'). For athletes, this is when they can grip it and rip it, look and automatically react, and enter a state of flow.
Motor learning thus involves stages from a more cognitive in nature to a less cognitive, but more sensorimotor. In the language of memory and learning, this is a shift from declarative to procedural processes.

How to Measure Learning?
In a typical motor learning experiment, two or more groups of subjects practice a task under a different level of an independent variable, i.e. the behavior task. The most common method for analysis is using a learning curve. The learning curve can take different measured variables (dependent, response), which depend on the type of studies conducted, e.g. in terms of a reduction in reaction time in sec., increase in accuracy in cm, or the number of correct scores received. There are a few considerations in averaging learning curves. First, the learning curve should depict learning not merely performance. Finding the average is likely undermining between-subject differences and differences in strategy, the latter being a more difficult confound to take care of.  Another aspect to consider is the within-subject variability caused by motor noise that may directly impact the measure of learning. Ceiling and floor effects also impair the measurement of learning such that further improvement or increase in performance is impossible. This is when learning has reached an asymptotic level or plateau.

The power law of practice states that the learning scores or index for a particular task increases linearly with the logarithm of the number of practice trials. The consequence of this law is that our trial-to-trial improvement isn't all the same or linear. This improvement is generally very quick during the first few trials, then it slows down. The law is generally true for all motor learning tasks if you take the average. There is an on-going debate whether the equation to model this is exponential or logarithmic, etc., and whether individual differences occur. One subject may reach a plateau more quickly than others.

There are a few practice paradigms to study motor learning. The first paradigm trains subjects at different levels of the independent variable then transferred to a common level of that variable. The design provides a separation between a relatively permanent effect (learning) and a task-dependent effect which is temporary and related more to performance. What happens when the ceiling or flooring is easily reached? One way is to incorporate a secondary task or measure related but different dependent variables. Another way is to measure motor automaticity and effort. After the learner reaches an asymptote, further improvement in accuracy is impossible. Another dependent variable is required, e.g. we can measure whether the reaction time improves (more automatic), or oxygen consumption reduces (less effort required).

Perhaps, as mentioned, the most fundamental way to probe motor learning is by studying it in terms of retention and transfer. Both tests are performed following a reasonable break or interval, after an initial "acquisition" or learning process. Retention is a measure of how well the changes persist following the initial learning of the same task. It is tested by calling the same subjects again after a long break of e.g. 5 or 10 days, to do the same task learned during "acquisition" earlier. Retention is related to saving, e.g. Ebbinghaus (1913), Nelson (1985), that is, a faster relearning. Motor generalization is related to the transfer of learning, that is, the effect of learning observed in another context. Technically, when generalization is beneficial, it is termed transfer (if not, it acts as interference). A practical example of transfer: if you learn to play tennis after a while, how is that skill useful for you in learning badminton? The specificity of the learning hypothesis says that we should attempt to match those conditions in practice with those used during the test or retention period.

Off-task and On-task Practices
Practice or training is the most important determinant of the so-called motor learning. There are some consequences of this statement. First: learners have to be motivated to practice. Second: to be motivated, they have to understand the goal or purpose. Psychologists found that goal-setting is a widely used motivational technique (e.g. Locke & Latham 1985). In sports psychology, specific and moderately difficult goals are more beneficial to learners. Third: in order to achieve a clear and unambiguous set of goals, verbal instructions are necessary. Instructions influence a certain level of attention that is task-specific. E.g. in a balancing study with both hands holding a tube, subjects that receive instructions to keep their hand horizontal (body parts, internal focus) and control group have the largest error, while instructions explicitly asking them to hold the tube horizontal yield the smallest error (Wuff et al. 2007). 

Factors mentioned above are called off-task practice conditions as they are indirect practices, that is, they aren't about actively performing the task itself. Other forms of off-task practice include mental practice and perceptual learning. In perceptual learning, a learner goes through a period of "experiencing" sensory events related to the task, e.g. what he/she will see, feel, and touch available during the task performance. A more specific form of perceptual learning is observational learning. This learning was originally more of a form of social learning in children proposed by Bandura and has been confirmed through animal studies on mirror neurons. In motor skill acquisition, observational learning happens when the learner watches an ideal model, skilled performer, or his/her instructor performing a demonstration of the task (for review, see a book chapter by Maslovat et al., 2010). Spatial structure and timing are the two most important components learned during learning by observing.

On-task practice conditions cover various types of practice with the aim of maximizing learning. For example, in terms of structure, distributed practice (broken up into a few shorter sessions over a long period of time) tends to show better learning than massed practice (done with longer sessions, or without an apparent break or rest in between) does, although these effects are seen to be stronger for the learning of continuous tasks. Refer to the figure below.

Other condition includes practice variability. It refers to the variety of movement and context characteristics the learner experiences while practicing a skill. Varying task sequences from trial to trial are more effective than constant practice conditions. In Shea & Kohl's experiments, subjects practiced creating a goal force by squeezing a handgrip connected to a force transducer. One group, the "constant group", experienced 100 trials of a constant task goal of 150 N. A second group, the "variable group", experienced a series of different task goals (100, 125, 175, 200 N, including 150 N!), hence a total of 500 trials. A third group practiced 150 N for the same # trials with the variable group. Although the variable group did the task rather poorly, they performed well during the retention period after a break. Shapiro & Schmidt observed an important phenomenon where children are always benefited from the variable task sequence, presumably because the schemas are not established yet.

In real life, motor skill learning usually involves various task goals that mimic the "variable" group mentioned above, e.g. physicians practice different motor skills related to surgery, musicians practice multiple songs at a time, tennis players practice serving and volleying as well as the more usual groundstrokes during a single session, and etc. Suppose a doctor has to learn suturing skills 1, 2, 3, and 4 How can we schedule them so as to maximize learning? There are two ways to do this: random practice (interleaved tasks 1, 2, 3, ...) and blocked practice. (complete practicing task 1 first, then proceed to 2, ... and so on). Note that these skills have one similarity or context, i.e. applying sutures. Although they have the same context, the motor components are not.

A term called contextual interference was introduced by W. Battig to name the effect of task differences while maintaining the same context (Shea & Morgan, 1979; a review by Brady 1998). Random practice design has high contextual interference. Several studies have shown that random practice has an impact on reducing the performance during the acquisition or learning phase but leads to more effective learning than blocked practice, as measured by the retention and transfer tests. Why is this so? Learning motor skills involve working memory of how to do the task well ("when forgetting improves remembering at a later time"?). Contextual interference has been replicated to a certain extent in more complex tasks and motor skills outside of the laboratory (e.g. Wulf & Shea, 2002; Goode & Magill, 1986; Albaret & Thon, 1998). Nevertheless, there is a limit to the generalization of contextual interference to motor tasks that are relatively simple.

Extrinsic Feedback: KR and KP
One of the most important features of practice or learning is the information the learner receives about their attempts to produce a movement. This is called movement-produced feedback and it tells the quality of our produced movements, the error or mistake, etc so that we can learn to correct them. There are basically two types of feedback:
     (1) Internal or intrinsic feedback: somatosensory, visual, and other sensory feedback.
     (2) External or extrinsic feedback or augmented feedback e.g. reward, verbal feedback.

The second type, augmented feedback, can be divided into KR and KP. A focus of this discussion is a type of external feedback called the knowledge of results (KR) where it provides post-movement information about the outcome of the movement in the environment. In practice, KR can appear in the form of a more abstract binary signal (right/wrong), visual or verbal reward signal ("Good job!"), or the amount of error produced ("the speed was too quick", "the endpoint was 2 cm too long"), etc. This is in contrast with the knowledge of performance (KP),: the information about how you perform the movements, e.g. "You bent your arm", "Your body was too stiff". In a practical sense, KP deals with how well you perform the movements, but what makes KR more popular than KP in studying motor learning? Because changes by KR is more easily measured.

The KR paradigm is used heavily in the field of behavioral and experimental psychology such as in studies by Pavlov, Thorndike, Tolman, etc. on conditioning and shaping. Thorndike is probably remembered for his KR/no-KR paradigm in motor learning. Essentially. the paradigm lets a participant learns a task with KR and then the same person is subjected to a transfer test where the KR is removed. This makes sense. For example, in the rehabilitation setting, patients are trained with KR given but tested without KR to simulate the real-life situation outside the clinical setting. For a review on KR in the rehabilitation setting, see Winstein C. (1991) and van Vliet &Wulf (2006).

Studies from the last century have shown how KR influences motor performance by giving "energizing" state, rewarding stimulus, and attentional reference. It provides guidance on what to do next (Salmoni et al, 1984) but has a limited effect on learning itself (e.g. see Szalma et al, 2006). But earlier works by Bilodeau et al. show that KR does not only influences performance but in itself a learning variable (an indirect way of saying KR causes learning). The authors trained a group of subjects without KR, another group with KR throughout, and the other group in between. During the practice period, the KR group had a rapid reduction in absolute movement error. The no-KR group consistently had a much higher error. Following this, the no-KR group performed another extra 5 trials with KR and their error performance is similar to the first 5 trials of the KR group.

KP may appear in various forms. An instructor can let the novice players see their own performance via video feedback. Dance instructors can include kinematic feedback such as "Move your arm to face sideways more quickly!". Kinematic feedback talks about motion trajectories and velocity, while kinetic feedback takes into account how much force exerted to perform the task. Some argue that KP directs the learners to focus on a more internal state of information, e.g. how they control their arm, versus KR which is more on the external state of information, e.g. "You are 10 mm undershoot the target location", "That's a good shot!".  The effectiveness of kinematic KP compared to KR depends on the task goals. The importance of kinematic KP is seen when some movement patterns are otherwise too difficult to perceive. There is evidence showing that KP can contribute to learning specificity. E.g. a study by Levin et al (2006) shows that a group of patients that received KR improved in their aiming (spatial) accuracy but not speed. On the other hand, patients that received KP on their shoulder/elbow velocity improved the velocity accuracy. Lastly, another form of KP is biofeedback, the most popular of which uses EMG signals.

More about KR
Suppose one is doing a ballistic movement to a target. The KR can be in the form of endpoint error (quantitative) or the direction (more to the left, undershoot, etc) or both. Another KR type, called bandwidth KR (Sherwood, 1988), is determined by bandwidth or range about the target or movement goal. If the error is within the target bandwidth, a binary KR (qualitative) is sufficient. If the performance is so bad for a prolonged time, then the instructor would give both the amount of error and the direction that is shown to enhance learning. In his study, Sherwood asked participants to make rapid elbow movement within the desired movement time, (MT = 200 msec). One group was told the exact MT as a KR following each trial. A second and third group received a KR if their MT exceeds ±5% and ±10% bandwidth of the target MT = 200 msec. Note that KR here means an indication of a negative or error KR. After blocks of 25 trials, all groups went through a retention test. He found that the 10% bandwidth group showed the smallest temporal error. It seems that less frequent error KR helps motor learning better.

But is the improvement in retention due to less frequent error KR? Lee & Carnahan (1990) studied a similar paradigm using two groups: a bandwidth group and a yoked control group. The control group received error KR exactly on the same trials as the bandwidth group. The difference is that in the actual group, no KR indicates that the previous trial was correct, while in the control group, no KR has got anything to do with their movement outcomes. They found that the bandwidth group performed better in retention. This suggests that no KR (or correct KR) provides an additional boost to learning on top of the less frequent error KR. Moreover, bandwidth KR facilitates learning in the observational learning task, supporting a high cognitive component to the provision of “correct” feedback (Badets & Blandin, 2005).

The way learners interpret error or correct KR may differ across time. During acquisition or practice blocks, as learning continues, the proportion of correct KR and error KR ideally increases. Thus, providing a constant correct/error KR ratio is not an effective paradigm (Lai & Shea, 1999).

How often do we provide KR? Motor learning researchers contrasted the relative and absolute frequency of KR. The absolute frequency of KR refers to the number of trials KR is given. Suppose, there are 50 trials and of those trials, 35 trials are with KR. The absolute frequency is 35 but the relative frequency is 35/50 = 70%. If the total trials doubles, the absolute frequency becomes 70 but the relative frequency remains the same. Early researchers thought that the relative frequency of KR is an irrelevant variable of learning. Later studies using the transfer paradigm show that both frequencies are important factors, which means no KR trials contribute somewhat to learning. Decreasing relative frequency does not suppress learning but enhances it (Winstein & Schmidt, 1990). In their study, the authors found that 50% and 100% groups don't differ in their performance during acquisition but their 5 min and 24 hours retention test favored the 50% group. Giving KR that is too frequent deter the performance because while on the one hand, it gives a motivational and information boost, the learners become so dependent on it until they neglect other inherent feedback e.g. somatosensory information. The over-reliance on KR is detrimental during motor tests later when KR is absent.

When should we provide KR? Since Thorndike's era, people thought that delaying reinforcement degraded learning in general. It appears this is not the case in motor learning, where delaying the KR presentation has a non-significant impact. Interestingly, it was found that if KR is presented too early, it can have a detrimental effect to learning (Swinnen et al, 1990). Filling the gap with extra stimulus or activity during both KR delay and post-KR interval (the short gap between the presentation of KR and the next trial) also degrades learning. It is thought that this occurs because the learners are not able to process the information provided by the KR and also by their inherent feedback. This is when KR blocks other critical information required for learning. Another instance where learning is degraded is when KR causes maladaptive correction. This is the case, e.g. in fast-reaching to a target when learners have reached the asymptotic phase and no further accuracy can be achieved. Our motor system is noisy (motor variability) and KR is thought to introduce unnecessary correction of an error caused by this noise.

Two Major Theories of Motor Learning
Let's move back to the 80's. One major theory of motor learning comes from J. Adams who used a set of empirical laws of motor learning based on slow, linear-positioning movements. He believed that on-going feedback from the limb is a key to learning, making motor learning inherently a type of close-loop process. This feedback -- called the perceptual trace -- provides a reference of correctness that is stored in the memory. During practice, KR serves a purpose to strengthen this perceptual trace in the memory. The sense of correctness, or thus a sense of wrong directions, get accumulated as the trial continues. KR also helps to guide subsequent movements. The learner strives to close the gap between the on-going inherent feedback and the prior perceptual traces. He claimed that the error-detection capability occurs through comparing the on-going inherent feedback and the perceptual trace.

Soon after Adams, Schmidt proposed an improved model that can be applied to both slow and rapid movements. According to Schmidt, motor learning involves creating rules or motor schemas. Learning first starts by selecting a generalized motor program (GMP) containing muscle commands that have invariant features. Then, the learner adjusts various parameters to produce a necessary movement. After the movement is performed, there are 4 types of information available for storage in the memory:
    1)  Information about initial conditions: posture, the weight of an object thrown, etc.
    2)  Parameters assigned to the GMP.
    3)  Augmented feedback of the movement outcome.
    4)  Inherent feedback from the body: proprioception, visual, audio etc.
From these 4 sources of information, the learner then continuously build and update two schemas:
    1)  Recall schema: to produce subsequent movements (updating GMP parameters with time).
    2)  Recognition schema: to evaluate movements just performed (sensory consequences).

Thus, the theory consists of 3 components: GMP, recall, and recognition schemas. Evidence of schema theory in real life has been outlined by Schmidt (chapter 2, Motor Control: Issues and Trends, 1976). Both major theories mentioned are not without criticisms. I think this is because scientists strive to refine the motor learning models that are able to explain all behavioral principles. Another popular approach to model motor learning is through using cognitive principles and degree-of-freedom problems (that of Bernstein's).

Retention and Transfer
In cognitive science, the concepts of learning, memory, retention, and transfer are very closely related. Motor memory is the persistence of the acquired capability for doing the motor task. The retained portion can be measured directly through recall and recognition tests. Such tests are usually performed after a certain time interval. The retention of motor learning can also be measured indirectly by looking at saving. For example, if one requires 50 trials to reach a criterion performance during early learning, but 25 trials during the retention test, the saving is computed to be 50%.

By definition, losses in memory are called forgetting. While learning can be measured directly, forgetting is measured indirectly through the performance loss following a retention interval. Studies have shown the absolute-retention measure is the most useful one. But the interpretation of such measures can be not as straightforward. For example the decaying effect of forgetting following a retention interval of person-A is slower than that of person-B. But is unclear whether it's due to a slower learning of person-A (and thus slower forgetting) or more retention capability.

Is retention profile uniform across different motor tasks? It appears that continuous skills are retained nearly perfectly over a long retention interval, whereas discrete skills can exhibit marked losses during the same interval. To recap, examples of a discrete skill are kicking a ball, throwing a dart, rapid reaching to an object. Continuous skills include swimming, jogging, and tracking task. Why is this so? Perhaps this is due that continuous skills are more basic and low-level and learned more completely.

The loss of motor memory can be triggered by passive decay processes. It can be due to active interference in the form of proactive and retroactive. Consolidation studies suggest that the interfering effects of learning a competing task are time-dependent.

A variant of a learning experiment is a transfer experiment where the effect of the practice of one task on the performance of some other task is evaluated. A common paradigm would be as follows: there are 2 independent groups (treatment group versus control). In the treatment group, subjects practice task-A, but tested on task B. In the control group, subjects do not do any practice but tested on task B.
Transfer of learning can be near (among almost similar tasks) or far (more apparent differences in tasks). Although the general consensus is that motor transfer is small, it is still debatable whether the transfer characteristics of all types of motor learning are the same. In real life, e.g. practicing tennis aids you in learning badminton as both sports are using a racket. The transfer is often measured as a percentage, indicating the proportion of performance improvement in one task that was achieved by practice on the other task. Based on this, a positive transfer means the practice of a task helps in learning another task. A negative transfer means the practice of a task hinders the learning of another task.


Reference: The writing and diagrams are based on a textbook, "Motor Control & Learning: a Behavioral Emphasis, 5ed" by RA Schmidt & T Lee.

Tuesday, March 14, 2017

Logistic Regression: with MLE and more

The following post outlines the concept of logistic regression with the well-known Generalized Linear Model or GLM. Detailed maths is out of scope and we'll just make use of R to aid our learning. To begin, logistic regression is used to model and predict data with the following properties:
     • A categorical dependent variable, e.g. cure vs sick, old vs young (binary values)
     • A continuous, quantitative independent variable

In the discussion herewith, the response variable has 2 possible values, e.g. male/female, success/failure, dead/alive, etc. Our main goal is to predict the probability of either value occurring. A series of such binary responses (Bernoulli process) follows a binomial distribution. Probability density function (pdf) tells us the probability of a particular value for a fixed or known model or distribution, e.g. binomial distribution. Suppose the probability of having Y is denoted as P(Y) or p, where 0 ≤ p ≤ 1. The ratio of having Y and not having Y, p/(1 – p) is called odds. Thus, instead of simply predicting Y from X as in the linear regression, we predict the odds of Y for a given X. 

Logistic Regression as a Special Case of GLM
A simple logistic regression related to P(Y) can be expressed in the form of Equation (1). The word 'logistic' comes from the logistic function of probability shown on the left-hand side. This function can be transformed to the natural logarithmic function of odds called logit as shown on the right-hand side. The transformed model now takes logit as the dependent variable, which reminds us of the linear regression equation. Although the assumption of linearity (and others e.g. normality, homogeneity of variance) does not hold for logistic regression, the linearity between independent variable and logit has to be maintained! An extended version with multiple predictors can be seen in Equation (2).
Fig-1: (Left) An example of a logistic function σ(t) that takes any value between 0, 1. The function has a unique sigmoidal shape. (Right) An example of logistic regression with different values of betas [figure taken from ips8e, supp-Chapter 14, Macmillan].


When we talk about the simple linear regression, we assume X and Y to be linearly related more or less, the random error or residual follows a normal distribution, and the assumption of independence holds true. What happens when some of the assumptions collapse? "Generalized Linear Model" attempts to provide a more general form of regression when such assumptions are not fulfilled. It's sometimes confused with the general linear model found in the ANOVA and neuroimaging analyses. Now, there are three components of GLM:
     • An exponential family distribution for residuals.
     • A set of linear predictors that represents a systematic component.
     • A function that links or connects the expected value or mean of (1) and (2).

In GLM, the outcome variable takes up an exponential family distribution, e.g. normal or Gaussian, beta, gamma, exponential, chi-squared, Poisson, Bernoulli, and Dirichlet distribution. In the case of the ordinary linear regression, it is the Gaussian distribution. Suppose we predict the expected value of the outcome variable as shown by its mean μ. GLM allows a function of μ rather than the mean itself. This function, called the link function, is denoted as g(μ) because it links the mean of the outcome variable to the set of k predictors or explanatory variables.
      In summary, GLM formula states that:    g(μ)  =  a +  b1X1  + b2X2  + b3X3 + .... +  bkXk
In the case of linear regression, the link function is the identity link where g(μ) = μ. For binary data such as success/failure, a logit link is used instead. This is the case of the logistic regression. In R, logistic regression can be done by using glm command, instead of lm( ). Now, of so many exponential families, is there a unified way we estimate the parameters in GLM? Yeah, using a technique described in the next section, MLE.

Maximum Likelihood Estimation (MLE)
How do we estimate b? Whereas linear regression uses the least-squares method to minimize the errors, logistic regression uses the maximum likelihood method to arrive at the solution. Key to MLE is really just the "likelihood". Here, we first assume there is a model and this model is valid/correct. If it's correct, we can find the most likely estimate that explains the observed data.
Probabiliy means given a model of an event, what's the distribution of the outcome? Likelihood means given a set of observed outcomes, what are model parameters? 
Undergraduate stats course teaches us binomial probability. For example, suppose we toss an unfair coin 20 times and trials or tosses are independent of each other. In each trial, the probability of a head is p = 0.6, so the probability of not getting a head is 0.4. We try to describe the probability of getting heads r times out of n tosses. In fact, we may obtain a total of r = 0, 1, 2, ..., up to 20 heads from tossing that coin 20 times. We compute,  0C20p0(1 – p)20 ,  1C20p1(1 – p)19  , ... up to 20C20p20(1 – p)0. Refer to Fig-2. The plot on the left is the binomial probability density function. The shaded area under the curve equals 1.


Fig-2: (Left) Discrete probability density function of tossing a coin 20 times, with the probability of obtaining head in each toss p = 0.6. The X-axis is the number of possible heads from 0 to 20. Obviously, if observing a head in each toss is 0.6 then I would expect 0.6 x 20 = 12 heads in 20 tosses. (Right) Suppose we were to estimate p here based on the observation that there are 13 heads out of 20 tosses. The best estimate of p is the one that has highest likelihood, in this case, the peak is at p = 0.65.





The step shown above is actually generating a possible data set from known parameters (n, p). What if the process is reversed? Twist your mind! We want to estimate p for a given series of n observations or outcomes. A likelihood function is exactly this, i.e. the opposite of probability density function. So, if we observed 13 heads in tossing the coin 20 times, what is the probability of a head in each coin toss? We try to estimate the value of p such that our observation is most likely to happen. Theoretically, if our model is correct, the most likely value of p = 13/20 = 0.65. The right panel Fig-2 shows this is the case. The highest likelihood is when the X-axis or p equals 0.65.
L(data | model) is a function that describes the probability of a model given our observed data set, p(model | data). If independence among trials is assumed, the likelihood function equals the joint-probability density which is none other than the product of each and every p(yi | xi), where yi takes up a binary value.
As the end result of multiplication may be too small to manage, people take the natural logarithm of the likelihood which is then called log-likelihood (LL). Twist your mind again, taking the log of multiplication yields a summation of logs. Look at Equation (3). Now, since the numerical optimizer is designed to find the minimum, not the maximum, we multiply this by –1. Equation (4) is the final equation we want to minimize, just like minimizing sum of squares residuals in linear regression! Numerical optimization is used as there is no closed-form analytical solution. In R, this can be achieved by using optim() function.

Fitting a model to our data in logistic regression means to estimate the coefficient b0 and b. The first is the intercept, the value of Y when X = 0. But how do we interpret b? The logit part is our dependent variable in the logistic regression. The logit is also easily converted back into the odds, where odds = exp(b0 + bX). Suppose we have X1 and X2 = X1 + 1; we compare both odds by finding the ratio of odds(X2) = exp(b0 + bX1 + b) and odds(X1) = exp(b0 + bX1), which is just equal to exp(b). Hence, the coefficient b can be interpreted as the log of the relative increase in the odds for a unit increase in score X. 

Assessing the model fit
There are a few common measures of model fit in the logistic regression. The first is the log-likelihood statistics as mentioned previously, i.e. summing the probabilities of the predicted and actual outcomes. This is analogous to the residual sum of squares (SSR) in multiple regression in the sense that it is an indicator of how much unexplained information there is after the model has been fitted. The larger the value, the poorer the fitting is. A similar measure called deviance statistics is more popular, where deviance = –2 LL. Again, a higher number means a bad fit. We'll talk more about deviance later.

The second measure is Pearson's goodness of fit test. This is a statistical test to investigate how close values predicted by the model are to the observed values. The null hypothesis here is that our data follow the logistic regression model. We first find the chi-squared (c2) value and then the degree of freedom to find the p-value. The third performance assessment is by R2, which is related to both the deviance and z2 (Wald statistics). This measure is similar to the one we saw in the linear regression but with a different formula. R2.can be unnecessarily inflated when we add more predictors. Like in multiple linear regression, we can use BIC or AIC to compute the parsimony measure of R2, penalizing for useless extra predictors in the model. AIC/BIC is useful for choosing a model, with the preferred model is the one with a lower AIC/BIC index.

Software packages such as SPSS and R make use of the deviance statistics in the report. More importantly, there are two types of deviance. The null deviance shows how well the response is predicted by a reduced or baseline model. In the case of the simple logistic regression, the model assumes that the outcome can be predicted by nothing but a constant value, that is, the intercept b0. In other words, it agrees with the null hypothesis that the estimated b = 0. The residual deviance shows how well the response is predicted by the full model when all predictors are included, implying that this is against the null hypothesis. Each deviance carries its own # dof.
The difference between the null and residual deviance is called likelihood ratio, another goodness of fit that also has a chi-squared distribution. The degree of freedom is defined as the difference between two deviances. We then easily get the p-value that suggests whether the model is better than chance at predicting the outcome.
How to assess the importance of the individual predictor in the logistic model? Similar to linear regression, it is assessed by carrying out statistical tests of the significance of the coefficient. Whereas t-test is used in linear regression, Wald test (z2) is used to evaluate the statistical significance of each predictor. It is calculated by taking the ratio of the square of the regression coefficient to the square of the standard error of the coefficient. The null hypothesis is similar, i.e. whether the estimated b is different from zero. As a side note: Wald test has been shown to be less reliable for small sample sizes!

Fig-3: Two R outputs from a logisitc model operation where you have one predictor (intervention) and two (intervention + duration) to predict the same dependent variable. Click to enlarge!


Let's take a look at Fig-3. Suppose there is a data frame containing three fields: whether the disease is cured (dependent variable), whether any intervention is performed (first categorical predictor), and the sickness duration (another, but quantitative, predictor). The model on the left fits the outcome variable (cured vs. not-cured) with whether or not treatment has been performed. What does R show us? First, the coefficients portion suggests there is a significant intervention predictor with estimated b = 1.23, z = 3.07, p < 0.005. This is Wald statistics. A z-value that is sufficiently far from 0 means that the estimate is both precise enough to be statistically different from 0 and large to have an effect on the response.

Next, the model fit. The null deviance = 154.08 and residual deviance = 144.16. The fact that residual deviance is lower suggests that one more predictor is better than just a constant in reducing the residual error. So adding "intervention" improves the fit by 9.93 unit. The difference in # dof of 1 so we can find the p-value of the chi-squared. Accordingly, c2(1) has p-value = 0.00163. The model on the right introduces another predictor, i.e. total duration. Based on the deviance values, there is no benefit of adding one more predictor. The book says that anova() can be used to compare both model fits. The test yields no significant difference between the left and right models.

Special example: Psychometric curve
A fundamental concept in psychophysics, the psychometric function relates a parameter of a sensory stimulus to a subjective response of a participant. We learnt previously how the psychometric curve has a sigmoidal shape and it's getting much clearer why. Take an example of a two-alternative-forced-choice (2AFC) task. What we're interested in is how the probability of responding with one of the two choices varies with the stimulus. Here, we will consider a situation where the varying sensory stimulus is the movement direction θ of a participant's arm to the left or right of the body midline. This becomes the continuous independent variable X. The desired response is binary, either left or right. This response is our dependent variable Y.

The model suitable for this is the simple logistic regression shown in Equation (1). Let p = P(y = right | X) is the probability of responding "right" given a certain direction X. So, 1 – p is equivalent to P(y = left | X). We typically represent the response in numerical format, 1 = right and 0 = left. The value X in which p = 0.5 corresponds to the perceptual threshold that in ideal case is the body midline, but one may have a perceptual bias. From our experiment, we obtain an array of direction X (in degree) and an array of binary responses R (in number 0, 1). First construct our cost function, name it NLL. This is then fed into a numerical optimizer to estimate b0 and b using MLE. An initial guess for each estimate has to be given, e.g. (0.1, 0.1).
NLL <- function(B,X,R) {
          y = B[1] + B[2]*X
          p = 1/(1+exp(-y))
          NLL = -sum(log(p[R==1])) - sum(log(1-p[R==0]))
  }

out = optim(par=c(-.1,.1),NLL,X=X,R=R)    #Numerical optimization! 
Bfit = out$par    #Retrieve parameter estimates of our model

#Let's construct our model with Bfits above and plot the data...
Xp = seq(-15,15,1)
myModel = logistic(Bfit[1] + Bfit[2]*Xp)
PS: I have recently found another nice R package called "quickpsy" for computing with psychometric functions.


Main references
(1) Andy Field, et al, "Discovering Statistics using R" (2012).
(2) Gribble's note on MLE, Psychology_9041B course.

Wednesday, February 22, 2017

Motor Control: Behavioral Emphasis (Part I)

An Overview
I have briefly presented the modern theories of motor control & learning many months ago. Although such concepts appeared after the '90s, the field has been in existence ever since the beginning of the last century. The current post is meant like a historical summary that stems from behavior perspectives. People began to ponder the basis and characteristics of movement productions as far back as 19th century. The field was heavily influenced by two separate but related fields: psychology (which was dominated by behaviorists) and the birth of neurophysiology (the study of the nervous system through electrophysiological recordings in animals).

Woodworth (1899) was one of the earliest pioneers in studying rapid arm movements and laid down the foundation of measurement such as movement speed, performance error, etc. Edward Thorndike (1914), in his Law of Effect, proposed how actions that are rewarded tend to be repeated. His ideas gave the foundation of instrumental conditioning in the field of psychology. Almost during the same period, Charles Sherrington talked about the concepts of reflexes, the final common pathway (alpha motor neurons), and sensory receptor of movements in which he coined the term proprioception. Moving to Eastern Europe, a Soviet scientist N. Bernstein made an influential contribution during 1930s where he called motor control as a degree-of-freedom problem (redundancy problem). This is because our limbs consist of various joints and each joint is connected to hundreds of muscle fibers that can be active separately. Correspondingly, different sets of motor activities capable of producing the same behavior are called motor equivalence. How does the brain know which muscle(s) to control among the various possible combination? A decade later, K. Lashley did studies on handwriting (1946) where he suggested the concept of a motor program inherent in each voluntary movement. This idea suggests the movements are based on an open-loop concept, undermining the role of sensory feedback. After WW-II ended, advancement in mathematics and information theory helped the formulation of a speed-accuracy trade-off by Paul Fitts (1954, 1964). Using a neat methodology, he discovered the link between movement speed and accuracy, now known as the Fitts' Law.

At the end of 1950's, the field of psychology shifted to themes in cognitive science that talked about attention and memory, a new euphoria for the scientific community. The concept of higher-order brain functions emerged and the flavor of motor behavior research shifted. Soon after, a strong interest focusing on "learning" appeared. In 1971, Jack Adams proposed a concept of closed-loop theory of motor learning in addition to other studies, e.g. short-term memory of movements. Mike Posner (1969) studied the role of attention, short-term memory, and movement control, expanding the concept of short-term memory storage and motor behavior. Fitts and Posner (1967) perhaps were known for their concept of three stages in motor skill acquisition. The role of attention on motor control and learning was also studied by S. Keele. His motor control thesis on the motor program was also quite influential (1968, 1986).

By the end of 1980s, integration of motor behavior and sports science gained momentum with goals of understanding motor skill learning, maintaining, and maximizing performance (e.g. F. Henry, John Whitting). At the same time, the field of neurophysiology also gained maturity in both animal studies and clinical works. This is the precursor of modern neuroscience. Rather than focusing on observable or products of behavior, scientists tried to elucidate the role of the brain or nervous system in performing movements. For example, Merton & Merton, Ian Boyd studied muscle spindles; Evarts and Georgopoulos respectively studied the neural discharge of a single neuron and ensemble of neurons in the motor cortex in awake behaving monkeys; Milner, Tulving, Tolman studied long-term memory; and Teuber for modern neuropsychology.

Interests in Bernstein's muscle coordination and redundancy problem reappeared in 1980's. On two separate occasions, Feldman and Bizzi came forward with his equilibrium point hypothesis to explain motor control. Another scientist, Latash proposed an improved theory of muscular coordination (or synergy). He said that the degree of freedom problem is solved in the brain by controlling a set of muscles doing the same job. The theory has been expanded: rather than the synergy of different neural circuits controlling movements, it refers to the synergy of the pattern of coordination such that the movement outcomes are stable and at the same time flexible. See Latash et al. (2007) for a nice summary.

Open-loop Processes and Motor Program
William James (1890) said that movement control is born out of a muscular contraction in response to either an external or internal event. This contraction produces a set of sensory feedback (now known as proprioception) from the muscles which in turn triggers another muscular contraction and so on. This sequence of events is also called the response-chaining hypothesis. In skilled movement, attention is needed for the initiation of the first action and subsequent series of actions can be 'automatically' running. The fundamental element of learning is by associating given feedback with the next action. This is the first proponent of an open-loop motor process. Our brain creates the very first muscular contraction in an effector or limb. There is no output monitoring, in the sense, no error correction. If something goes wrong or the environment changes, open-loop control can do no corrective actions. In contrast, a closed-loop process is used when we perform the online correction. In this particular case, afferent feedback provides information on the movement outcome. The discrepancy between the actual and planned (reference) movement is the basis for error correction.

Studies with deafferented animals and patients have shown that sensory feedback from the muscles is not critical for motor control. Although the trajectory isn't as smooth and accurate, movement production is still possible. James' theory is therefore not universally complete. Still, though, there are other experiments that may point to the existence of an open-loop executive controller. For example, the central pattern generator (the most popular experiment is the one involving decerebrated cats on a treadmill [FV Severin, et al. 1966]) and reflex responses. Furthermore, because sensory processing is slow, how can rapid movements be executed other than through an open-loop process?
Another example includes a study involving rapid elbow extension a 2-dof structure (Wadman, et al. 1979). Subjects were asked to perform an arm extension each trial. Such a simple but rapid action was shown to involve two agonist-antagonist muscles: biceps and triceps. To cause extension, the triceps muscle contracted as shown by the EMG burst, then the biceps muscle followed suit. At this point, the forearm slowed down. Subsequent on/off activity served to stabilize the arm position to a final stop. In a second condition, the structure was locked such that no movement was possible. Still, EMG activities appeared to be synch in time. Why is it so? It seems the control center (brain) was able to produce such stereotyped actions without the need to wait for the sensory feedback. This is a salient example of a motor program, that is, actions are pre-programmed in the brain. The motor program originated in the brain becomes the basis of the 'centralist' group, e.g. Lashley himself. This idea also suggests that the brain does not have to solve Bernstein's degree of freedom problem one by one, but rather the specific action born out of multiple joints or muscles.

Challenges to the concept of a motor program include storage space and producing novel movements never learned before. In terms of speech, e.g., if there are 100 types of sound produced, how many motor programs should a person have? For these reasons, Schmidt (1975) proposed a generalized motor program or GMP, that contains parameters that can be adjusted depending on the situation and purpose. Invariant features of certain movements are thought to be as a result of GMP. Each movement has its own invariant features or signatures. Parameters to be adjusted include: relative timing or duration, the sequence of events, force (impulse) generation, and thus muscles recruitment. Accordingly, the motor program tells the muscles when to turn on, how much force to produce, and when to turn off.

Equilibrium-point Hypothesis
According to this theory, the movement end-points are programmed by the brain and biomechanical properties of the muscle determine the trajectory. In other words, to produce a movement, the brain has to only specify where, not how/when. The model sees our musculoskeletal system as a mass-spring mechanism with stiffness, a force-length relationship. Inherently, muscle fibers are behaving like a spring where certain tension (unit: Newton) is associated with a certain muscle length or elbow angle (cm or degree). The length and tension become two invariant characteristics of the model. The best example comes from using biceps-triceps of the upper arm, an example of antagonistic muscles. Extension occurs because there is a "force" acting to stretch to the biceps. This force increases tension in the biceps but reduces tension in the triceps. Upon sudden "removal" of the force, the musculoskeletal structure goes back to an equilibrium point.
How does this theory explain movements? The model explains that the limb moves to a position defined by an equilibrium point between forces (or torques actually!) of opposing muscles. Let's use the same biceps-triceps example. Suppose at first, the elbow is at a 110º angle and the equilibrium point is defined as two length-tension curves, each for flexor and extensor muscle. The X-axis is the muscle length that defines an elbow angle, the Y-axis is the tension. Threshold length λ is defined as muscle length in "subthreshold state", in which the muscle begins to contract. When the flexor muscle is activated (biceps contract!), its length-tension curve shifts from line 1 to 2. This shift causes the equilibrium point to move towards flexion, i.e. from 110º to 80º. Moreover, the threshold length also shifts from λ1 to λ2. There are two versions of the model. The alpha model (Polit & Bizzi, 1978) says that this process involves no sensory feedback. The lambda model (Feldman, 1966, 1986) says that there is the involvement of muscle spindles to ensure accurate stiffness.

Close-loop Processes
Open loop (left) and close-loop (right) processes
Contrary to the open-loop motor control, the closed-loop model says that sensory feedback plays a crucial role in movement execution. According to this theory, an executive controller continuously monitors the difference between the actual movement produced by the effector or limb, and the intended movement goal (reference). If there is a discrepancy (error), the executive controller sends the command to the limb to correct for this error. This process is generally slower as it requires information processing, e.g. attention control. This is even so as the reference point may change from time to time. In addition, we now know that sensory afferents contain noise and require time (~200 msec) to travel to the cerebral cortex.

One source of sensory feedback comes from the somatosensory and vestibular systems. Although our limbs contain somatosensory feedback in the form of proprioception or kinaesthesia (spindles, joints, tendon organ), the visual system is arguably the most predominant source of perception. This has been shown in the classic literature (e.g. by Gibson, 1943; Adams, 1975; and Jordan, 1972 ). The notion "perception drives action" is the key to Gibson's theory. While the information carried by the proprioception is limited to own-body, the visual system tells us information about both own-body and external world or environment. Such a feedback source is called exteroceptor. For example: while performing an action, we understand that the body moves and at the same time the surrounding moves. Specifically, the apparent motion of objects in the visual scene caused by the relative motion between an observer and a scene is known as optical flow. We now know that the brain contains two visual streams: ventral visual stream (for object recognition or visual perception), and dorsal visual stream (for movement or motion, vision for action!).

The greatest strength of the closed-loop model is the ability to explain movements that are slow, e.g. tracking and tracing. Woodworth's throwing task shows how vision is useful especially when movement duration is > 200 msec. Further, visual feedback is useful during anticipatory activity if the stimulus is available long enough. Close-loop control is important for muscle stiffness by making use of information from the muscle spindle (Houk, 1976). Daily activities that require corrective actions are also considered, e.g. rotating a ball using the index finger. Unfortunately, many other movements that we make are much faster, e.g. aiming to grab a falling item, throwing or kicking a ball, etc. Such quick movements are termed ballistic and require rapid muscle contractions producing high movement velocity. In such a scenario, it is impossible to continuously process sensory feedback (Schmidt, 1972).

One evidence of the feedback influence to rapid movement comes from studies involving long-loop or transcortical reflex. Studies in primates and humans suggest the existence of a long-loop reflex following sudden perturbation of arm/hand position. The existence of two distinct EMG or electromyograph spikes, for example, indicates that there is another signal, M2 (~50-70 msec latency) to the motor neurons following the first, involuntary, short-latency monosynaptic reflex M1 (~30-50 msec latency). Once a source of controversy, this M2 signal is thought to be supraspinal, requires no conscious attention, and can be influenced by prior instructions. The next EMG burst following M2 is the voluntary movement itself and it is regulated by the cerebral cortex (> 150 msec latency).

Ballistic Movement: Reaching 
In principle, a motor task can be divided according to the start and endpoint into discrete, continuous, and serial tasks or skills. A discrete task is a motor task where there are a clear start and endpoint. A continuous task, on the other hand, is a motor task that is continuous and repeated (cyclical). A serial task is a special form of a discrete task in which movements are performed according to a certain sequence. In terms of speed, the motor task can be divided into slow and rapid or ballistic movements. We'll talk briefly about reaching movement here.

Woodworth first discussed this topic in detail a century ago. Ballistic reaching is prevalent in everyday lives, e.g. from movements performed during boxing, playing tennis, and daily voluntary movements such as reaching and grasping (prehension). Such movements are characterized by fast muscle contractions that yield very high speed/acceleration. Scientists thought this mechanism is too fast to be managed by a closed-loop controller in the brain.

While reaching primarily involves the muscular control that rotates shoulder and elbow joints, grasping involves more complex coordination among the five fingers. In most cases, both movements typically occur almost within one intended action in daily lives. The opening of the hand to grasp an object often happens well before the whole arm reaches the object. Does this mean reach-and-grasp is managed by the same motor GMP? According to Jeannerod (1984), reach-and-grasp consists of two behavior phases that utilize two independent neural channels working in parallel, both require a visual system to make visuomotor coordination possible. The two paths are:
1)  A fast initial, transport phase that brings the hand closer to the object,
     - Requires a channel that processes the object's extrinsic properties.
     - E.g. object location w.r.t the body, orientation, direction; viewer-centered coordinates.
2)  A slow, open-the-hand phase during which the hand makes contact with the object.
     - Requires another channel that uses the object's intrinsic properties.
     - These properties include e.g. object size, shape.

In later studies, Arbib et al. proposed that both channels are interdependent through temporal coordination which is necessary to ensure both phases are fulfilled correctly. Opponents to this theory (e.g. by Wing et al.) suggested another concept where spatial, instead of temporal, coordination is required. In another premise, Smeets and Brenner proposed that reach-and-grasp has no distinct temporal components and is nothing but pointing using thumb and finger.

Speed and accuracy 
"Haste makes waste". The fact that faster movements lead to lesser accuracy is known as the principle of a speed-accuracy trade-off. It was Woodworth who first did an experiment studying the relationship between voluntary aiming and its accuracy. He said that manual aiming consists of two phases. The first is an initial open-loop phase (initial adjustment) that propels the hand towards a target. The second phase is a current-control phase using visual feedback to land into the target. This and the following experiment make up a cornerstone of human motor control.

Half a century later, Fitts revisited the concept and showed a formulation of this trade-off. The formula says that average movement time (MT) is linearly related to the index of difficulty. This index is log [2A/W] where A is movement amplitude and W is target width to aim. Both A and W are tightly controlled during the experiment. His law suggests an inverse relationship between 'difficulty' of a movement and the completion time or speed. In the mathematical equation, a and b are Y-intercept and slope respectively. The slope represents a measure of 'controllability', the additional MT caused by an increase in the index of difficulty by a unit. Movements utilizing an upper limb exhibits a steeper slope than fingers. A higher slope is also shown in older adults. The log relationship is presumably due to information load (related to e.g. Hick's study). Subsequent studies have suggested that Fitts' Law can be found in many different situations: cyclical and discrete tasks, in children, adults, and used in neurological patients.

Experiment setup used by Fitts in his early days

Some models attempt to explain Fitts' Law. The most important one is probably the impulse variability model (Schmidt et al., 1979), a model of simple rapid aiming movements in which the variability of the impulse of forces leads directly to variations in the movement end-point of a limb. With this model, the initial phase (initial-impulse) of a ballistic movement is therefore crucial. Related to GMP, this impulse is like a code that tells arm muscles to produce forces within a certain time. The theory maintains that as the distance from a target increases, more force must be exerted, leading to greater variability in movement trajectory, decreasing the chances of hitting the target. To compensate for this, the movement time can be slowed down. A variant of this is by Meyer et al. called the optimized impulse variability model. According to this model, if an error occurs while performing an extremely rapid movement (e.g. the target overshoot or undershoot), a quick corrective movement follows after.

Although Fitts' Law was originally studied in terms of spatial accuracy, later studies discussed the characteristics of a temporal trade-off. This is relevant for example movement reaction time, preparatory time, etc, in baseball (see, e,g. Wing-Kristofferson timing model)


Reference: The writing and diagrams are based on a textbook, "Motor Control & Learning: a Behavioral Emphasis, 5ed" by RA Schmidt & T Lee.