In laboratory settings, motor learning is often studied in the context of motor adaptation paradigms, in which subjects must learn to compensate for a systematic perturbation, either in the through manipulating visual feedback (Krakauer et al. 2000) or a change in the dynamics of the arm manipulandum (Shadmehr and Mussa-Ivaldi 1994). In the first case, visual feedback is altered resulting in a change in movement kinematics. In the latter, change in dynamics is created through the introduction of the opposing force field which is perpendicular to the movement direction. In motor adaptation, what is typically observed is a monotonic improvement in performance that is initially rapid, and then slows to an asymptotic level close to initial baseline levels of performance. The progress of learning is well described by an exponential fit, implying that the amount of improvement on each trial is proportional to the error (Thoroughman and Shadmehr 2000; Donchin et al. 2003). According to these adaptation paradigms, motor learning is predominantly mediated by a specific mechanism that is based on changing an internal forward model.
In real life, we may see the motor adaptation in the case, for example, when one learns to adapt when holding a heavy tennis racket or counteracting the fluid dynamic under-water. However, not all motor learning falls under the definition of motor adaptation. For example, when we learn to synthesize entirely novel movements in the absence of perturbation, adaptation fails to explain these learning processes. Performance improvements take shape from being incompetent as a naive learner to full proficiency. However, such improvements are far slower than in adaptation paradigms: while tens of trials are usually enough to reach an asymptotic level after perturbation, performance in these more complex tasks continues to improve over hundreds of trials or even across a few days.
Haith & Krakauer (2013) define this long-term reduction in movement variability as skill learning and argue that such learning is associated with incrementally improving the quality of one's movements with practice. Skill learning has been studied in the laboratory setting using a few tasks, e.g. maneuvering a cursor along a constrained path (Shmuelof et al. 2012) or through a series of via points (Reis et al. 2009). overall variability in task performance reduces substantially over days of practice, even though subjects immediately exhibit near-perfect performance at slow speeds.
Learning a "Skill"
Skill is another dimension of motor learning. Krakauer et al. defined skill as a shift of speed-accuracy trade-off function when there is no systematic perturbation to movement trajectory. This can be measured before and after a series of training protocols, e.g through tracing an arc or reaching a point in space. Skill learning involves a slower process, or improvement and it is distinct from adaptation. Adaptation to error is not a skill because the performance reaches plateau once the trajectory returns to the baseline trajectory, i.e. there is no further improvement possible. Krakauer mentioned that performance improvement can occur through:
(a) Better state estimation (improved forward models, or processing of sensory feedback);
(b) Better motor execution (improved signal-to-noise ratio in motor output).
Which one is the limiting factor is unknown.
In light of skill learning, an obvious question is this: how can we model skill learning? As mentioned above, error-based motor adaptation has some limitations in real-life cases. Some studies have posited that adaptation also cannot fully explain motor learning processes such as saving, and learning in the absence of error signal. In Huang et al. (2011), subjects experienced visuomotor perturbation. Following adaptation, they went through a sufficient washout block to reset the adaptation. After imposing the perturbation the second time, there is no indication of faster relearning (saving) than the initial learning of naive subjects. However, with the introduction of reward feedback during a task with the same paradigm in another group of subjects, saving is observed following a sufficient washout block.
Model-based and Model-free?
Much of the work on error-based learning uses an internal model to explain motor learning. This is also called model-based learning because it requires building, updating, or adapting an internal model through performing the task. It turns out that another process is possible to support motor learning, called model-free learning. In this model-free type, the goal is to learn directly through a process of trial and error, to explore the space of potential actions in each state, and keep track of which states and actions lead to successful outcomes (rewarded). Indeed, reward-based tasks (reinforcement learning) is an example of model-free learning. Huang et al (2011) and Krakauer (2013) posited both model-free and model-based adaptation are working hand-in-hand to cause saving or faster relearning.
Building an internal model as learning progresses is a salient feature of the model-based approach. Once found, learning can be generalized to multiple but similar task structures. This, however, bears heavy computational costs. In contrast, we usually talk about control policy in reinforcement learning, or model-free learning. A control policy here means the selection of a single action per trial, or describes an ongoing stream of motor commands in continuous time according to the instantaneous state. No intermediate model representation and no calculation required to transform a forward model into motor commands in the control policy. Model-free learning tends to deliver superior performance on a particular task in the long run because they do not rely so heavily on noisy computations each and every time a movement must be made. But the disadvantage is that the learning scope is rather restricted to the task performed during training. Even if the reward structure of the task changes in a known way, one must start from scratch (or at least from some previous but incorrect control policy). This is in sharp contrast to the flexibility offered by model-based learning.
Optimal Feedback Control
While in daily lives, the internal models require constant adjustment especially in the context of learning or adaptation. Todorov and Jordan proposed that the key to such adjustment is the presence of feedback. There are primarily two types: own sensory feedback (proprioception, vision) and the predicted sensory consequences from the internal model (efference copy).
The computation framework that links internal model, feedback, motor cost and reward, and optimization is called the optimal feedback control (OFC) (Todorov and Jordan, 2002) championed by scientists such as Todorov, Kording, etc. It has been regarded as a comprehensive theory of motor coordination in redundant systems, such as our joints that have multiple degree-of-freedom. The theory says that, for a given motor task there is this optimal control policy to update the motor plan using a cost function of effort and accuracy. An OFC uses an optimal estimate of the state of the system, generated through sensory feedback and efferent copy, and uses this feedback to adjust its output towards a specific goal.
Lastly there are several theories suggesting the neural substrates underpinning motor control and learning according to this framework:
References
[1] Haith and Krakauer (2013). "Model-based and Model-free Mechanisms in Human Motor Learning". Adv Exp Med Biol. 782:1–21..
[2] Krakauer and Mazzoni (2011). "Human sensorimotor learning: adaptation, skill, and beyond." Curr Opin Neurobiol, 21(4): 636-644.
In OFC, a cost function of effort and accuracy terms can be optimized, assuming that unbiased estimates of some key parameters (e.g. for forward dynamic model) are available a priori, to derive a feedback control policy for a given task goal. It is thought that during motor learning, a person first learns to adapt to the task dynamic, not the task-irrelevant goal. This step is then followed by the optimization of a cost function to do the task. In other words, there is no clear one-to-one mapping between tasks and actions in the brain. Instead, the naive learner will aim to "converge" into the best solution.
How does OFC relate to skill learning? What is supposedly occurring in the training period is a set of process that leads to better performance: convergence on the optimal policy, or improved execution of the control policy itself, perhaps through an increased signal-to-noise ratio via expanded neural representations. Either of these possibilities could be the explanation for shifts in the speed-accuracy tradeoff, and reductions in variability described in motor skill learning studies. Ultimately, the authors suggest that the process of converging can be both model-based and model-free.
How does OFC relate to skill learning? What is supposedly occurring in the training period is a set of process that leads to better performance: convergence on the optimal policy, or improved execution of the control policy itself, perhaps through an increased signal-to-noise ratio via expanded neural representations. Either of these possibilities could be the explanation for shifts in the speed-accuracy tradeoff, and reductions in variability described in motor skill learning studies. Ultimately, the authors suggest that the process of converging can be both model-based and model-free.
Lastly there are several theories suggesting the neural substrates underpinning motor control and learning according to this framework:
- Basal ganglia (striatum) helps to monitor the reward and cost of the motor commands generated (reward-based learning).
- Cerebellum helps to predicting sensory consequences and monitor mismatch, otherwise known as error monitoring.
- Parietal cortex combines the expected sensory consequences and the actual sensory feedback, a process analogous to state estimation, a place for multisensory integration and the body awareness.
- Premotor and primary motor cortex assign the feedback gain for the visual and proprioceptive states, transforming the belief about the states into the actual motor commands downstream.
References
[1] Haith and Krakauer (2013). "Model-based and Model-free Mechanisms in Human Motor Learning". Adv Exp Med Biol. 782:1–21..
[2] Krakauer and Mazzoni (2011). "Human sensorimotor learning: adaptation, skill, and beyond." Curr Opin Neurobiol, 21(4): 636-644.
[3] Shadmehr & Krakauer (2008). "A computational neuroanatomy for motor control". Experimental Brain Res, 185: 358-381.

No comments:
Post a Comment