Saturday, November 20, 2021

What is Hierarchical Linear Modeling?

This is going to be a very brief conceptual discussion on hierarchical linear modelling (HLM), also known as multi-level modelling in social sciences. People sometimes call this mixed-effects modelling or mixed modelling, but be careful that not all mixed-effects models have a hierarchy. 

Note: Mixed-effects models by definition are models which have both random and fixed effects (see below).

(1)  When do we use it? 
In social sciences, the data types are usually hierarchical in nature, i.e. they have a certain hierarchy that is based on grouping or nesting. We use HLM precisely in cases where observations are nested or clustered in some ways at a different level. It has many applications in cross-sectional and longitudinal research (repeated-measures design). Research studies create data hierarchies through the way the data are sampled. For example:
  • Students (S) are nested within a particular Class, which again is nested in a School.
  • Time points are nested within participants in a longitudinal study.
Fig-1: Pictorial representation of (a) two-level hierarchy and (b) three-level hierarchy (Source: SAGE Ency. 2018)
Suppose we wanted to examine A-level exam scores this year to predict the university entrance test in the following semester. So entrance exam is the outcome variable [DV] and the A-level score is the predictor [IV]. One way is to randomly sample students from the whole population such that each individual is guaranteed to be independent. However, what if the students recruited only belong to a particular high school? In HLM, the students are treated as the first level or the base, and the school as the second level (2-level hierarchy). In other words, we performed sampling primarily from four different local schools only. To advance further, we have to introduce 'Class' in between both levels (now, 3-level hierarchy), since each school surely has different classes. 

In HLM, independence can no longer be freely assumed. Why? E.g. think of how students of Teacher John of School ABC have a different profile, intelligence, general ability than those attending the class of Teacher Mary from School XYZ. In such a nested condition, assumptions fulfilled for the OLS will be violated: observations are no longer independent of each other, and residual errors are correlated within clusters. The error variance will be different for different clusters. As a result, the standard errors of the regression coefficients will generally be underestimated and the significance level will be wrong, leading to misinterpretation. One way is to bring the students' data of Level 1 up to the next level through some averaging, but this is a poor practice. HM treats levels as something to take into account. It is basically an extension of the ordinary least square (OLS) of the linear regression.   

(2)  How to construct the model? 
For a given A-level exam score xi in a simple linear regression, we estimate the intercept b0 and slope b1. The term ei is the residual. The equation states that for a unit increase in the A-level exam score, the score on the entrance test yi increases by the value of b1. But wait! The outcome dataset (entrance test scores yi) is clustered by the school. If this is not taken into account, then the data from all students is treated as unique observations.  
Suppose we adopt a 2-level HLM comprising Students and Schools, i.e. the students are grouped or nested within a school. We introduce a group-level variation by changing the suffices, such that yij is the score on the entrance test for student i in school j, and uj is the group-level residual for that j-th school. This uj is a new random variable, referred to as the Level-2 or group residuals, and is assumed to follow a normal distribution, ~ N(0, σ2). The new HLM equation has the following terms:
  • We see b0 as the grand mean of the university entrance test y.  
  • Level 1 contains the basic form of a linear regression, yij = b0j + b1xij + eij  ; in turns, Level 2 consists of b0j = b0 + uj.
  • The mean of y for the school group j is b0 + uj, where uj is the "school effect", i.e. the difference between the mean of the school group j and the grand mean. With this, each school can have different mean scores on the entrance test. 
  • The individual-level residual eij is the difference between the value of y for a student i and the individuals group mean b0 + uj. This reflects differences in students’ individual test scores from their respective school (or group) mean.
What can this improved model tell us? Some schools will have means that are higher than the grand mean, suggesting that their students perform better on the A-level exam on average, and some schools will have lower cluster mean values. 

(3)  Defining fixed and random terms 
The core of hierarchical models is the assignment of fixed effects and random effects, two terms that already appeared in another blog post on group-level fMRI analysis. You have to define each term of a mixed-effects model to be either a fixed or random effect.
  • A fixed effect is a parameter that is fixed across all groups and does not vary in the model. It is the term of interest in experimental manipulation. Estimating a fixed-effect of a term is like estimating a regression slope b. Differences in the slope If all terms in the model above are fixed, it behaves as the usual linear regression.
  • In contrast, a random effect allows each group to have a different estimate. In HLM, random effects represent a higher level variable under which data points are grouped. This implies that random effects must be categorical (but cannot be continuous!). For example, the residual errors (uj and eij) are considered random effects. Estimating a random effect is like looking for an effect in our data to come from a large group of normally distributed datasets. 
The most basic model has only an intercept without any predictors xij, which is also called the intercept-only model. Another name for it is the Unconditional Model or Null Model. Usually, the model has only participants eij as a random error.

The next basic model is the one shown in the equation above. It allows for different groups or schools j to have a varying intercept or mean b0j = b0 + uj. Thus, the model is also known as the random intercept model. 

(4)  The more complete model 
A random intercept model shown above assumes that the relationship between the university test and the A-level exam score (the predictor) is the same for each group. This assumption can be relaxed by allowing for different slopes for the predictor in each group, making it a random slope model as shown in the right panel of Fig-2 below. Each slope will be estimated separately for each group. In this new equation, there is a new term u1j xij , where:
  • The intercept for school group j is now b0 + u0j. 
  • The slope for school j is now b1 + uij , where b1 is the average slope across groups.
  • Subscript “0” differentiates the random effect for the intercept u0j from the random effect for the slope u1j. Both random effects are assumed to be normally distributed. 
  Fig-2: Hierarchical models [Source: Univ Bristol]

The variance of the intercept and slope are assumed to be correlated; the covariance between the intercept and slope is estimated as part of the random slopes model. The random slopes model is also commonly known as the random coefficient model or a growth curve model when using repeated measures or longitudinal analysis. 

(5)  Final notes 
How is the implementation in practice? Apart from defining fixed/random terms, we have to know which one nested under which variable.
  • When we deal with a dataset we have to begin with the simplest model, which is the intercept-only model.
  • Then we add additional terms, the simplest being the random intercept model, i.e. each group has its own group mean. 
  • If we have more than one predictor or IV, define the fixed/random effect. A mixture of fixed and random effect slopes is possible. 
  • The most complete model is the random slope model, where intercepts and slopes are treated as random effects. Often in a longitudinal dataset, we can add the quadratic term "Time" (Level-1).
So how do we know which model is best? We can use the usual performance metrics such as Bayesian Information Criterion (BIC) or Akaike Information Criterion (AIC). The two most common libraries for HLM analysis in R are lme4 and nlme. In the commands, you have to indicate which ones are the fixed and random effects, and which variable is nested within what. The outputs of the HLM in R is quite similar to what we expect from OLS in linear regression. 

The random slope model often mimics reality, but this requires a large sample size. Parameter estimation in HLM is achieved using either the Maximum Likelihood (ML) or Restricted Maximum Likelihood (ReML). ReML works better if the sample size is small.


Some good references:

Sunday, November 7, 2021

Introduction to Artificial Neural Network (1)

Everyone is talking about AI and machine learning recently. But the concepts of AI goes back to more than 40 years ago when scientists took steps to model the brain and cognition. In its original concept, there lies a perceptron, an artificial neuron that can form a neural network to learn and make decision (after F. Rosenblatt). Just as a neuron has an all-or-none firing characteristic, a neural network is able to provide a binary response or yes/no. While in the past it required advanced C++ and expensive hardware to run modeling, today knowing Python and having a high-speed computer are sufficient.

1. Principles of a perceptron
A perceptron takes several binary inputs and produces a single binary output. How does it compute the output? Each input is associated with a certain weight wi and the total value of wi xi will be compared against a certain threshold to decide the output. Look at the figure below.
We can take a simplified real life example for how a perceptron works. Suppose there is a jazz festival and you need to decide whether you want to go ('1') or not ('0'). There seems to be 3 important factors that influence your decision, each has its own influence value:
      - Is the weather good (yes/no)? Influence scale = 10.
      - Is your girlfriend going with you (yes/no)? Influence scale = 3.
      - Is the place near a metro station (yes/no)? Influence scale = 4.
Note that the influence scales here represent the weights you personally assign to these factors. For example, you really hate being caught in the rain so much; or you don't really mind if you go alone without your girlfriend and make new friends. You should also infer a threshold to make your decision. That's it, a perceptron has just been used to model how you make a decision. What happens when we combine multiple perceptrons? This structure is called a neural network, as shown in the figure below. In practice, such network is used to take into account more factors and make a better decision.

Let us describe the perceptron more formally. First, the summation of weights and inputs can be rewritten as a dot product of w.x = Σ wi xi . Second, the threshold is in fact a bias, b = −threshold. This bias can be thought as a measure of how easy it is to get the perceptron to give an output '1'. Or to put it in more biological terms, it is a measure of how easy it is to get the perceptron to fire. The formula to get the output thus becomes, y = w . x + b. Thirdly, the binary output of '0' or '1' beyond a certain threshold mimics a type of function called a step function.


2. Simple neural network
A neural network is basically a multi-layer perceptron or MLP as shown in the figure above. The leftmost layer is called the input layer and the neurons inside it are called input neurons. The rightmost layer is called an output layer with just a single output neuron. In between, there exists 2 layers of what is called hidden layers, which take inputs from the input layer and project outputs to the output layer. A neural network in such structure is also popularly known as a feedforward neural network since the outputs from one layer are projected to the the next layer to the right and so on, with no feedback allowed. There are a few other neural network architecture, which will be said at a later time.

The power of a neural network lies in the capability to be trained. Neural networks can be 'trained' to behave in a certain way, or to produce a certain output given some patterns of inputs. Engineers can call this an iterative fine-tuning process. We use neural networks to help us decide something. For example, you want to decide whether a blurred image is a picture of a handwritten digit (number "0" to "9"). Training the neural network here means we let the network to readjust the weights and biases within the network according to the given inputs, in this case, different images.

How can we readjust the weights and biases? We let a small change in a weight or bias to cause only a small change in output Δy (like fine-tuning). In this way, we would get our network to behave more in the manner we want. For example, suppose the network was mistakenly classifying an image as an "8" when it should be a "9". We could figure out how to make a small change in the weights and biases so the network gets a little closer to classifying the image as a "9". This process is then repeated over and over again to make the output more and more accurate to decide "9". The network is said to be learning.

But there is one problem. With the current setup, a small change in an input xi will yield quite a big jump in the binary output, it can completely be flipped (yes/no, '1' or '0'). This is primarily due to the all-or-none nature of the perceptron output. Scientists were not satisfied with such characteristic, so they defined a new type of perceptron where the output follows a sigmoid rather than a step function characteristic. Thus, a neuron with a sigmoid function don't just produce output '1' or '0'. It turns out that with this characteristic, Δy is a linear function of the changes Δw and Δb. This linearity makes it easy to choose small changes in the weights and biases to achieve any desired small change in the output. Formally, output characteristic of an artificial neuron or perceptron can be defined by the activation function, where a sigmoid is one of the examples.
As mentioned, training a neural network means allowing the weights w and biases b to readjust themselves such that a desired output is obtained. We can track how well the training progresses through a metric called a cost function, sometimes also a loss or objective function as a function of weights and biases. Remember that the goal of a neural network is to help us make decision or prediction. Traditionally, the cost function defines the difference between the predicted and the actual output of the network. Our end goal is to find w and b such that this difference is minimized.
3. Minimizing cost function
Suppose we have our multi-layer perceptron and this is so-called our model. Mathematically, the method to "train" the model of a neural network is called backpropagation algorithm. This term is coned after Rumelhart et al who proposed an efficient numerical solution of the training problem. Training or learning in terms of backpropagation here means to minimize discrepancy or error (or something bad) by adjusting weights and biases. 

To portray the discrepancy or error, we need a cost function, which defines how 'good' our model is at the moment as learning progresses. Now we wish to obtain a set of parameters (weights and biases) such that the discrepancy between the predicted and actual output of the model (neural network) as defined by the cost function is minimized. To minimize this so-called cost function, we can use an optimization algorithm called gradient descent, by iteratively moving in the direction of steepest descent as defined by the negative of the gradient (or slope).

Suppose we have a cost function F that depends on imaginary parameter x (x1 and x2). Let us imagine a valley defined in the dimension of x1 and x2, such as the one shown below, and we want to roll a ball down this valley. Making the ball roll down at different dimension of x1 and x2 is akin to saying that the change in F is negative, ΔF < 0. We can build a relationship such that: ΔF ≈ ∇F ⋅ Δx ; where ∇F is a the gradient vector of F, which carries partial differentiation operators. At the moment, it is enough to think a gradient vector simply as something that relates changes in x to changes in F, just as we would expect something called a gradient to do. Suppose we make the change in x to be Δx = −η∇F, where η > 0. Then mathematically, ΔF ≈ −η ∇F⋅∇F = −η∥∇F∥2 , which means that ΔF is guaranteed to be negative and F will forever decrease not increase. The ball is for sure rolling down the valley!


The whole idea of iteration is as follows. From an arbitrary ball position of in dimension x, we first compute the change in Δx. so that to find a new position of x → x' = x − η∇F. Note that the arrow denotes an update rule, the variable takes up a new value. This update rule can be thought as defining the gradient descent algorithm. It gives us a way of repeatedly changing the ball position in order to find a minimum value of the function. If we keep doing this over and over again, we will keep decreasing F, until theoretically we reach a global minimum. Once this global minimum has been found, it is said that the training algorithm has converged. Note that η is called learning rate where it is usually kept small, to control learning and to prevent the update behavior to be chaotic.

To summarize, the way the gradient descent algorithm works is to repeatedly compute the gradient vector ∇F, and then to move in the opposite direction step by step, so as to "fall down" the valley. Updates of weight and bias parameters occur in an efficient way until the error is minimized.

Some notes:
1) In more complex networks for a multiclass prediction, softmax is used as the activation function instead of a sigmoid.
2) The following website is informative to understand backpropagation.

Sunday, February 28, 2021

10 Most Common Cognitive Bias

Hello! Learning about human cognition is indeed fascinating. While humans like to believe that they are rational and logical, the fact is that on a daily basis people are continuously under the influence of cognitive biases, mental or psychological phenomena that distort thinking, sway beliefs or decision, and influence their decisions that lead to poor choices. Some of the most cognitive biases are given below. 

  1. Confirmation Bias: the tendency to listen more often to information that confirms our existing beliefs. In other words, people tend to favor information that reinforces the things they already know or believe, refusing the opposing side. Some examples: following social media that you only like, choosing news sources that present stories that support your views, etc. Why is this happening? Perhaps it fuels the ego-centric satisfaction, it also protects self-esteem by making people feel that their beliefs are accurate. 
  2. Hindsight Bias: the tendency to see events, even random ones, as more predictable than they are. Sometimes, it is also commonly referred to as the "I knew it all along" phenomenon. Some examples: you claim to know who will win the Election or World Cup, or what the exam answers are, after these events are just over. This bias occurs can be due to the tendency to view events as inevitable, and our tendency to believe we could have foreseen certain events. The effect of this bias is that it causes us to overestimate our ability to predict events. 
  3. Anchoring Bias: the tendency to be overly influenced by the first piece of information that we hear or see, taking it as the final truth. Some examples: the first price produced during a price negotiation, or doctors can be susceptible to first impressions when diagnosing patients. Why does this bias occur? Perhaps the source of the information plays an influential role. Other factors such as priming and mood also appear to have an influence.
  4. Misinformation Effect: the tendency for our memories to be heavily influenced by things that happened after the actual event itself. For example: a person who witnesses a car accident or crime could believe that their recollection is crystal clear. But studies have shown that simply asking questions emphasizing some keywords, or letting them watching a misguided video about an event can change someone's memories of what actually happened. This bias may occur because of interference, new information may get mixed with older memories.
  5. Actor-observer Bias: The actor-observer bias is the tendency to attribute our actions to external influences and other people's actions to internal ones. When it comes to our own actions, we are often far too likely to attribute things to external influences. For example: You argued that the reason for missing an important meeting is because of a jet lag, or you failed an exam because the teacher posed too many trick questions. Conversely, however, when a colleague screwed up an important presentation we say it's because he’s lazy and incompetent (not because he also had a jet lag); when a friend failed a test, it happens because they lack diligence and intelligence (and not because they took the same test as you with all those trick questions). Our reasoning and perspective play a key role in forming this bias. 
  6. False-Consensus Effect: the tendency people have to overestimate how much other people agree with their own beliefs, behaviors, attitudes, and values. For example: one may think that other people (friends and family members) share your opinion on controversial topics, or believing that the majority of people share your preferences. Why is this? We tend to generalize by thinking that the amount of times spent together make us easy to think the views of others conform with our own. This bias cause us to overvalue own opinions. It also means that we sometimes don't consider how other people might feel when making choices. 
  7. Halo Effect: the tendency for an initial impression of a person to influence what we think of that person overall. It is also known as the "physical attractiveness stereotype". Some example: we may think that a good-looking and confident woman is also smarter, kinder, and funnier than less attractive people, or an attractive job applicant is seen as likable and more likely to be viewed as competent and qualified for the job. It is thought that this bias is due to our desire to be correct. If our initial impression of someone was positive, we want to look for proof that our assessment was accurate. It also helps people avoid experiencing cognitive dissonance, which involves holding contradictory beliefs. This bias is being used to advertise commercial products with attractive supermodels.
  8. Self-Serving Bias: the tendency for people to give themselves credit for successes they make, but lay the blame for failures on outside causes. For example: when you do well on a project, you probably assume that it’s because you worked hard. But when things turn out badly, you are more likely to blame it on circumstances or bad luck. This bias is closely linked to self-esteem, age, and culture. 
  9. Availability Heuristic: the tendency to estimate the probability of something happening based on how many examples readily come to mind. Some examples: after watching the news of car thefts in the neighborhood, you then start to believe that such crimes are more common than they are; or that plane crashes are more common than they really are because you can easily think of several examples. This bias is basically a mental shortcut designed to save us time when we are trying to determine risk, resulting in overthinking and worrying. 
  10. Optimism Bias: the tendency to overestimate the likelihood that good things will happen to us while underestimating the probability that negative events will impact our lives. Some examples: we may assume we will get salary increment or bonuses, and that we won't be sacked in our current job; or we tend not to wear seatbelts because we feel the accident won't happen any time soon. Why is this bias? Because we may have over-confidence, or think that bad things often occur to other people and it seems more likely that others will be affected by such bad events.

Sunday, February 7, 2021

I may stop writing posts

Writing a blog post is taking too much time. I recently discovered that using OneNote by Microsoft Office365 seems like a much better idea to compile notes (yes, I'm old-fashioned!). So I will stop spending time writing blog posts and use OneNote instead. Another thing is that I just learn how to write a manuscript with LaTex, a nice alternative to MS Word. I'm not sure when to totally stop this. 

Meanwhile, the world has faced a serious biological threat stemming from a new coronavirus, the COVID-19 pandemic. The disease has caused a worldwide chaos since early 2020, and the situation has not improved as newer variants has emerged and hospital manpower is overstrained. The pandemic has also affected my research work and I do not feel that good about it. 

Let's see............. Stay safe! 

Okie, ciao!

Sunday, January 31, 2021

Notes on Somatosensation

By and large, somatosenses or bodily sensations are mostly concerned with sensations that arise from stimulation of the skin of the body. Primarily referred to as touch sensation, this was later known to encompass more properties, e.g. pressure, vibration, warmth, and cold. When we interact with an object, we are usually aware of its many different attributes such as the shape, texture, plasticity, hardness, and temperature. One distinctive feature of touch is that it arises from specialized receptors distributed throughout the whole skin known as the mechanoreceptors. There are also specialized receptors within the muscles and joints that provide sensory feedback known as proprioception and kinesthesia. These sensors convey muscle length, change in length, and force (tension), allowing us to be aware of the limb location and movement without vision in space. Receptors encapsulating nerve endings are stimulus-specific and have their unique physiological properties. 

Historically, researchers have viewed the senses as generally subserving a primarily discriminative role (Mountcastle, 2005). Touch and proprioception are naturally related to motor control. Touch is also related to affective and social neuroscience, that is, how one feels and interacts with each other.

1. Peripheral receptors
Mechanoreceptors beneath the skin surface give rise to the perception of touch, pain, and temperature. Free nerve endings are known to provide information of pain and temperature to the brain. There are several properties worth reporting, i.e. receptor size, myelination, location, conduction speed, and response adaptation. More information can be found in Table 3.1 and 3.2 below and also Fig-1.
  1. Mechanoreceptors with small fields can distinguish two closely spaced objects better (higher spatial resolution) than receptors that have a large field size (low spatial resolution).
  2. Some receptors are located superficially near the epidermis and are called Type-I receptors. Conversely, receptors located in deep skin are of Type-II receptors. 
  3. Fast conduction is important to provide afferent signals that build the withdrawal response or reflex arc, e.g. to detect adverse event. 
  4. Response adaptation, in this case, refers to how the fibers respond to continuous touch stimulation: slow-adapting fibers (SA) and fast-adapting (FA) fibers. FA-fibers are only sensitive to the onset and offset of the stimulus. As such, their response property is called phasic, dependent on on/off time. SA-fibers, however, respond continuously but gradually to a stimulus, that is, with a tonic response. See below!
Fig-1: Properties of mechanoreceptors and free nerve endings in the peripheral system. Take note the bottom left figure. Rapidly adapting receptors exhibit phasic response, while slow-adapting receptors show a more tonic response.

All somatosensory information is sent to the central nervous system through spinal nerves. In the case of the head, neck, and face region, information is sent through a set of cranial nerves out of the brain stem. The entire body surface can therefore be divided into discrete areas that are represented by a dermatome, and the map representing the skin surface devoted to all spinal nerves is called the dermatomal map. Although it appears that the boundaries are exact in this representation, there is actually some overlap in innervation between adjacent spinal nerves. Dermatomal maps are valuable as a clinical tool in the event of injury or infection to a particular dorsal root.
Fig-2: Graphical illustration of a dermatomal map in relation to the spinal cord.

As mentioned on top, an incoming stimulus makes contact with receptor organ on the skin. An action potential is generated in the immediate vicinity of the receptor organ itself, unlike conventional multipolar neurons where the depolarizing signal must first reach the cell body to produce an action potential. With the mechanoreceptor, action potentials then flow along the peripheral to reach the spinal cord. The cell body of a mechanoreceptor neuron is located in the dorsal root ganglion next to the spinal cord. There, the neuron carries synapse with interneurons and cortical neurons, some of which send information to the cerebral cortex. 

The topic on ascending pathways and somatosensory cortices have been discussed in the earlier blog post. To add on, the primary somatosensory cortex or S1 in the postcentral gyrus is arranged in a colunnar organization (Mountcastle, 1997). Somatosensory inputs from different parts of the body are arranged as columns of neurons that run from the surface to the white matter and encompass all six layers of the cortex. An example is shown here for a part of Area 3b that represents the digits. The expanded view at the bottom shows that each digit is represented by a different column, which in turn is subdivided into inputs from FA- and SA-type afferents.

2. Perceptual aspects of tactile sensation
Among the different somatosenses, the perception of touch was the first to be studied and remains the most widely studied topic. The perception of touch became the main interest of Fechner and Weber in their pioneering work in psychophysics, and was later developed further by Weinstein in 1960s. Tactile sensation can be examined based on the intensity, spatial, and temporal aspects. 
  • Tactile intensity. To measure the absolute detection threshold of touch on the various parts of the body, a device called aesthesiometer was invented using nylon filaments. In practice, the filaments will produce skin deformity that in turn gives rise to a perception of touch. See the figure below. The facial regions have the lowest absolute threshold (most sensitive), whereas the foot area has the highest (least sensitive). Another measure of psychophysical properties of touch is the difference threshold but it warrants careful and systematic experiments as it can be influenced by two main factors: intensity of the stimulus and the location of the contact.
  • Place of contact. According to Weber, the two-point limen is the smallest separation of two points applied simultaneously to the skin that can still be discriminated, i.e. they evoke the sensation of two separate points. Traditionally, the two-point limen was seen to improve steadily from the shoulder to the fingertip, resulting in a near twice reduction of the threshold. What is the physiological basis for this? The size and density of the mechanoreceptors. The smaller the size of the receptive field, the greater the ability to discriminate two different contact points. Such property is apparent for FA-I and SA-I type afferents that are more superficial on the skin. Differences in packing density also play an important role in determining tactile acuity (Vallbo & Johansson). The superiority in acuity (discriminative ability) of the fingertips allow visually-impaired people to read Braille alphabets.
Fig-3: Although the facial regions are most sensitive to touch (*), it is the fingers that show the highest spatial discrimination ability. Similarly, much of the upper torso is quite sensitive to touch but shows poor spatial resolution. The feet display just the opposite pattern: relatively good acuity or discrimination ability (#) but poor sensitivity.
  • Temporal aspect. Much of the research done with the temporal aspects made use of either a prolonged stimulation or a short but repetitive vibrotactile stimulation. Prolonged application of vibrotactile stimulation is shown to cause adaptation by reducing the sensitivity of vibration at the skin stimulation site (Gescheider and Wright, 1969). In contrast, vibratory stimuli produce a different threshold value than static/prolonged stimuli do. Temporal changes in touch can best be detected by fast-adapting mechanoreceptors. The relation between vibration frequency and skin displacement is therefore a U-shaped recruitment curve. Our perception is determined by the property of the mechanoreceptors active: 
    • Very small vibrations between 200–300 Hz are captured by the FA-II receptors (Pacinian) that respond to vibrations > 50 Hz, 
    • Whereas FA-I receptors (Meissner) respond to vibrations between 20–40 Hz. 
    • For stronger vibrations well above these threshold levels, more than one type of mechanoreceptor typically responds to movements of the skin. 
Lastly, vibrotactile stimulation has been widely used to examine sensory processing and memory in primates (e.g. Romo et al.). Touch is ... an intermediary sensory system in that its spatial resolving power is poorer than that of vision but superior to that of audition, and its temporal resolving capacity is better than vision but inferior to audition (Lynette Jones, MIT).

3. Perceptual aspects of limb position
Proprioception and kinesthesia are critically important to the motor system in guiding our movement through the environment. Unlike tactile sensation, proprioceptive stimuli are primarily internal that are generated by the position or movement of a body part. Static forces on the joints, muscles, and tendons, which maintain limb position against the force of gravity, indicate the position of a limb. The movement of a limb is indicated by dynamic changes in the forces applied to muscles, tendons, and joints. More recent studies have looked into how proprioception and kinesthesia are also responsible for sensing effort and force exertion through tension placed on the muscles. 

Receptors associated with the limb position sense generally can be divided into 4 groups depending on the discharge properties, type of stimulus, and the neurons innervating the organ. Interestingly, some mechanoreceptors are found to give rise to the sensation of limb position and movement, e.g. presumably due to skin stretch, vibration. For historical reason, the naming convention is different from the receptors innervating the skin. Compare the table above and below. Both Group I and II afferents are analogous to the Aα and Aβ fibers respectively, that have large diameters and heavily myelinated. Muscle spindles are arranged in parallel with the extrafusal fibers that make up the main body of the muscle. This parallel arrangement is the best for detecting a change in fiber length, and the speed of that change. Consequently, these fibers are among the fastest in terms of signal transmission to the spinal cord.

From a clinical and sports science perspective, it is common to associate proprioception to primarily balance and postural control; and kinesthesia to the sensation of movement. In one of the clinical tests, for example, we look into how the proprioceptive components are working properly when the visual cues are missing and proprioceptive cues are the major sources of information.


The roles of proprioception and kinesthesia in motor control and learning cannot be denied. Every time we consciously move one limb, a command is issued from the motor centers of our brain to the appropriate muscles. If the perceptual centers in our brain could somehow get a copy of that command, then it would have the means to know what movement is taking place. This information would be independent of that being generated by the proprioceptors in the muscles and joints. Such is the principle of corollary discharge. Not only that these receptors sense movement and position, but they also allow us to perceive the sensation of effort and force. A comprehensive review of this system has been written by Proske & Gandevia, 2012. 

The sensitivity of the limb position and movement in space is quite remarkable. Among the joints, the hip joint appears to be the most sensitive to detecting a movement as little as 0.2°. Among the major limb joints, the following order has been reported in terms of decreasing sensitivity: shoulder, knee, ankle, elbow, wrist, and finger base. Unfortunately, the data from kinesthetic experiments is complicated by several factors such as the direction of the joint movement, the degree to which the limb is stretched, and the precise way in which the measurements are performed. From psychophysics, the Weber fraction for weight discrimination is about 0.02. On the other hand, the perceived effort (force generation) can be fitted to a power function with an exponent value of 1.7.

One popular paradigm to test proprioception is the joint position matching test. Experimental factors that affect matching errors (Goble, 2010): (1) Ipsilateral matching, (2) Left arm advantage - right hemisphere damage people are more prone to proprioceptive deficits, (3) Tau effect - longer time to complete reference movement by experimenter leading to targets being perceived as further from the starting point, (4) Age of participants (matching errors increase in older population) (5) Left workspace bias in joint position matching task. To elaborate on the second factor: a study by Naito and colleagues (2005, 2007) used a tendon vibration paradigm in combination with neuroimaging to map regions of the brain responsible for processing input from key proprioceptors—the muscle spindles. Hence, they concluded proprioceptive performance lies more within the right hemisphere.