{"id":739,"date":"2018-01-10T13:54:30","date_gmt":"2018-01-10T03:54:30","guid":{"rendered":"https:\/\/www.cognav.net\/?p=739"},"modified":"2018-01-14T16:21:05","modified_gmt":"2018-01-14T06:21:05","slug":"population-encoding-and-decoding","status":"publish","type":"post","link":"https:\/\/braininspirednavigation.com\/?p=739","title":{"rendered":"\u3010Excerpt Note\u3011Population Encoding and Decoding"},"content":{"rendered":"<p><span style=\"font-size: 12pt;\">The content is from the lecture notes (<a href=\"http:\/\/www.caam.rice.edu\/~caam415\/Ma_lectures\/M1_Slides_Population_coding.pdf\">M1_Slides_Population_Coding<\/a>, <a href=\"http:\/\/www.caam.rice.edu\/~caam415\/Ma_lectures\/M1_Notes_Population_coding.pdf\">M1_Notes_Population_Coding<\/a>) of &#8216;Theoretical Systems Neuroscience&#8217; by <a href=\"http:\/\/www.weijima.com\/\">Professor Wei Ji Ma<\/a> (Baylor College of Medicine) in 2013. He is now leading the <a href=\"http:\/\/www.cns.nyu.edu\/malab\/index.html\">Wei Ji Ma&#8217; lab <\/a> at Center for Neural Science and Department of Psychology at New York University. <\/span><\/p>\n<p><span style=\"font-size: 12pt;\">For further more info on the course website <a href=\"http:\/\/www.caam.rice.edu\/~caam415\/Ma_lectures\/\">CAAM\/NEUR 415: Theoretical Neuroscience <\/a><\/span><span style=\"font-size: 12pt;\">Fall 2013, Biophysical Modelling and Computation from Cell to Network at Rice University.\u2003 <\/span><\/p>\n<p><span style=\"font-size: 12pt;\">Some useful <a href=\"http:\/\/booksite.elsevier.com\/9780123748829\/pictures\/code\/popcodes\/\">MATLAB example of population coding and decoding<\/a> from the Book, <a href=\"http:\/\/booksite.elsevier.com\/9780123748829\/pictures\/code\/popcodes\/\">Gabbiani, Cox: Mathematics for Neuroscientists , 1st Edition<\/a>. <\/span><\/p>\n<p><span style=\"font-size: 12pt;\"><strong>Table of Contents <\/strong><\/span><\/p>\n<p><span style=\"font-size: 12pt;\">1 Population encoding and decoding <\/span><\/p>\n<p><span style=\"font-size: 12pt;\">1.1 Encoding <\/span><\/p>\n<p><span style=\"font-size: 12pt;\">1.2 Decoding <\/span><\/p>\n<p style=\"margin-left: 36pt;\"><span style=\"font-size: 12pt;\">1.2.1 Winner\u2010take\u2010all decoder <\/span><\/p>\n<p style=\"margin-left: 36pt;\"><span style=\"font-size: 12pt;\">1.2.2 Center\u2010of\u2010mass or population vector decoder <\/span><\/p>\n<p style=\"margin-left: 36pt;\"><span style=\"font-size: 12pt;\">1.2.3 Template\u2010matching decoder <\/span><\/p>\n<p style=\"margin-left: 36pt;\"><span style=\"font-size: 12pt;\">1.2.4 Maximum\u2010likelihood decoder <\/span><\/p>\n<p style=\"margin-left: 36pt;\"><span style=\"font-size: 12pt;\">1.2.5 Bayesian decoders <\/span><\/p>\n<p style=\"margin-left: 36pt;\"><span style=\"font-size: 12pt;\">1.2.6 Sampling from the posterior <\/span><\/p>\n<p><span style=\"font-size: 12pt;\">1.3 How good are different decoders? <\/span><\/p>\n<p style=\"margin-left: 36pt;\"><span style=\"font-size: 12pt;\">1.3.1 Cram\u00e9r\u2010Rao bound <\/span><\/p>\n<p style=\"margin-left: 36pt;\"><span style=\"font-size: 12pt;\">1.3.2 Fisher information <\/span><\/p>\n<p style=\"margin-left: 36pt;\"><span style=\"font-size: 12pt;\">1.3.3 Goodness of decoders <\/span><\/p>\n<p style=\"margin-left: 36pt;\"><span style=\"font-size: 12pt;\">1.3.4 Neural implementation <\/span><\/p>\n<p><span style=\"font-size: 12pt;\">1.4 Decoding probability distributions <\/span><\/p>\n<p style=\"margin-left: 36pt;\"><span style=\"font-size: 12pt;\">1.4.1 Other forms of probabilistic coding <\/span><\/p>\n<p style=\"margin-left: 36pt;\"><span style=\"font-size: 12pt;\">1.4.2 Discrete variables <\/span><\/p>\n<p><span style=\"font-size: 12pt;\">A population code is a way of representing information about a stimulus through the simultaneous activity of a large set of neurons sensitive to the feature. This set is called a population. A population code is useful for increasing the animal&#8217;s certainty about a feature, as well as for encoding multiple features at once. Encoding is how a stimulus gives rise to patterns of activity (in a stochastic manner), decoding is the reverse process, by which a neural population is &#8220;read out&#8221;, either by an experimenter or by downstream neurons, to produce an estimate of the stimulus. <\/span><span style=\"font-size: 12pt;\">Population codes are believed to be widespread in the nervous system. <\/span><\/p>\n<p><span style=\"font-size: 12pt;\"><strong>1. Population Encoding and Decoding <\/strong><\/span><\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/www.braininspirednavigation.com\/wp-content\/uploads\/2018\/01\/011018_0237_PopulationE1.png\" alt=\"\" \/><\/p>\n<ul>\n<li>\n<div><span style=\"font-size: 12pt;\">Stimulus: physical feature of the world (often 1D) <\/span><\/div>\n<ul>\n<li><span style=\"font-size: 12pt;\">Orientation of a contour <\/span><\/li>\n<li><span style=\"font-size: 12pt;\">Direction of self-motion <\/span><\/li>\n<li><span style=\"font-size: 12pt;\">Number of students in class <\/span><\/li>\n<li><span style=\"font-size: 12pt;\">Whether object A is bigger than object B <\/span><\/li>\n<li><span style=\"font-size: 12pt;\">\u2026\u2026 <\/span><\/li>\n<\/ul>\n<\/li>\n<li><span style=\"font-size: 12pt;\">Neural Representation: spike activity of neurons in response to stimulus (population code) <\/span><\/li>\n<li><span style=\"font-size: 12pt;\">Stimulus judgment: often motor response <\/span><\/li>\n<\/ul>\n<p><span style=\"font-size: 12pt;\"><strong>1.1 Encoding <\/strong><\/span><\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/www.braininspirednavigation.com\/wp-content\/uploads\/2018\/01\/011018_0237_PopulationE2.png\" alt=\"\" \/><\/p>\n<p><span style=\"font-size: 12pt;\">Let s be the stimulus (or a specific value of the stimulus). We will denote the response of a neuron to the stimulus by r. For a brief stimulus, this is the total number of spikes elicited. For a sustained stimulus, it can be the total number of spikes in a certain time interval. When s is presented many times, different values of r will be recorded. We will denote the mean response by f(s). Unlike r, f(s) is not necessarily an integer. As a function of s, f(s) is called the tuning curve of the neuron. It typically is bell\u2010shaped (for stimuli like orientation) or monotonic. When it is bell\u2010shaped, then the mode of the function is called the preferred stimulus of the neuron. <\/span><span style=\"font-size: 12pt; background-color: white;\">The variability of r around its mean in response to s can often reasonably be described as a Poisson process with mean f(s). That means that it is drawn from the following distribution: <\/span><\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/www.braininspirednavigation.com\/wp-content\/uploads\/2018\/01\/011018_0237_PopulationE3.png\" alt=\"\" \/><\/p>\n<p><span style=\"font-size: 12pt; background-color: white;\">Note that this is a conditional probability distribution: we are not interested in the distribution of responses in general, but only in response to a specific stimulus. The variance of a Poisson-distributed variable is equal to its mean, which is not completely consistent with measurements. In most cortical neurons, the Fano factor, which is the ratio between variance and mean, is found to be more or less constant over a range of mean activities, but with a value anywhere between 0.3 and 1.8. <\/span><\/p>\n<p><span style=\"font-size: 12pt; background-color: white;\">Different neurons will in general have different tuning curves. Suppose we have a population of n neurons. We label the neurons with an index i, which runs from 1 to n. The tuning curve of the i&#8217;th neuron is denoted fi(s). Instead of choosing a Poisson distribution to describe neural variability, one can use a normal distribution: <\/span><\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/www.braininspirednavigation.com\/wp-content\/uploads\/2018\/01\/011018_0237_PopulationE4.png\" alt=\"\" \/><\/p>\n<p><span style=\"font-size: 12pt; background-color: white;\">For large values of the mean, a Poisson distribution is very similar to a normal distribution. A normal distribution can be made into a more realistic description of variability by taking its variance to be proportional or equal to its mean, just as is the case for the Poisson distribution. However, for small values of the mean, any normal distribution runs into problems since it is defined on the entire real line, including negative values, whereas spike counts are always nonnegative. One can cut it off at zero, but then the distribution loses the nice properties of the normal distribution. <\/span><\/p>\n<p><span style=\"font-size: 12pt;\">On a single trial, the response of a single neuron is denoted <em>ri<\/em>, and the population response can be written as a vector <strong>r <\/strong>= (<em>r<\/em>1,\u2026,<em>rN<\/em>). We can characterize the variability of the population response upon repeated presentations of stimulus <em>s <\/em>by a distribution <em>p<\/em>(<strong>r<\/strong>|<em>s<\/em>). This is called the <em>response distribution <\/em>or the <em>response distribution<\/em>, although the word &#8220;noise&#8221; might <\/span><span style=\"font-size: 12pt;\">be misleading. In more generality, a mathematical description of how the observations are <\/span><span style=\"font-size: 12pt;\">generated probabilistically by a source variable is called a <em>generative model<\/em>. <\/span><\/p>\n<p><span style=\"font-size: 12pt;\">The simplest assumption we can make to proceed is that variability is independent between neurons. This means that for a given <em>s<\/em>, the probability distribution from which a spike count in one neuron is drawn is unrelated to the activity of other neurons. In that case, the response distribution of the population is a product distribution: <\/span><\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/www.braininspirednavigation.com\/wp-content\/uploads\/2018\/01\/011018_0237_PopulationE5.png\" alt=\"\" \/><\/p>\n<p><span style=\"font-size: 12pt;\"><strong><span style=\"background-color: white;\">Tuning<\/span> curve of a single neuron <\/strong><\/span><\/p>\n<p><span style=\"font-size: 12pt;\">Mean response as a function of the stimulus <\/span><\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/www.braininspirednavigation.com\/wp-content\/uploads\/2018\/01\/011018_0237_PopulationE6.png\" alt=\"\" \/><\/p>\n<p><span style=\"font-size: 12pt;\"><strong>Population activity on a single trial <\/strong><\/span><\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/www.braininspirednavigation.com\/wp-content\/uploads\/2018\/01\/011018_0237_PopulationE7.png\" alt=\"\" \/><\/p>\n<p><span style=\"font-size: 12pt;\"><strong>1.2 Decoding <\/strong><\/span><\/p>\n<p><span style=\"font-size: 12pt;\">Based on a pattern of activity <strong>r<\/strong>, the brain often has to reconstruct what was the stimulus <em>s <\/em>that gave rise to <strong>r<\/strong>. This is called decoding, estimating, or reading out <em>s<\/em>. There are many ways of doing this, some of which are better than others. <\/span><\/p>\n<p><span style=\"font-size: 12pt;\"><strong>Decoding population activity <\/strong><\/span><\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/www.braininspirednavigation.com\/wp-content\/uploads\/2018\/01\/011018_0237_PopulationE8.png\" alt=\"\" \/><\/p>\n<p><span style=\"font-size: 12pt;\"><strong>1.2.1. Winner-take-all decoder <\/strong><\/span><\/p>\n<p><span style=\"font-size: 12pt;\">Suppose that each neuron in the population has a preferred stimulus value. Then, a simple <\/span><span style=\"font-size: 12pt;\">estimator of the stimulus is the preferred stimulus value of the neuron with the highest <\/span><span style=\"font-size: 12pt;\">response: <\/span><\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/www.braininspirednavigation.com\/wp-content\/uploads\/2018\/01\/011018_0237_PopulationE9.png\" alt=\"\" \/><\/p>\n<p><span style=\"font-size: 12pt;\"><strong>1.2.2. Center-of-mass decoder <\/strong><\/span><\/p>\n<p><span style=\"font-size: 12pt;\">A better decoder is obtained by computing a weighted average of the preferred stimulus values of all neurons, with weights proportional to the responses of the respective neurons: each neuron &#8220;votes&#8221; for its preferred stimulus value with a strength proportional to its response: <\/span><\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/www.braininspirednavigation.com\/wp-content\/uploads\/2018\/01\/011018_0237_PopulationE10.png\" alt=\"\" \/><\/p>\n<p><span style=\"font-size: 12pt;\">While this method performs quite well in many cases, it does not take the form of neuronal variability into account. On a circular stimulus space \u2013 for instance when the stimulus is orientation or motion direction \u2013 the equivalent of this method is called the population vector. It has been applied to experimental data, for instance by Georgopoulos et al. (Georgopoulos,Kalaska et al. 1982) to decode movement direction from population activity in primate motor cortex. <\/span><\/p>\n<p><span style=\"font-size: 12pt;\"><strong>1.2.3. Template-matching decoder <\/strong><\/span><\/p>\n<p><span style=\"font-size: 12pt;\">We can ask the question for which stimulus value <em>s <\/em>the observed population response is closest to the mean population response generated by <em>s<\/em>. That is, we match the observed population response with a set of templates (mean population responses for different <em>s<\/em>). As an error measure, we use the sum\u2010squared difference. This gives the decoder <\/span><\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/www.braininspirednavigation.com\/wp-content\/uploads\/2018\/01\/011018_0237_PopulationE11.png\" alt=\"\" \/><\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/www.braininspirednavigation.com\/wp-content\/uploads\/2018\/01\/011018_0237_PopulationE12.png\" alt=\"\" \/><\/p>\n<p><span style=\"font-size: 12pt;\">This decoder also does not take the form of neuronal variability into account, but uses more <\/span><span style=\"font-size: 12pt;\">than only the preferred stimulus values of the neurons. <\/span><\/p>\n<p><span style=\"font-size: 12pt;\"><strong>1.2.4. Maximum-likelihood decoder <\/strong><\/span><\/p>\n<p><span style=\"font-size: 12pt;\">An important method that does take the form of neuronal variability into account is the <\/span><span style=\"font-size: 12pt;\">maximum\u2010likelihood decoder. This decoder computes the probability that a stimulus value <\/span><span style=\"font-size: 12pt;\">elicited the given population response, and selects the stimulus value for which this probability is highest: <\/span><\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/www.braininspirednavigation.com\/wp-content\/uploads\/2018\/01\/011018_0237_PopulationE13.png\" alt=\"\" \/><\/p>\n<p><span style=\"font-size: 12pt;\"><strong>1.2.5. Bayesian decoders <\/strong><\/span><\/p>\n<p><span style=\"font-size: 12pt;\">Bayesian decoders use Bayes&#8217; rule to express the probability of a stimulus given a response, <\/span><span style=\"font-size: 12pt;\"><em>p<\/em>(<em>s<\/em>|<strong>r<\/strong>), as the normalized product of the probability of this response given a stimulus, <em>p<\/em>(<strong>r<\/strong>|<em>s<\/em>), and the prior probability of the stimulus, <em>p<\/em>(<em>s<\/em>): <\/span><\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/www.braininspirednavigation.com\/wp-content\/uploads\/2018\/01\/011018_0237_PopulationE14.png\" alt=\"\" \/><\/p>\n<p><span style=\"font-size: 12pt;\">The prior probability reflects knowledge about the stimulus before the population response is elicited, and can have been generated on the basis of previous experience. The probability distribution obtained in this way is referred to as the posterior probability distribution over the stimulus. When the number of neurons is large, this distribution is usually a narrow normal distribution. It can be collapsed onto an estimate by taking the value that has the highest posterior probability; this is called the maximum\u2010a\u2010posteriori (MAP) decoder (also the mode of the posterior): <\/span><\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/www.braininspirednavigation.com\/wp-content\/uploads\/2018\/01\/011018_0237_PopulationE15.png\" alt=\"\" \/><\/p>\n<p><span style=\"font-size: 12pt;\"><strong>1.3 How good are different decoders? <\/strong><\/span><\/p>\n<p><span style=\"font-size: 12pt;\">Given that there are so many possible decoders, how can we objectively evaluate how good <\/span><span style=\"font-size: 12pt;\">each of them is? Imagine you have a large number of population patterns of activity all generated by the same value of <em>s<\/em>. For each pattern, you apply your decoder of interest to obtain an estimate <em>s<\/em>\u02c6 . Now look at the distribution of estimates, <em>p<\/em>(<em>s<\/em>\u02c6 | <em>s<\/em>). There are several criteria for what makes a decoder good. First, you would like the mean estimate to be equal to the true stimulus value, i.e. <em>s<\/em>\u02c6 = <em>s <\/em>. Here, the average <span style=\"font-family: Cambria Math;\">\u22c5<\/span> is in principle over <em>p <\/em>(<em>s<\/em>\u02c6 | <em>s<\/em>) but can also be regarded as one over <em>p<\/em>(<strong>r<\/strong>|<em>s<\/em>), since \u02c6<em>s <\/em>is a function of <strong>r <\/strong>for all deterministic decoders (but not for the sampling decoder). The difference between the mean estimate and the true stimulus value is called the <em>bias <\/em>of the estimator, and it may depend on <em>s<\/em>: <\/span><\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/www.braininspirednavigation.com\/wp-content\/uploads\/2018\/01\/011018_0237_PopulationE16.png\" alt=\"\" \/><\/p>\n<p><span style=\"font-size: 12pt;\">An estimator is called unbiased if <em>b<\/em>(<em>s<\/em>)=0 for all <em>s<\/em>. It is not very difficult for a decoder to be approximately unbiased. Most of the decoders we described in the previous section are unbiased in common situations. However, not all unbiased decoders are equally good, since not only the mean matters, but also the variance. The smaller the variance, the better the decoder. Whereas bias can be equal to zero, this is not true for the variance \u2013 because of variability in patterns of activity, it is impossible for an unbiased decoder to have zero variance. (It is easy for a <em>biased <\/em>decoder to have zero variance: just take one that ignores the data and always produces the same value. This decoder has no variability, but it is severely biased for all values of <em>s <\/em>but one.) It turns out that there is a fundamental lower bound on the variance of an unbiased decoder. This is a famous result in estimation theory, known as the Cram\u00e9r\u2010Rao inequality. We will go through it in some detail because it is an important notion in population coding. <\/span><\/p>\n<p><span style=\"font-size: 12pt;\"><strong>Good decoders <\/strong><\/span><\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/www.braininspirednavigation.com\/wp-content\/uploads\/2018\/01\/011018_0237_PopulationE17.png\" alt=\"\" \/><\/p>\n<ul>\n<li><span style=\"font-size: 12pt; color: #000000;\">Unbiased <\/span><\/li>\n<li><span style=\"font-size: 12pt; color: #000000;\">Low variance <\/span><\/li>\n<li><span style=\"font-size: 12pt; color: #000000;\">Can be implemented by a neural network. <\/span><\/li>\n<\/ul>\n<p><span style=\"font-size: 12pt;\"><strong>1.3.1. Cram\u00e9rRao bound <\/strong><\/span><\/p>\n<p><span style=\"font-size: 12pt;\">Based on the response distribution <em>p<\/em>(<strong>r<\/strong>|<em>s<\/em>), one can define a quantity called the <em>Fisher <\/em><\/span><span style=\"font-size: 12pt;\"><em>information <\/em>that the population contains about <em>s<\/em>: <\/span><\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/www.braininspirednavigation.com\/wp-content\/uploads\/2018\/01\/011018_0237_PopulationE18.png\" alt=\"\" \/><\/p>\n<p><span style=\"font-size: 12pt;\">where \u2202 denotes a partial derivative and the average <span style=\"font-family: Cambria Math;\">\u22c5<\/span> is now over <em>p<\/em>(<strong>r<\/strong>|<em>s<\/em>). Fisher information can alternatively be expressed as <\/span><\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/www.braininspirednavigation.com\/wp-content\/uploads\/2018\/01\/011018_0237_PopulationE19.png\" alt=\"\" \/><\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/www.braininspirednavigation.com\/wp-content\/uploads\/2018\/01\/011018_0237_PopulationE20.png\" alt=\"\" \/><\/p>\n<p><span style=\"font-size: 12pt;\">This is the Cram\u00e9r\u2010Rao bound (or inequality). It states that no estimator \u02c6<em>s <\/em>can achieve a variance that is smaller than the inverse of the Fisher information. An estimator that has the smallest possible variance is called an <em>efficient estimator<\/em>. The Cram\u00e9r\u2010Rao bound can be generalized to multidimensional (vector) variables. <\/span><\/p>\n<p><span style=\"font-size: 12pt;\"><strong>1.3.2. Fisher information <\/strong><\/span><\/p>\n<p><span style=\"font-size: 12pt;\">Since Fisher information puts a hard limit on the performance of any possible decoder, it is <em>a decoder\u2010independent measure of the information content of a population of neurons <\/em>(which is why it is called &#8220;information&#8221; in the first place). Here, the understanding is that the population is characterized by its response distribution. Fisher information reflects the maximum amount of information that can be extracted from a population. <\/span><span style=\"font-size: 12pt;\">The variance of an estimator determines the smallest change in the stimulus that can be <\/span><span style=\"font-size: 12pt;\">reliably discriminated. If the variance is small, the estimator can be used to detect tiny changes in <em>s<\/em>. Accordingly, there is a link between Fisher information and discrimination threshold: Fisher information is inversely proportional to the square of the discrimination threshold of an ideal observer of the neural activity, or equivalently, it is proportional to the square of the sensitivity <em>d <\/em>&#8216; of an ideal observer: <\/span><\/p>\n<p style=\"text-align: center;\"><img decoding=\"async\" src=\"https:\/\/www.braininspirednavigation.com\/wp-content\/uploads\/2018\/01\/011018_0237_PopulationE21.png\" alt=\"\" \/><\/p>\n<p><span style=\"font-size: 12pt;\">where <em>d <\/em>&#8216; is the sensitivity (a measure of performance) and \u03b4 <em>s <\/em>is the distance between the two stimuli to be discriminated. Fisher information is subject to the data processing inequality, which states that no operation on the data (in our case population patterns of activity) can increase Fisher information. <\/span><\/p>\n<p><span style=\"font-size: 12pt;\">Note: <span style=\"background-color: white;\">the Fisher information (<a href=\"https:\/\/en.wikipedia.org\/wiki\/Fisher_information\">https:\/\/en.wikipedia.org\/wiki\/Fisher_information<\/a> ) may be seen as the curvature of the support curve (the graph of the log-likelihood). Near the maximum likelihood estimate, low Fisher information therefore indicates that the maximum appears &#8220;blunt&#8221;, that is, the maximum is shallow and there are many nearby values with a similar log-likelihood. Conversely, high Fisher information indicates that the maximum is sharp. <\/span><\/span><\/p>\n<p><span style=\"font-size: 12pt; background-color: white;\"><strong>1.3.3. Goodness of decoders <\/strong><\/span><\/p>\n<p><span style=\"font-size: 12pt;\">There are some general results that are helpful here. It turns out that in the limit of a <\/span><span style=\"font-size: 12pt;\">large number of observations (in our case, many spikes), the maximum\u2010likelihood estimate is the &#8220;best possible&#8221; one in the sense that it is both unbiased and efficient. Moreover, its distribution is approximately normal in this limit. <\/span><\/p>\n<p><span style=\"font-size: 12pt;\">Other decoders will have a larger bias and\/or a larger variance than the maximum\u2010likelihood decoder, although in simple situations, some of them may come very close. Typically, the winner\u2010take\u2010all decoder and the sampling decoder are rather poor, but one might choose them for their computational advantages. <\/span><\/p>\n<p><span style=\"font-size: 12pt;\"><strong>1.3.4. Neural implementation <\/strong><\/span><\/p>\n<p><span style=\"font-size: 12pt;\">So far, we have discussed many decoders from an abstract perspective. However, the brain <\/span><span style=\"font-size: 12pt;\">itself also has to do decoding, for example in generating a response to a stimulus. Therefore, if we want to know whether a particular decoder is used in performing a perceptual task, the question needs to be answered how it can be implemented in neural networks. Fortunately, this problem has been solved for the maximum\u2010likelihood decoder. Under certain assumptions on the form of neural variability, a line attractor network can turn a noisy population pattern of activity into a smooth pattern that peaks at the maximum\u2010likelihood estimate (Deneve, Latham et al. 1999). <\/span><\/p>\n<p><span style=\"font-size: 12pt;\">The winner\u2010take\u2010all decoder can easily be implemented using a nonlinearity (to enhance the <\/span><span style=\"font-size: 12pt;\">maximum activity) and global inhibition (to suppress the activity of other neurons). Neural <\/span><span style=\"font-size: 12pt;\">implementations of the other decoders we discussed are not known. <\/span><\/p>\n<p><span style=\"font-size: 12pt;\"><strong>1.4 Decoding probability distributions <\/strong><\/span><\/p>\n<p><span style=\"font-size: 12pt;\">Further info in the <a href=\"http:\/\/www.caam.rice.edu\/~caam415\/Ma_lectures\/M1_Notes_Population_coding.pdf\">lecture notes<\/a>. <\/span><\/p>\n<p>&nbsp;<\/p>\n<p><span style=\"font-size: 12pt;\">Deneve, S., P. Latham, et al. (1999). &#8220;Reading population codes: a neural implementation of ideal observers.&#8221; Nature Neuroscience <strong>2<\/strong>(8): 740\u2010745. <\/span><\/p>\n<p><span style=\"font-size: 12pt;\">Foldiak, P. (1993). The &#8216;ideal homunculus&#8217;: statistical inference from neural population responses. Computation and Neural Systems. F. Eeckman and J. Bower. Norwell, MA, Kluwer Academic Publishers<strong>: <\/strong>55\u201060. <\/span><\/p>\n<p><span style=\"font-size: 12pt;\">Georgopoulos, A., J. Kalaska, et al. (1982). &#8220;On the relations between the direction of twodimensional arm movements and cell discharge in primate motor cortex.&#8221; Journal of Neuroscience <strong>2<\/strong>(11): 1527\u20101537. <\/span><\/p>\n<p><span style=\"font-size: 12pt;\">Pouget, A., P. Dayan, et al. (2003). &#8220;Inference and Computation with Population Codes.&#8221; Annual Review of Neuroscience. Sanger, T. (1996). &#8220;Probability density estimation for the interpretation of neural population codes.&#8221; Journal of Neurophysiology <strong>76<\/strong>(4): 2790\u20103. <\/span><\/p>\n<p><span style=\"font-size: 12pt;\">Zhang, K., I. Ginzburg, et al. (1998). &#8220;Interpreting neuronal population activity by reconstruction: unified framework with application to hippocampal place cells.&#8221; Journal of Neurophysiology <strong>79<\/strong>(2): 1017\u201044. <\/span><\/p>\n","protected":false},"excerpt":{"rendered":"<p>The content is from the lecture notes (M1_Slides_Population_Coding, M1_Notes_Population_Coding) of &#8216;Theoretical Systems Neuroscience&#8217; by Professor Wei Ji Ma (Baylor College of Medicine) in 2013. He is now leading the Wei Ji Ma&#8217; lab at Center for Neural Science and Department of Psychology at New York University. For further more info on the course website CAAM\/NEUR [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[114,96],"tags":[232,229,233,234,231,227,230,228],"_links":{"self":[{"href":"https:\/\/braininspirednavigation.com\/index.php?rest_route=\/wp\/v2\/posts\/739"}],"collection":[{"href":"https:\/\/braininspirednavigation.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/braininspirednavigation.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/braininspirednavigation.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/braininspirednavigation.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=739"}],"version-history":[{"count":3,"href":"https:\/\/braininspirednavigation.com\/index.php?rest_route=\/wp\/v2\/posts\/739\/revisions"}],"predecessor-version":[{"id":960,"href":"https:\/\/braininspirednavigation.com\/index.php?rest_route=\/wp\/v2\/posts\/739\/revisions\/960"}],"wp:attachment":[{"href":"https:\/\/braininspirednavigation.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=739"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/braininspirednavigation.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=739"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/braininspirednavigation.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=739"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}