The Most Beautiful Statistic: The History and the Science of the Humble Mean

The Most Beautiful Statistic: The History and the Science of the Humble Mean

, 1976, the Viking 1 spacecraft entered Mars orbit. A month later, as a midsummer night stretched pleasantly ahead of them in Pasadena, California, dozens of NASA controllers helplessly watched a feed delayed by 19 minutes as the Viking 1 lander separated from the orbiter. Three hours later, the lander plunged into the Martian atmosphere at 10,000 miles per hour. Ten minutes later, following a fiendishly complex approach, and with zero help from JPL, the lander made a perfect touchdown on Mars. For Mission Control, it hadn’t helped that, when Viking 1 entered Martian orbit and pointed its camera down at the surface, the originally selected landing site turned out to be “at least twice as rough as the Mars average” prompting a frantic search for a new site. Fortunately, a safer site was found which, in the scientists’ words, was still “rougher than the Martian average,” but “near the Martian average for elevations accessible to Viking” and “near the Mars average in reflectivity.” Even at 212 million miles from Earth, the humble mean turned out to be the statistical device guiding Viking 1 to a safe spot in an unfriendly world. The mean keeps making its usefulness felt in all sorts of situations, often in truly non-obvious ways. Consider, for instance, another happenstance. Five days after the Viking lander touched down on Mars, on July 25, 1976, the Viking orbiter transmitted to Earth, the following photograph. “The Face on Mars”: A close-up of the Cydonia region of Mars (Source: NASA. Copyright: Public domain) The photo was immediately dismissed as a mere play of light and shadow coupled with missing data (the dark “speckles” in the image). Nevertheless, it became one of the most famous examples of pareidolia: our tendency to detect patterns, especially faces, and ascribe meaning to them. And it is not just faces on Mars or the Moon. People have spotted animal shapes and UFOs amid clouds. Down on Earth, people have spotted cups and saucers, saucepans with handles, flagpoles, hammers, people hanging from poles, triangles, rectangles, and above all, skyrocketing trends in charts and plots that have been comfortably out on little more than a random walk. It’s not that these patterns do not exist. Often, they do exist, if only lightly and ephemerally, and when they are there, our vision system has evolved to detect their presence almost every single time. Sadly, the patterns often never mean what we think they do. So, where in all this is the connection with the mean? A Representative for Every Occasion The connection with the mean emerges from the counterfactual, i.e., from data in which you do not instantly see saucepans or people hanging from flag poles. In other words, data in which trends and patterns cannot be easily detected. Often, trends in data aren’t expressed in ways that give you an evolutionary advantage to detect them. So, your eyes fail to see them. Your brain fails to interpret them. That is when you sorely need devices like the mean, median, mode, and percentile scores. Such statistical measures as the mean can compress the noise and pull out the pattern in a measurable and actionable form, provided it exists. Consider the following three charts showing daily absenteeism in NYC public schools in school years 2015-16, 2016-17, and 2017-18. The X-axis represents school days, and the Y-axis represents percentage daily absenteeism. Each bar represents the aggregate daily absenteeism rate across all NYC schools for a particular school day such October 22, 2015. Daily absenteeism in NYC public schools in school years 2015-16, 2016-17, and 2017-18 (Data source: NYC OpenData. Data Tems of Use) This is a classic example of a panel data set in which the units, i, are school days, the time variable, t, is the school year, and the response variable, yit, is the aggregate daily percentage absenteeism across all schools on school day, i, in year, t. From the charts alone, can you tell if absenteeism has worsened, improved, or simply meandered without a clear trend over the three years that constitute the data panel? Here is another example showing the state-level annual House Price Index (HPI) values in the United States from 2020 to 2025. The X-axis represents a U.S. state, and the Y-axis represents the annual HPI for that state. Each bar represents a state’s annual HPI value for the corresponding year, assuming a base of 100 in 1975. Annual House Price Index by US state (2020 – 2025). Each bar represents the state’s annual House Price Index value for the corresponding year, assuming a base of 100 in 1975 (Data source: FHFA. Data copyright: Public domain) What is your take on the house price trend expressed in this 6-year data panel? Have prices increased, decreased, or meandered listlessly? It is hard for us to pick up trends and patterns in such data. So, what would be a good representative that would efficiently bring out the trend, if one exists? Would the maximum or the minimum value be a good candidate? How about the mode (the most frequently occurring value), or the median (the value around the 50th percentile)? The statistic that you choose will depend entirely on how you intend to use it. Let’s revisit the daily absenteeism data panel. The U.S. Department of Education defines chronic absenteeism as students missing “10% or more of school”. Strictly speaking, chronic absenteeism is a student-level yearly measure. But if a school system had an additional operational goal of keeping the daily absenteeism rate also below 10%, then they’d want to monitor the maximum daily absenteeism rate. The following chart uses the data in the absenteeism data panel shown earlier. For each school year, it shows the maximum daily absenteeism across all schools. It is immediately evident that plotting the daily maximum value paints a troubling picture of daily school attendance. Max percentage absenteeism by school year in NYC public schools Now let’s revisit the House Price Index data panel. If federal economic policy requires the monitoring of the median annual state-level House Price Index, policymakers will want to keep an eye on a chart like the following. The following chart uses the data from the HPI data panel and shows the median annual HPI across the 50-states-plus-DC panel from 2020 to 2025. Again, the trend is immediately visible when we use a suitable summary statistic. Median annual state-level HPI. Each data point is the median of the annual HPIs of the 50-states-plus-DC data panel for the corresponding year. How about using the sum as a summary statistic? There are two problems with using the sum. First, its value is often proportional to the sample size. The sum confounds the “typical” (summary) value with the number of observations, making it tricky to compare samples of different sizes. Second, often the sum doesn’t mean anything. Adding observations is sensible only if they are additive. Observations such as indexes, rates, and percentages are not additive. For example, the sum of House Price Indexes across all states is a real head-scratcher. But let us not be so quick to dismiss the sum. The Mean as a Probability-Weighted Sum Suppose we tweak the sum of the observed values so that the importance of each distinct value is weighted by just the right factor, a factor that reflects how frequently that value is observed. Normally, when you multiply an integer A by another integer B, B acts as a numerical magnifying glass for A, and vice versa. You are magnifying the effect of A by B. But here, we’d like to do the opposite. We would like to “weaken” the importance of each distinct value by a factor. The factor we choose is that value’s probability of occurrence, p. Let us develop this idea. Suppose there are k distinct values x1, x2, …, xi, …, xk in a population of size N. Notice that k ≤ N since each xi can occur more than once in the population. Let Ni be the frequency of occurrence of xi. What is the probability of observing xi? It is the following ratio. Using the probabilities, pi, we express the probability-weighted sum of the k distinct values in the population as follows. Equation (1) is the formula for the population mean μ (“mu”), and it uses the population probabilities p1, p2, …, pi, …, pk. The population probability represents the true frequency of occurrence of the value in the entire population. Expected Value Eq. (1) is also the formula for the expected value or the unconditional expectation of the discrete random variable X which can take one of k distinct values x1, x2, …, xi, …, xk with corresponding probabilities p1, p2, …, pi, …, pk. The Arithmetic Mean Suppose all distinct values x1 to xk in Eq. (1) occur at the same frequency, m. Notation-wise, we represent this situation as follows. Then, each probability, pi, is the following constant value. In this case, the mean can be expressed as follows. Since m times xi is the same as (xi + xi + xi …m times), Eq. (1a) can be written as a simple sum over all N values in the population, as follows. Eq. (1b) is the formula for the arithmetic mean of N values. In the simplest case, m = 1, that is, each distinct value in the population occurs exactly once. In other words, every value in the population is distinct. Then, we have the following situation. In this case, k = N, and the arithmetic mean of N values takes the following familiar form. Eq. (1c) is the familiar formula for the arithmetic mean of N values, where every value in the population has the same probability 1/N. The Sample and the Population In most cases, you wouldn’t have access to the entire population. Often, the population size is unknown. Think about a country’s population: for most countries the exact population is unknown. In practice, you would be working with a finite sample of size n in which you have observed q distinct values x1, x2, …, xi, …, xq where q ≤ k ≤ N, where k is the number of distinct values in the underlying population of size N. In this finite sample of size n, you will note down how frequently each distinct value xi occurs. For the ith distinct value xi, this gives you the sample probability pi_hat (pronounced “p i hat” or “p hat i”). pi_hat is also called an empirical probability because it arises from an experiment involving the selection of an appropriate sampling strategy and taking measurements. Consequently, a statistic computed using empirical probabilities is also empirical in nature. Using pi_hat, we write the formula for the sample mean as follows. The n in X_bar_n denotes the sample size. While the population mean in Eq. (1) is usually a theoretical quantity, the sample mean in Eq. (2) is practically computable. The sample mean serves as your working estimate of the population mean. All of us use the sample mean because we rarely, if ever, have access to the population. Consider restaurant ratings, for example. The 3.7-or-so rating that you saw on Google for the new restaurant in town was the mean of several 1-star, 2-star, 3-star, 4-star, and 5-star ratings left by a fraction of all patrons who dined there. Even Google cannot coax every patron into leaving a rating. That makes the star rating a sample mean and, therefore, a mere estimate of the true rating. In fact, things get slightly worse. The rating you saw was the mean of the ratings left by only those patrons who decided to give a rating. If you believe that people who had either an exceptionally fabulous or an exceptionally lousy experience are more likely to leave a rating, then 1-star, 2-star, and 5-star ratings will dominate the sample. This dominance will necessarily come at the expense of the 3- and 4-star ratings. Such over-representation of the extreme values can tilt the mean toward one or the other end of the scale, causing the mean to lose accuracy. The mean becomes a biased estimate of the “true” rating. The following chart illustrates this situation. The sample (shown by the orange bars) contains an over-representation of 1- and 2-star ratings, causing the sample mean to become biased. In this case, the sample mean is biased downward. Platforms such as Google are surely aware of such biases, of course. And they almost certainly use every trick in the statistician’s handbook to tame them, though perhaps at the cost of increasing the variance of the estimated rating. In other words, the displayed rating becomes less biased, but it may also become a less precise estimate of the true rating. This tension between systematic error and random variation (the bias-variance tradeoff) influences the design of many statistical and neural-net models. Biased or not, the sample is all you will have, and the sample value of the statistic, whether the sample mean, the sample median, the sample mode, or any other statistic, is what you will use as a working estimate for making decisions. In fact, you would be surprised how much data science, nay, science itself, is devoted to estimating the population mean, population median, population whatever, based on the sample value while beating the bias out of that estimate. The Role of Probability, p, in the Mean’s Formula In the probability-weighted sum formula for the mean, why should the weight be specifically p, and not some function of p like p2 or in general pr (r ≠ 1), or ep? There are two strong reasons why p is the perfect choice for the weight in the formula for the ordinary mean. The Philosophical Reason The first reason is philosophical (or common-sensical, depending on your perspective), and is best explained with an example. Suppose your data contains only three distinct values. Suppose the first value is observed 70% of the time, the second one only 20% of the time, and the third the remaining 10% of the time. You would want a fair sample containing all three distinct values to give more weight to the first value, much less weight to the second, and even less weight to the third. Specifically, you would want the three values to carry weights in the ratio: 70 : 20 : 10, or 70/100 : 20/100 : 10/100. Thus, weighting each distinct value by its frequency, or standardized frequency, in other words, the probability, p, agrees pleasantly with our sense of fairness. The Connection of the Mean with the Weak Law of Large Numbers The second reason for using p is based on something profound. To know what it is, we must go way back in time. Specifically, we must travel to the city-state of Basel in the old Swiss confederacy in the 1680s, when a little discovery, later to be known as the Weak Law of Large Numbers, was made by the one and only Jakob Bernoulli (1655-1705). Jakob Bernoulli (1654 – 1705) (Base Image: Wikipedia. Public domain) I am going to indulge you in a slight detour here to give you a short explanation of the Weak LLN, just enough for us to address the question at hand: why is p the correct choice for the weight? If you are eager to know the answer right away, it is this: If the target is the ordinary population mean μ, p is the only choice of weight in the formula for the mean that respects the Weak Law of Large Numbers. If you are happy to take that answer at face value, you don’t need to know how the Weak LLN works and why p, and only p, respects the Weak LLN. But in case you are curious, read on. A Short Introduction to the Weak Law of Large Numbers The Weak Law of Large Numbers says, basically, the following. First, assume that your observations are independent and identically distributed, or i.i.d. “Independent” means the probability of observing a value isn’t influenced by the probability of observing any other value. “Identically distributed” means that your experimental setup doesn’t vary from one observation to the next, and all observations are generated by the same underlying probability distribution. Also assume that the population mean exists and the population variance is finite. Second, even with i.i.d. observations, and no matter how hard you’ve worked to remove bias from your sample, your sample mean will generally contain an error with respect to the population mean. In other words, your sample mean will generally be an estimate of the population mean. Now comes the payoff. Suppose you choose some positive threshold value, ϵ, for the error you can tolerate. Bernoulli’s big idea was that as you collect larger samples, it becomes increasingly hard to come across a sample sporting an absolute error larger than your tolerance level, ϵ. And this behavior does not depend on how vanishingly small an error threshold you select. This behavior is driven by a concept in statistics called convergence in probability. The bottom line is that the sample mean becomes progressively more reliable with larger samples. This is the Weak Law of Large Numbers, and Jakob Bernoulli quantified precisely how its machinery works. The exact statement of the Weak LLN runs as follows. Consider a random sample made up of n independent, identically distributed random variables X1, X2, …, Xn. Let the population mean of the underlying distribution on which the random variables are defined be μ. Let Xn_bar be the sample mean. For any positive real number ϵ, the probability that the sample mean is more than ϵ away from μ approaches zero as the sample size approaches infinity. Under these assumptions, the Weak LLN can be stated as follows. The Weak Law of Large Numbers implies that, under suitable assumptions, the sample mean converges in probability to the population mean. It follows that the sample probability, p_hat, also converges in probability to the population probability, p. To see why this is so, assume that X1, X2, …, Xn are n binary (1/0) random variables, each one denoting the outcome of a random trial. Let 1 denote the observation of a particular value, such as 5, and 0 denote any other observation, such as any value other than 5, in a single trial. In a random sample of size 10, i.e., 10 i.i.d. trials, suppose you see three 5s. So, your sample might look like the following set. [1, 0, 1, 0, 0, 0, 0, 0, 1, 0] The sample probability pi_hat of observing a 5 is 3/10. The sample mean also happens to work out to the same value 3/10, as follows. Thus, in this case, the sample mean is also the sample probability, and the Weak LLN says that both converge to the population mean μ, whatever that might be, which for a binary indicator variable is exactly the population probability, p. With this background, let’s return to showing why probability p is the right choice for the weight in the formula for the mean. In an i.i.d. sample of size n containing q distinct values x1, x2, …, xq, suppose the ith distinct value xi is observed ni times. Then the sample probability pi_hat = ni /n, and the sample mean can be expressed as follows. Next, we take identical limits on both sides of the above equation as follows. Now let’s take the limit on the right-hand side inside the summation. We’ve seen how sample probabilities converge to their population counterparts. Substituting (6) inside (5), we arrive at the following result. Thus, the choice of p as the weight causes the sample mean to converge to the population mean, thereby respecting the Weak LLN. Now, suppose we use p2 as the weight instead of p. The resulting statistic is obviously not the ordinary sample mean X_bar_n mentioned in Eq. (2). Let’s call it Tn. The summation of pj2over {1,…,q} in the denominator is essential to standardize the weights. Incidentally, we also performed this standardization while using pi via ∑ pj. But there, we did it invisibly because ∑ pj=1. That’s ’cause, in any random trial, exactly one of xi where i ϵ {1,…,q} is guaranteed to be observed. Therefore ∑ pj = 1 when summed over q is always 1. Therefore, mentioning this sum explicitly in the denominator wasn’t needed. Let’s return to Eq. (7) and, as before, let’s take identical limits on both sides. Also as before, we’ll take the right-hand side limit inside the summation. Remember that the Weak LLN causes the sample probability pi_hat to converge to the population value pi. So, a continuous transformation of pi_hat such as p2i_hat will also converge to the corresponding population value p2i. Thus, Eq. (8) will converge in probability to the corresponding population value as follows. But the population value on the right-hand side of the above equation is not the ordinary mean μ. Thus, p2 weighting fails for estimating the ordinary mean as it converges to the wrong target. It estimates an entirely different statistic. The same issue arises with any other standardized transformation of p such as ep. So, You Want to Build Your Own Universe Let us not get ahead of ourselves. Suppose we still use p2, or ep as the weight. So what if the sample statistic Tn doesn’t converge to the ordinary population mean μ? What’s so bad about that? Well, to quote the line popularized by actor Tony Shalhoub in the detective series Monk: “Here’s the thing.” Tony Shalhoub (Wikipedia CC BY-SA 4.0) Suppose your stated target is some well-known population statistic such as the ordinary mean, μ. If your chosen sample statistic does not converge to your stated target, then you have made yourself a fractured promise. In doing so, you have also succeeded in constructing a teacup-universe, a make-believe world of your own liking, a Tolkienish existence if you like (no offense meant to Tolkien fans). A world that the fundamental laws of the universe, as we know them, barely bother to visit. There is a somber ending to this story. In 1705, Jakob Bernoulli died of tuberculosis when he was just 50 years old. This was an era when TB was rampant, yet doctors didn’t know what caused it or how to cure it. Diagnosis often amounted to a death sentence, and the decline could be long and painful, lasting from months to a few years. The patient gradually wasted away, “consumed” as it were by something mysterious inside the body. It would be nearly two more centuries before Robert Koch discovered the cause of TB: the microbe Mycobacterium tuberculosis. When Jakob Bernoulli died, he also left unfinished his magnum opus, Ars Conjectandi (“The Art of Conjecturing”), which contained his discovery of the Weak LLN. Fortunately for future generations, his nephew Nicolaus Bernoulli (1687–1759) managed to publish his uncle’s work in 1713, a full eight years after Jakob’s death. The Mean’s Connection to Squared Loss There happens to be a third, indirect and non-obvious reason for using probability p as the weight. The role of p in the mean emerges from the mean’s deep historical connection with squared loss and the method of least squares estimation. To know about this connection and the circumstances that gave it shape, we must once again tunnel back into history, this time to a point roughly a century after Jakob’s discovery of the Weak LLN in the late 1680s. We now enter Napoleonic Europe of the early 1800s, and the professional lives of two mathematicians on opposite sides of Napoleon Bonaparte’s conquests: the prolific Parisian mathematician Adrien-Marie Legendre (1752 – 1833), and the extraordinarily gifted polymath Johann Carl Friedrich Gauss (1777 – 1855) who had a knack of plucking brilliant ideas out of the ether, and who lived in the Duchy of Brunswick-Wolfenbüttel in what is now modern-day Germany. Left: A highly exaggerated caricature of Adrien-Marie Legendre (1777 – 1855) painted by Julien-Léopold Boilly in 1820. It is also the only surviving authentic portrait of the great man (Wikipedia, Public domain)Right: An 1803 portrait of Johann Carl Friedrich Gauss (1777 – 1855) by Johann Christian August Schwartz (Wikipedia, Public domain) The Method of Least Squares and its Connection with the Arithmetic Mean In 1805, exactly 100 years after Bernoulli died in Basel leaving Ars Conjectandi unfinished, Legendre, one of mathematics’ all-time greatest heroes, published a book in Paris whose appendix would become the focal point of statistical discourse for the next 100 years. The book was Nouvelles Méthodes Pour La Détermination Des Orbites Des Comètes (New Methods for Determining Comet Orbits), and the appendix contained the method of least squares estimation. It so happens that the humble mean owes a deep debt to the method of least squares. So deep is this debt that had it not been for the method of least squares, the mean would have lost quite a lot of its usefulness. The following example will illustrate this deep connection. Suppose you wish to estimate an unknown quantity x. x could be the position of a new star you discovered in the night sky. To pin down x, you point your telescope at the presumed location of the star on three successive nights and make three error-ridden (a.k.a. “noisy”) observations x1, x2, and x3 such that x1, x2, x3 can be expressed as the true value x plus an error, as follows. x1 = x + ε1x2 = x + ε2x3 = x + ε3 Remember that the true x is unknown. The best you can do is compute a smashingly good estimate of the star’s true location using the three observations x1, x2, x3. In 1805, Legendre had a groundbreaking brainwave for creating such an estimate: he postulated that the best estimate of x would be one that heavily penalizes large errors, both negative and positive. The following are Legendre’s words, translated into English. “Among all the principles that can be proposed for this purpose, I think there is no one more general, more exact, and more easy to apply than that which we have made use of in the preceding researches, and which consists in making the sum of the squares of errors a minimum. In this way there is established a sort of equilibrium among the errors, which prevents the extremes to prevail and is well suited to make us know the state of the system most near to the truth.” (Legendre, 1805, pp. 72–73) Unfortunately, it’s not possible to calculate these errors directly since each error εi is the difference between the observed xi and the actual x and the actual x is unknown. Therefore, in practice, we must use the residual ei which is the difference between the observed xi and the estimated value of x, that is, ei = (xi – x_hat). Note the difference in notation between the error and the residual. And note the difference in meaning between the two. εi denotes the unknown observation error, while ei denotes the residual which forms the estimate of the observation error. Effectively, Legendre had proposed that the best estimate of x would be one that minimizes the sum of squared residuals between the estimate and each observation xi. Thus, the best possible estimate x_hat would be one that minimizes the following sum. In Eq. (9), S(x_hat) is a function of x_hat. To find the value of x_hat that minimizes the value of this function, we must take its derivative with respect to x_hat, set it to 0, and solve for x_hat as follows. By this time, you might have developed at least a vague premonition of where this derivation is headed. For look at what happens when we set the derivative in Eq. (10) to 0 and solve for x_hat. Out pops the arithmetic mean of the three observations! Now, tell me, is that neat or what? Also note that the second derivative of S(.) is 6 > 0 (see below), indicating that at the arithmetic mean of observations, the squared loss in the observations is minimized. In general, given an unknown x, and a set of noisy observations x1, x2, x3, …, xn, of that quantity, the simple arithmetic mean of those observations is the estimate that minimizes the sum of squared residuals between the estimate and the observations. This arithmetic mean is the optimal estimate of x under least squares loss. Conversely, the sum of squared residuals between the estimate and the observations is minimized at a point where the estimate is the arithmetic mean of the observations. Note once again that we are minimizing the sum of squared residuals, not errors, even though the method is usually called the method of least squared errors, or simply, least squares. Now, let’s dial up our ambitions a bit. Expected Squared Loss and its Connection with the Mean Going beyond the special case of the arithmetic mean, if our goal is to minimize the expected value of the squared error (a.k.a. expected squared loss), then once again, the mean is the only representative of the observed sample that will meet this design goal. This can be proved as follows. Let x1, x2, x3, …, xn be several observations of a random variable X. To illustrate, X could represent the air temperature recorded at 12 PM CEST on July 1 by the historic Parc Montsouris weather station in Paris. Then, x1, x2, x3, …, xn are the readings taken on July 1 in n randomly selected years. Suppose we wish to find a suitable constant representative c for X. The squared residual between xi and c is the following. Next, we’ll use the following formula for the expected value of a discrete random variable X. Using the above formula and empirical probabilities pi_hat, the empirical expected squared loss E_hat(ei2) for a sample of size n containing q distinct observed values is given by the following formula. As before, to find the value of c that minimizes the empirical expected squared loss, 𝐿(c), we’ll differentiate 𝐿(c) with respect to c, set the derivative to 0, and solve for c as follows. The second derivative of 𝐿(c) with respect to c is as follows. Since 𝐿”(c) is positive, setting Eq. (11) to zero will give us the value of c that minimizes the squared loss. Let us set Eq. (11) to zero and solve for c. Thus, if our design goal is to find a constant representative for a random variable X that minimizes the empirical expected squared loss, that is, the following. Then, the solution is the empirical expectation. In other words, the following. In simple terms: The empirical expectation E_hat(X) minimizes the empirical expected squared loss between the observed values of a random variable X and a constant representative c of that random variable. Analogously, the population mean, μ, minimizes the population expected squared loss, E[(X – c)2]. Conversely, The empirical expected squared loss between the observed values of X and a constant representative c of that random variable is minimized at a point where the estimate c is the empirical expectation E_hat(X) of X. Analogously, the population expected squared loss, E[(X – c)2] is minimized at a point where the population representative is the expected value E(X) of X. So far, we’ve been working in 1-dimensional space where we’ve seen the connection between the mean and the squared loss for a single random variable X. In 1805, when Legendre published his method of least squares, he went way beyond the single unknown. His goal was to show how minimizing the sum of squared errors could be used to estimate the optimal values of all the coefficients in a system of linear equations in which the number of observations exceeded the number of coefficients. Effectively, in the spring of 1805 when Legendre published “Nouvelles Méthodes…”, he had taken aim at nothing short of the Linear Regression model, and he presented a disarmingly simple least-squares based technique of estimating all the coefficients of such a model. Legendre, Gauss, and a Matter of Priority Meanwhile, five hundred miles east of Paris, the young and extraordinarily talented Gauss was also busy working on a book: a tome on celestial mathematics. In the early 1800s, Gauss could not have been unaffected in his work by the political climate of Europe. Napoleon Bonaparte’s armies were sweeping across the European continent with ever gathering speed. The Duchy of Brunswick-Wolfenbüttel he lived in was strongly aligned with the Kingdom of Prussia, and by Autumn of 1806, Napolean Bonaparte was at Gauss’s doorstep. For Gauss, 14 October 1806 was likely a life-changing day. On this day, Napolean’s armies engaged with the Kingdom of Prussia in two decisive battles at Jena and Auerstedt. The Duke of Brunswick commanded the troops at Auerstedt. It was an all-out disaster. The duke was shot through the eyes and left blind, helpless, and unable to lead. In subsequent weeks, he would die of his injuries. By the end of the day on 14 October, Prussia had lost both battles. It would soon loose half of its territory to France. With the Duke defeated and dead, Gauss’s homeland of Brunswick also came under Napoleonic occupation. In the following year, Napoleon dissolved the Duchy, incorporated its lands into a freshly minted territory called the Kingdom of Westphalia, and handed it over to his youngest brother, the wildly extravagant Jérôme Bonaparte who somehow managed to empty Westphalia’s treasury despite the punishing war taxes imposed by the Emperor on the occupied lands. The French empire, occupied territories, dependents, and alies in 1812 (Base image source: Wikipedia Copright: CC BY-SA 3.0) Meanwhile, in 1806, Gauss had finished working on the German language version of his book but he found it hard to find a local publisher. Eventually, he found one who agreed to publish the work under the condition that it be translated into Latin – probably to reach a pan-European scientific audience. After all, given the heavy war levies coupled with Mr. J. Bonaparte’s spending stunts, it might be safe to assume that academic institutions and celestial observatories in Gauss’s homeland were not brimming with free cash to buy scientific tomes. Finally in 1809 the Theoria motus corporum coelestium in sectionibus conicis solem ambientium (Theory of the Motion of the Heavenly Bodies Moving about the Sun in Conic Sections) saw the light of day. In his book, Gauss introduced his own version of the method of least squares. While Legendre’s exposition was purely algebraic, Gauss couched his method in the principles of probability and inverse probability which had been developed by Pierre-Simon Laplace in France. While introducing his own method, Gauss made a passing reference to Legendre’s 1805 paper: “…our principle, which we have made use of since the year 1795, has lately been published by LEGENDRE in the work Nouvelles méthodes pour la détermination des orbites des comètes, Paris 1806, where several other properties of this principle have been explained, which, for the sake of brevity, we here omit.” The emphasis is all mine. It’s hardly surprising that when the above claim reached Legendre, it triggered an almost immediate response from the older, highly principled mathematician: a response that got the young Gauss into a lot of hot water. On May 31, 1809, Legendre shot off a letter to Gauss from Paris. After a polite preamble dripping with high praise for the younger mathematician’s achievements, Legendre came to the real point of his missive: “…I felt some regret on seeing that, while citing my memoir on page 221, you say principium nostrum quo jam inde ab anno 1795 usi sumus, etc. There is no discovery that one could not claim for oneself by saying that one had found the same thing some years earlier. But if one does not provide proof of this by citing the place where one published it, such an assertion becomes pointless and is nothing more than something hurtful to the true author of the discovery.…You have treasures enough of your own, Sir, to have no need to envy anyone;…” The emphasis is all mine. What emotions lay behind that withering, reproachful prose, no one will ever know. Although, one thing was clear: Legendre had pointed out to Gauss in no uncertain terms that one cannot simply claim priority for a discovery by saying that they thought of it first. They must be able to cite where and when they published it. In effect, the seemingly unanswerable question that Legendre also asked the scientific community was this: if Gauss had discovered the method of least squares in the 1790s, what was he doing keeping it a secret for over a decade? Why didn’t he publish it in the 1790s? Gauss’s correspondences with his friends and with the legendary Pierre-Simon Laplace reveal that his response to such questions was, in essence, the following: in Gauss’s opinion, the method of least squares was a rather simple concept that he had discovered while working on another problem. Besides, he thought it had probably already been used by others before him. In other words, it was no big deal, really, and Gauss would mention it in a future publication along with other discoveries (which, incidentally, is exactly what he did in his 1809 book “Theoria motus…”) Unfortunately for Gauss, Legendre took a very poor view of this line of logic, and he made Gauss the target of bitter rancor and name-calling for the next two decades. Laplace, a close contemporary of Legendre’s in Paris, and forever the astute and diplomatic scientist, did his part to douse the controversy by publicly giving both Legendre and Gauss credit for the discovery. But to no avail. Even after Legendre’s death in 1833, Gauss found it difficult to shake himself free of Legendre’s accusations. Whatever Gauss thought of the method of least squares, it was to dominate statistical discourse throughout the 1800s. And the primacy that least squares gave to the mean propelled the mean to superpower-status. Drawbacks of the Mean For all its advantages, the mean also suffers from a few serious drawbacks. Failure to Capture the Distributional Shape of the Data Look at the following set of distributions. Despite being remarkably different in shape, all of them have the same mean! The mean, as a point estimate, has zero ability to capture distributional shape, in other words, the very essence of the data it represents. Thus, it loses much of the information contained in the shape of the data set. From an information-theoretic point of view, the mean is a lossy estimate. This has important implications. For example, suppose your data changes over time. Despite different groups within the data moving in opposite directions, the mean might remain the same thereby projecting a deceptive picture of stability. If your decisions depend on the shape of the distribution, the mean is a poor guide. Incidentally, the median, the mode, the max, and the min suffer from the same drawback, in that they too don’t capture the distributional shape of the data. Unbounded Sensitivity to Outliers The second problem with the mean is that a single rogue outlier can have an outsized influence on the statistic. This sensitivity to outliers emerges directly from the formula for the mean which is a probability-weighted sum. In a sample of size n containing q distinct values x1, x2, …, xq with frequencies n1, n2, …, nq and sample probabilities p_hat1, p_hat2, …, p_hatq respectively, the sample mean is given by the following formula. Now, suppose we extend this sample with an outlier, z. We will assume there is a single outlier. The new mean for the (n + 1)-sized sample containing (q+1) distinct values, i.e., the original set of q distinct values plus the new outlier, is the following. Inclusion of z has shifted the mean by the following quantity. Simplifying, we get the following. If z is greater than the old mean, the shift is positive, and vice-versa. Since z is an outlier, by definition, it lies far from the bulk of the sample, and often far from the old mean. So, the value in brackets in Eq. (12) will be large. Although the actual shift in the mean is the bracketed value scaled down by the new sample size (n + 1), for small samples (small n), the bracketed difference isn’t scaled down enough. So, the effect of an outlier on the old mean will be especially large in small samples. But even for large n, the shift in the mean can be substantial for extreme values of z. In fact, Eq. (12) shows that the mean has unbounded sensitivity to outliers. Even a single sufficiently extreme outlier is enough to thoroughly destabilize the statistic. A Totally Synthetic Statistic The third problem with the mean is that it’s a totally synthetic statistic. The mean is a computed value. The min, max, and mode are all observed values found in the sample. The median yields the observed value in an odd-sized sample, although it might need to be interpolated in an even-sized sample. The mean carries no such obligations and is little affected by such constraints. It is synthetic by design. The “Average Airman” The synthetic nature of the mean, if not promptly called out, can lead to costly consequences, as illustrated by an anthropometric survey of considerable proportions that was carried out by the U.S. Air Force in 1950. The Air Force’s survey team visited 14 airbases in the U.S. and took 132 types of body measurements from more than 4,000 Air Force personnel in different roles and ranks. For each measurement such as height, weight, eye-height, arm length, head circumference and so on, the survey authors published percentile tables reporting the 1st to 99th percentile values, along with mean, median, and standard deviation. The stated goal of the survey was to “provide a basis for the design of clothing, equipment, and other aspects of the flight environment”. Unfortunately, the survey failed to adequately highlight the simple fact that the average value of a measurement such as the “sitting height”, or the “neck circumference” used either by itself or in combination with other measurements cannot form an adequate basis for designing any aspect of clothing, equipment, or operating environment. Fortunately, this omission was promptly corrected when one of the study authors, G. S. Daniels, published a follow-on report highlighting the dangers of designing anything based on the concept of the “average man”. Daniels took pains to point out that the “average man” did not exist. He drove home his point by showing how, for any measurement, if the middle 30% of the measured values are considered a rough indication of the average for that measurement, then in the 4063-strong sample in the original survey, not a single airman possessed all of the following: average stature, average circumferences of chest, torso, hip, neck, waist and thigh, average sleeve and crotch lengths, and average crotch height. And this was just a tiny subset of the 132 measurement types covered in the original survey. To its credit, the Airforce survey did publish percentile scores for each measurement which provide the decision maker with a clear sense for the shape and the range of the measurements. Unfortunately, the Airforce example wasn’t the first, and won’t be the last, to research anthropometric averages. The chimerical allure of the “Average Man” has attracted scientists for centuries. Adolphe Quetelet and the Myth of the “Average Man” No discussion about the science of averages would be complete without mentioning Quetelet’s seminal contributions to the field. In the 1820s, the Belgian scientist and astronomer Adolphe Quetelet embarked upon quite possibly one of the largest systematic studies of averages in modern history. From his extensive labors spanning multiple years emerged (surprise, surprise!) the concept of the “Average Man”. In fact, he built a large collection of “average men” each one based on a different feature such as height or weight or a combination of features. Quetelet used these fictional average beings to draw comparisons between groups. For example, he drew seemingly benign comparisons conscripts from different kingdoms in Europe based on aggregate anthropometric measurements such as the average height, weight, and chest measurements. Then, Quetelet did something truly troubling. He turned his analytical mind toward “moral” characteristics such as the aggregate rates of crime and drunkenness for the different groups. Quetelet noticed that some of these group-specific rates were remarkably stable from year to year and took it as a sign that they reflected persistent group-specific causes and tendencies. He associated such group-specific rates with the “average” individual representing that group. In effect, Quetelet was anointing the average as the theoretically correct value, the value that nature intended for that group, and individual differences were simply random deviations from the average. Quetelet, who was an astronomer by profession, had constructed a false equivalence between noisy astronomical observations, and biological and socioeconomic diversity. A celestial object like a star can be said to have a theoretically true position in a specified coordinate system at a specified time. Multiple noisy observations taken through a telescope can be interpreted as deviations from this true value. But this approach doesn’t automatically transfer to a social group of living beings. When it comes to characteristics such as height, length, weight, intellect, or propensity for crime of living creatures, there is no such thing as a theoretically true, nature-intended value. The average value of such characteristics for a group of people is just that, an average. Any kind of elevation of this statistic to the level of a nature-intended value for a social group, with individual variation being mere deviation from this true value constitutes a gravity-defying intellectual leap of startling proportions. On one hand, Quetelet had gifted future generations with a large and rigorously developed body of influential statistical work on averages. On the other hand, his work proved to be dangerously susceptible to misinterpretation. Without adequate contextualization, Quetelet’s work could be, and sadly was, extended and twisted into pseudoscientific narratives to elevate or condemn an entire group of people based purely on the notion of a fictitious average representative of that group and the average rates of various qualities attributed to that representative. It seems, to Quetelet, such deadly side-effects were never enough of a consideration to begin with. His own words betray the extent to which he was enamored of the deceptively effective representational power wielded by the “average”. At one point, he wrote the following: “If an individual at any given epoch of society possessed all the qualities of the average man, he would represent all that is great, good, or beautiful.” (Anonymous 1835, p. 661; Quetelet 1835, Vol. 2, p. 276; Quetelet 1842, p. 100) The Mean Remains the First Among Equals Despite all their shortcomings, the mean, the median, and the mode have become the indisputable superpowers of statistical science. But even among the three, the mean is numero uno: the first among equals. The mean’s influence on data science runs deep and broad. From measures of dispersion such as the variance, to standardized observations such as the z-score; from moment-based statistics such as the mean and variance themselves, to shape statistics such as the skewness and excess kurtosis; from measures of association such as covariance and Pearson’s correlation, to statistical tests of significance such as the z-test,and economic statistics such as the Gini coefficient, there are literally hundreds of well-known statistical measures that use the mean in their formula. In regression analysis, the mean occupies a place of unparalleled significance. A regression model that is trained to minimize the squared error will generally estimate the conditional mean of the response variable. The most well-known example of such a model is the linear regression model whose fitting procedure minimizes the squared error. The following equation expresses the mean response of a linear regression model as a linear function of the input variables x1, x2, …, xp. In general, statistical models, including neural net models, that are trained to minimize the squared loss end up estimating the conditional mean of the response variable. For such models, Eq. (13) can be generalized as follows. In Eq. (14), the function fθ(.) captures the structure of the model. fθ(.) can range from a simple linear function as shown in Eq. (13) to a frightfully nonlinear and complex representation arising from a multilayer neural net model. The θ subscript denotes all the weights, biases, normalization parameters – everything that qualifies as a tunable or trainable value. From average commute times shown in Google Maps, to product ratings shown on Amazon, from HbA1c levels to average life expectancies that drive insurance premiums, from average standardized test scores that shape education policy, to average hourly earnings and consumer price indexes that inform economic policy, from average deal sizes, average retail order values, and average hotel room rates to the average amount of time audiences spend reading my work before moving on to other things, the list of areas in which the mean makes its presence felt is endless. The mean informs and guides every aspect of human existence. References Anonymous (1835), “On Man, and the Developement of His Faculties, &c.—[Sur l’Homme et le Développement de ses Facultés, &c.], by A. Quetelet,” The Athenæum, August 8, 593–594; August 15, 611–613; August 29, 658–661 [URL] Bernoulli, J. (2005 [1713]), On the Law of Large Numbers: Part Four of Ars Conjectandi, O. Sheynin, (translated to English)., Berlin: NG Verlag. [PDF] Bühler, W. K. (1981), Gauss: A Biographical Study, Berlin: Springer-Verlag, doi:10.1007/978-3-642-49207-5. Daniels, G. S. (1952), The “Average Man”?, WCRD Technical Note 53-7, Wright-Patterson Air Force Base, OH: Wright Air Development Center. [PDF] Gauss, C. F. (1809), Theoria Motus Corporum Coelestium in Sectionibus Conicis Solem Ambientium, Hamburg: F. Perthes and I. H. Besser. Gauss, C. F. (1857), Theory of the Motion of the Heavenly Bodies Moving about the Sun in Conic Sections: A Translation of Gauss’s “Theoria Motus,” with an Appendix, C. H. Davis, trans., Boston: Little, Brown and Company. Hald, A. (2007), A History of Parametric Statistical Inference from Bernoulli to Fisher, 1713–1935, New York: Springer, doi:10.1007/978-0-387-46409-1. Hertzberg, H. T. E., Churchill, E., and Daniels, G. S. (1954), Anthropometry of Flying Personnel—1950, WADC Technical Report 52-321, Wright-Patterson Air Force Base, OH: Aero Medical Laboratory, Wright Air Development Center, U.S. Air Force, and Antioch College, Contract AF 18(600)-30, PB111583, AD0047953. [PDF] Legendre, A.-M. (1805), Nouvelles Méthodes pour la Détermination des Orbites des Comètes, Paris: Firmin Didot, pp. 72–80. Niedersächsische Akademie der Wissenschaften zu Göttingen (n.d.), Complete Correspondence of Carl Friedrich Gauß, online correspondence database Plackett, R. L. (1972), “Studies in the History of Probability and Statistics. XXIX: The Discovery of the Method of Least Squares,” Biometrika, 59, 239–251, doi:10.1093/biomet/59.2.239. [PDF] Polasek, W. (2000), “The Bernoullis and the Origin of Probability Theory: Looking Back after 300 Years,” Resonance, 5, 26–42, doi:10.1007/BF02837935. [PDF] Quetelet, A. (1835), Sur l’homme et le développement de ses facultés, ou Essai de physique sociale, 2 vols., Paris: Bachelier. Quetelet, A. (1842), A Treatise on Man and the Development of His Faculties, English translation supervised by R. Knox and edited by T. Smibert, Edinburgh: William and Robert Chambers. Seneta, E. (2013), “A Tricentenary History of the Law of Large Numbers,” Bernoulli, 19, 1088–1121, doi:10.3150/12-BEJSP12. [PDF] Stigler, S. M. (1986), The History of Statistics: The Measurement of Uncertainty before 1900, Cambridge, MA: Belknap Press of Harvard University Press, doi:10.2307/2982057 Copyrights This article is © Sachin Date and is licensed under CC BY-SA. Unless otherwise indicated in an image caption, all images in this article are © Sachin Date and are licensed under CC BY-SA.

Original Source

Read the full article at Towardsdatascience →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.