Coefficient of variation

In probability theory and statistics, the coefficient of variation (CV), also known as normalized root-mean-square deviation (NRMSD), percent RMS, and relative standard deviation (RSD), is a standardized measure of dispersion of a probability distribution or frequency distribution. It is defined as the ratio of the standard deviation $\sigma$ to the mean $\mu$ (or its absolute value, $|\mu |$ ), and often expressed as a percentage ("%RSD"). The CV or RSD is widely used in analytical chemistry to express the precision and repeatability of an assay. It is also commonly used in fields such as engineering or physics when doing quality assurance studies and ANOVA gauge R&R, by economists and investors in economic models, and in psychology/neuroscience.

Not to be confused with Coefficient of determination.

Definition[edit]

The coefficient of variation (CV) is defined as the ratio of the standard deviation $\sigma$ to the mean $\mu$ , $CV={\frac {\sigma }{\mu }}.$ ^[1]

It shows the extent of variability in relation to the mean of the population. The coefficient of variation should be computed only for data measured on scales that have a meaningful zero (ratio scale) and hence allow relative comparison of two measurements (i.e., division of one measurement by the other). The coefficient of variation may not have any meaning for data on an interval scale.^[2] For example, most temperature scales (e.g., Celsius, Fahrenheit etc.) are interval scales with arbitrary zeros, so the computed coefficient of variation would be different depending on the scale used. On the other hand, Kelvin temperature has a meaningful zero, the complete absence of thermal energy, and thus is a ratio scale. In plain language, it is meaningful to say that 20 Kelvin is twice as hot as 10 Kelvin, but only in this scale with a true absolute zero. While a standard deviation (SD) can be measured in Kelvin, Celsius, or Fahrenheit, the value computed is only applicable to that scale. Only the Kelvin scale can be used to compute a valid coefficient of variability.

Measurements that are log-normally distributed exhibit stationary CV; in contrast, SD varies depending upon the expected value of measurements.

A more robust possibility is the quartile coefficient of dispersion, half the interquartile range ${(Q_{3}-Q_{1})/2}$ divided by the average of the quartiles (the midhinge), ${(Q_{1}+Q_{3})/2}$ .

In most cases, a CV is computed for a single independent variable (e.g., a single factory product) with numerous, repeated measures of a dependent variable (e.g., error in the production process). However, data that are linear or even logarithmically non-linear and include a continuous range for the independent variable with sparse measurements across each value (e.g., scatter-plot) may be amenable to single CV calculation using a maximum-likelihood estimation approach.^[3]

The data set [100, 100, 100] has constant values. Its is 0 and average is 100, giving the coefficient of variation as 0 / 100 = 0

standard deviation

The data set [90, 100, 110] has more variability. Its standard deviation is 10 and its average is 100, giving the coefficient of variation as 10 / 100 = 0.1

The data set [1, 5, 6, 8, 10, 40, 65, 88] has still more variability. Its standard deviation is 32.9 and its average is 27.9, giving a coefficient of variation of 32.9 / 27.9 = 1.18

In the examples below, we will take the values given as randomly chosen from a larger population of values.

In these examples, we will take the values given as the entire population of values.

Comparison to standard deviation[edit]

Advantages[edit]

The coefficient of variation is useful because the standard deviation of data must always be understood in the context of the mean of the data. In contrast, the actual value of the CV is independent of the unit in which the measurement has been taken, so it is a dimensionless number. For comparison between data sets with different units or widely different means, one should use the coefficient of variation instead of the standard deviation.

Anonymity – c_v is independent of the ordering of the list x. This follows from the fact that the variance and mean are independent of the ordering of x.

Scale invariance: c_v(x) = c_v(αx) where α is a real number.

[22]

Population independence – If {x,x} is the list x appended to itself, then c_v({x,x}) = c_v(x). This follows from the fact that the variance and mean both obey this principle.

Pigou–Dalton transfer principle: when wealth is transferred from a wealthier agent i to a poorer agent j (i.e. x_i > x_j) without altering their rank, then c_v decreases and vice versa.

[22]

Examples of misuse[edit]

Comparing coefficients of variation between parameters using relative units can result in differences that may not be real. If we compare the same set of temperatures in Celsius and Fahrenheit (both relative units, where kelvin and Rankine scale are their associated absolute values):

Celsius: [0, 10, 20, 30, 40]

Fahrenheit: [32, 50, 68, 86, 104]

The sample standard deviations are 15.81 and 28.46, respectively. The CV of the first set is 15.81/20 = 79%. For the second set (which are the same temperatures) it is 28.46/68 = 42%.

If, for example, the data sets are temperature readings from two different sensors (a Celsius sensor and a Fahrenheit sensor) and you want to know which sensor is better by picking the one with the least variance, then you will be misled if you use CV. The problem here is that you have divided by a relative value rather than an absolute.

Comparing the same data set, now in absolute units:

Kelvin: [273.15, 283.15, 293.15, 303.15, 313.15]

Rankine: [491.67, 509.67, 527.67, 545.67, 563.67]

The sample standard deviations are still 15.81 and 28.46, respectively, because the standard deviation is not affected by a constant offset. The coefficients of variation, however, are now both equal to 5.39%.

Mathematically speaking, the coefficient of variation is not entirely linear. That is, for a random variable $X$ , the coefficient of variation of $aX+b$ is equal to the coefficient of variation of $X$ only when $b=0$ . In the above example, Celsius can only be converted to Fahrenheit through a linear transformation of the form $ax+b$ with $b\neq 0$ , whereas Kelvins can be converted to Rankines through a transformation of the form $ax$ .

$\sigma ^{2}/\mu ^{2}$

Efficiency

$\mu _{k}/\sigma ^{k}$

Standardized moment

(or relative variance), $\sigma ^{2}/\mu$

Variance-to-mean ratio

$\sigma _{W}^{2}/\mu _{W}$ (windowed VMR)

Fano factor

Standardized moments are similar ratios, ${\mu _{k}}/{\sigma ^{k}}$ where $\mu _{k}$ is the k^th moment about the mean, which are also dimensionless and scale invariant. The variance-to-mean ratio, $\sigma ^{2}/\mu$ , is another similar ratio, but is not dimensionless, and hence not scale invariant. See Normalization (statistics) for further ratios.

In signal processing, particularly image processing, the reciprocal ratio $\mu /\sigma$ (or its square) is referred to as the signal-to-noise ratio in general and signal-to-noise ratio (imaging) in particular.

Other related ratios include:

Information ratio

Omega ratio

Sampling (statistics)

Variance function

: R package to test for significant differences between multiple coefficients of variation

Coefficient of variation

Definition[edit]

standard deviation

Comparison to standard deviation[edit]

Advantages[edit]

[22]

[22]

Examples of misuse[edit]

Efficiency

Standardized moment

Variance-to-mean ratio

Fano factor

Information ratio

Omega ratio

Sampling (statistics)

Variance function

cvequality