Definitions and terminology[edit]

In order to give the definition for something that is PAC-learnable, we first have to introduce some terminology.^[2]

For the following definitions, two examples will be used. The first is the problem of character recognition given an array of $n$ bits encoding a binary-valued image. The other example is the problem of finding an interval that will correctly classify points within the interval as positive and the points outside of the range as negative.

Let $X$ be a set called the instance space or the encoding of all the samples. In the character recognition problem, the instance space is $X=\{0,1\}^{n}$ . In the interval problem the instance space, $X$ , is the set of all bounded intervals in $\mathbb {R}$ , where $\mathbb {R}$ denotes the set of all real numbers.

A concept is a subset $c\subset X$ . One concept is the set of all patterns of bits in $X=\{0,1\}^{n}$ that encode a picture of the letter "P". An example concept from the second example is the set of open intervals, $\{(a,b)\mid 0\leq a\leq \pi /2,\pi \leq b\leq {\sqrt {13}}\}$ , each of which contains only the positive points. A concept class $C$ is a collection of concepts over $X$ . This could be the set of all subsets of the array of bits that are skeletonized 4-connected (width of the font is 1).

Let $EX(c,D)$ be a procedure that draws an example, $x$ , using a probability distribution $D$ and gives the correct label $c(x)$ , that is 1 if $x\in c$ and 0 otherwise.

Now, given $0<\epsilon ,\delta <1$ , assume there is an algorithm $A$ and a polynomial $p$ in $1/\epsilon ,1/\delta$ (and other relevant parameters of the class $C$ ) such that, given a sample of size $p$ drawn according to $EX(c,D)$ , then, with probability of at least $1-\delta$ , $A$ outputs a hypothesis $h\in C$ that has an average error less than or equal to $\epsilon$ on $X$ with the same distribution $D$ . Further if the above statement for algorithm $A$ is true for every concept $c\in C$ and for every distribution $D$ over $X$ , and for all $0<\epsilon ,\delta <1$ then $C$ is (efficiently) PAC learnable (or distribution-free PAC learnable). We can also say that $A$ is a PAC learning algorithm for $C$ .

Occam learning

Data mining

Error tolerance (PAC learning)

Sample complexity

M. Kearns, U. Vazirani. . MIT Press, 1994. A textbook.

An Introduction to Computational Learning Theory

M. Mohri, A. Rostamizadeh, and A. Talwalkar. Foundations of Machine Learning. MIT Press, 2018. Chapter 2 contains a detailed treatment of PAC-learnability.

Readable through open access from the publisher.

D. Haussler. . An introduction to the topic.

Overview of the Probably Approximately Correct (PAC) Learning Framework

L. Valiant. Basic Books, 2013. In which Valiant argues that PAC learning describes how organisms evolve and learn.

Probably Approximately Correct.

Littlestone, N.; Warmuth, M. K. (June 10, 1986). (PDF). Archived from the original (PDF) on 2017-08-09.

"Relating Data Compression and Learnability"

Moran, Shay; Yehudayoff, Amir (2015). "Sample compression schemes for VC classes". :1503.06960 [cs.LG].