ImageNet

ImageNet is a large-scale image database whose creation under the leadership of Fei-Fei Li became one of the key preconditions for the rise of deep learning. It contains more than fourteen million images manually assigned to tens of thousands of categories organised according to the hierarchy of the WordNet lexical database; the annotations were produced by splitting the work among tens of thousands of people on a crowdsourcing platform. The best known part is the ILSVRC competition subset with one thousand classes and roughly 1.2 million training images, on which image classification was measured annually from 2010. The breakthrough came in 2012, when the convolutional network AlexNet cut the error by more than ten percentage points against the best previous methods and showed unambiguously that the decisive factor is the combination of large data, deep networks and computation on graphics accelerators. By 2015 models had surpassed human-level accuracy. ImageNet’s historical significance is twofold: it proved that data matter at least as much as algorithms, and it became the starting point for transfer learning. At the same time it is a textbook example of benchmark problems – contamination, bias and annotation errors.


Before you can examine students, you need a textbook and a common exam. ImageNet was both: fourteen million photographs, each labelled by humans with what it shows, plus an annual public race that revealed whose method was better. Until then everybody boasted about results on their own data and nothing could be compared. And 2012 was the moment when one contestant won not by a whisker but by the length of the whole track – and it was immediately clear to everyone else that they had to take up a different discipline. From there a straight line leads to today’s models.

Is this article useful to you and are you citing it? Copy the citation