Pre-trained model

Artificial intelligence is like a student. A pre-trained model is a university graduate with years of general study behind him and a broad grounding – he can read, write and grasp basic connections. He is not a specialist, but he has firm foundations. When a company hires him (which is like using the model for a specific task), it does not have to teach him everything from scratch. A short induction for the specific job is enough, sorting customer e-mails, for instance. This ready-made “graduate” thus picks up a new skill quickly, because he does not have to go through basic education all over again.

A pre-trained model is a neural network that has already been trained on a large, typically general-purpose dataset. During that training the model acquired the ability to extract relevant features and representations from data – recognising edges and textures in images, say, or grammatical structures in text. It serves as a starting point for further specialised tasks, which markedly reduces the computing time and volume of data needed to train a new model. To adapt it to a specific task, techniques such as fine-tuning or feature extraction are then applied.

Is this article useful to you and are you citing it? Copy the citation