The Transformer: Born Half by Accident, and It Changed Everything

Summarise this article with AI

A transformer is like that friend at a party who can listen to everyone in the room at once, work out what matters, and then tell you a story that makes sense out of all the noise.

Transformers are a revolution in artificial intelligence on the scale of the motor car in transport. Their story began in the spring of 2017 at Google, where a team of eight researchers – including a young intern named Aidan Gomez – was working on a new neural network architecture. Much like Henry Ford and his team developing the Model T, they sensed they were building something groundbreaking, but they had no idea how far-reaching the effect of their work would be.

The original goal of transformers was to improve machine translation. Existing models processed text sequentially, word by word, which was rather like reading a book from beginning to end. And that takes time. The transformer changed the game by analysing the whole text at once – as if someone tore all the pages out of the binding, scattered them across a football pitch, flew over with a drone and grasped the entire book in a single glance.

The key innovation of transformers is the “attention” mechanism. Imagine you are at a noisy party trying to follow several conversations at once. Your brain automatically homes in on the most important parts of each one, ignores the noise and joins up the connections. Just like the friend from the opening line. That is precisely how the attention mechanism works inside a transformer. The capability turned out to be surprisingly effective: when the team asked their model to write Wikipedia articles about “the Transformer”, it generated five entirely fictitious but remarkably plausible entries, including a detailed history of a Japanese punk band of that name which had never existed.

The advantages of transformers are considerable. They are lightning fast compared with traditional models – the difference between flying and travelling by horse and cart. They keep learning and improving as the volume of data grows, much like a person who understands the world better the more they read. The Google team found their model reached a BLEU score of over 26 points translating from English into German, markedly better than previous models, and in a fraction of the training time.

Even though transformers represent a major advance, exactly how they work remains largely a mystery, even to the people who built them. It is as if we had a magic wand that does wonderful things without our knowing quite how. As Ashish Vaswani put it: “I think even talking about ‘understanding’ is something we are not ready for. We have only begun to define what it means to understand these models.”

As development and research into transformers continue, we can expect further breakthroughs. In the words of Noam Shazeer: “the more the models are trained, the better they get, with no end in sight.” Transformers have opened a new era in our relationship with technology and language, much as the internet changed the way we communicate and share information. We are only beginning to discover their full potential, and the future they bring will undoubtedly be fascinating, if perhaps a little frightening. As Jakob Uszkoreit observed: “Once we build these machines, we inevitably lose the ability to conceptualise what is happening, the ability to interpret it. And perhaps there is no other way.”

Source: https://www.newyorker.com/science/annals-of-artificial-intelligence/was-linguistic-ai-created-by-accident

Is this article useful to you and are you citing it? Copy the citation