, , ,

Attention Is All You Need: the 2017 paper that changed AI strategy and the IT industry

How the 2017 Transformer paper made AI parallel, turned AI budgets into GPU, memory, network and power decisions, and changed research, business, jobs and the wider IT industry.

Attention Is All You Need: the 2017 paper that changed AI strategy and the IT industry, by Dr. Pranay Jha

Generative AI · AI Infrastructure · Private AI · Architecture Strategy

In June 2017 eight researchers published a 15-page paper about translating English into German. It quietly removed the one design choice that stopped AI from using more hardware. Everything that followed, from ChatGPT to 16,000-GPU training clusters to the GPU shortage, comes from that change. This post reads the paper as an infrastructure architect would: what it changed, why it moved AI budgets from research to compute, and what that means for anyone designing a data centre today, and how far the effects reach beyond it.

Most articles about the Transformer explain how attention works. I have done that already in Attention, the Idea That Made Modern AI Work. This one asks a different question. Why did a change to how a model reads a sentence end up deciding what servers companies buy, how much power a rack draws, and why NVIDIA became one of the most valuable companies in the world?

Short answer: the Transformer made AI parallel. Once AI could be parallel, more hardware reliably meant a better model. From that point on, AI strategy became infrastructure strategy.

Before 2017AI that read one word at a time

Before the Transformer, the best language models were recurrent neural networks (RNNs) and their improved version, LSTMs. They read a sentence the way you read a news ticker: word one, then word two, then word three. That design had two problems, and both mattered.

Problem 1

Slow: one word at a time

Each step needs the result of the step before it. Word 50 cannot start until word 49 is finished, so a GPU’s thousands of cores mostly sit idle.

TheserverisdownGPU cores: one busy, the rest waiting
More hardware did not make it faster, only more expensive
Problem 2

Forgetful: long text fades

Meaning is passed along word by word, like a message in a long queue. By the end of a long paragraph, the beginning has mostly faded.

Serverinrack12isdownby the last word, the first has faded
Long documents and long conversations were out of reach

“This inherently sequential nature precludes parallelization within training examples.”

Vaswani et al., Attention Is All You Need (2017), Section 1
In VMware terms: an RNN is a single-threaded application. Give it a VM with 64 vCPUs and it still uses one. Every infrastructure person has met this application, and knows there is no sizing answer to it, only a rewrite.

So in 2017 the authors asked a bold question

BeforeRNN does the reading; attention is an add-on that lets it look back
2017Remove the step-by-step part. Attention is the whole model

Hence the title: attention is all you need.

2017What the paper changed, in plain words

In a Transformer, every word in the input looks at every other word at the same time and decides which ones matter to it. This is called self-attention. Nothing waits for anything else. A whole sentence, or a whole document, is processed in one parallel pass.

Plain example: an RNN is a game of Chinese whispers. The message passes person to person, slowly, and gets garbled along the way. A Transformer is a meeting where everyone gets the full document at once, reads it simultaneously, and each person underlines the parts relevant to them. Faster, and nothing is lost in transit.
RNN (before 2017) one word at a time The server is down step 1step 2step 3step 4 each step waits for the previous one most GPU cores sit idle Transformer (2017) every word sees every word The server is down one step all words processed in parallel every GPU core has work
Watch it run: the RNN lights up one word at a time; the Transformer does all four in a single beat. Same four words, two designs. On the left the work is a queue; on the right it is a grid that a GPU can fill all at once.

Table 1 of the paper puts this in one line. A recurrent layer needs a number of sequential steps that grows with sentence length. A self-attention layer needs one. That single row is the reason this post exists.

Nothing comes free. That same table shows the price: because every word compares itself with every other word, the work grows with the square of the input length. Double the context, quadruple the attention work. Remember this; it comes back later as a memory and cost problem.

Step by step: what happens inside

You do not need the maths to follow the idea. Here is what a Transformer does with one short sentence, in six steps:

How a Transformer reads a sentence same basic steps in the 2017 paper and in today’s chat models 1 Text goes in “The server is down” 2 Split into tokens [The] [server] [is] [down], roughly one per word 3 Each token becomes numbers a long list of numbers that captures its meaning 4 Attention: every token looks at every other “down” learns it is about “server”, all at once 5 Repeat in many layers 6 layers in the 2017 paper, over 100 in today’s largest 6 Output: a translation or the next word e.g. “Der Server ist ausgefallen” in German
Step 4 is the paper’s big idea. Steps 4 and 5 are also where nearly all the GPU work happens.

Reading the paperWhat it actually claimed, and how modest it was

It is worth reading the paper itself (link at the end). Its claims are small and specific. Nowhere does it say “this will change the world”. It says: we built a translation model that is better and cheaper to train.

arXiv 1706.0376212 June 2017NeurIPS 2017

Attention Is All You Need

Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser, Polosukhin
Google Brain, Google Research, University of Toronto

8NVIDIA P100 GPUs in one machine
3.5 daysto train the big model (base model: 12 hours)
213Mparameters in the big model (base: 65M)
28.4BLEU English to German, 2+ points above the previous best
TaskMachine translation, English to German and English to French (WMT 2014). 41.8 BLEU on French.
Training compute2.3 × 1019 FLOPs for the big model
Future workExtend it to “images, audio and video”

Look at the first number. One server, eight GPUs, under four days. A competent university lab could reproduce it. Seven years later, models built on exactly this architecture were trained on clusters of sixteen thousand GPUs. Nothing in the architecture had to change to make that jump. That is what “it scales” means.

What is BLEU? A score from 0 to 100 that measures how closely a machine translation matches human translations. Two points sounds small. In 2017 translation research, two points was a generation’s worth of progress.

Why it changed strategyMore hardware finally meant a better model

Parallelism alone would have been a nice speed-up. What turned it into a strategy was a second discovery three years later.

In January 2020 a team at OpenAI published Scaling Laws for Neural Language Models. They trained many Transformers of different sizes and found that model quality improved as a smooth, predictable power law in three things: model size, amount of data and amount of compute. The trend held across more than seven orders of magnitude.

Model size+Data+Compute=Predictable quality

Read that as a CFO would. If you can predict how much better a model gets for each extra unit of compute, then AI progress stops being a research gamble and becomes a capital investment with a forecastable return. Spend ten times more on GPUs, get a known improvement. Boards understand that sentence. Few of them understood “we have a promising new architecture”.

In VMware terms: this is the difference between an application that scales linearly with hosts and one that does not. Once you know a workload scales out cleanly, the capacity plan writes itself: more hosts, more throughput, in a straight line. The Transformer gave AI that property, and scaling laws drew the straight line on a chart.
How a design change became a GPU race each step made the next one possible 1 2017: work can run in parallel a model can now use thousands of GPU cores at once 2 2020: more compute means a better model scaling laws make the gain predictable 3 Money follows the forecast budgets shift from research teams to GPUs and buildings 4 Bottlenecks move to the data centre GPU memory, GPU-to-GPU network, power, cooling 5 AI strategy becomes infrastructure strategy whoever can build and run the compute sets the pace
Five links in one chain. Break any of them and the GPU race would not have happened.

Seven years, measured in hardware

Here is the original paper next to Meta’s Llama 3 405B (July 2024), which uses the same basic architecture and published its training details openly:

2017 · Transformer big2024 · Llama 3 405B
Parameters213 million≈ 1,900×405 billion
GPUs8 × P100≈ 2,000×up to 16,384 × H100 (80 GB)
Training compute2.3 × 1019 FLOPs≈ 1.6 million×3.8 × 1025 FLOPs
Training dataabout 4.5 million sentence pairsfar larger15.6 trillion tokens
Where it ranone servernew scalea purpose-built data centre cluster

Sources: the Transformer paper (Tables 2 and 3, Section 5) and Meta’s “The Llama 3 Herd of Models” (arXiv 2407.21783).

Seven years of scale, drawn to scale each small square = one 8-GPU server, like the one used in the paper 2017 1 server 8 × P100 GPUs 3.5 days 2024: Llama 3 405B 2,048 servers’ worth: up to 16,384 × H100 GPUs total compute used: about 1.6 million times more
Find the single square on the left, then look at the grid. Same architecture; every extra square is power, cooling, network and failures to manage.

A 1.6-million-fold increase in compute in seven years did not come from cleverer algorithms. It came from buildings, power contracts, GPUs and networks. Here AI strategy and infrastructure strategy became the same conversation.

TimelineFrom a translation paper to the AI industry

  1. Jun 2017
    ResearchTransformer paperTranslation model, trained on 8 GPUs.
  2. 2018
    ResearchGPT-1 and BERTGPT-1 (June) uses the decoder half to generate text; BERT (October) uses the encoder half to understand it. One pre-trained model, many tasks.
  3. Jan 2020
    ResearchScaling lawsQuality is predictable from size, data and compute.
  4. May 2020
    ModelGPT-3, 175 billion parametersScale alone produces abilities nobody explicitly trained for.
  5. Oct 2020
    ResearchVision TransformerImages cut into patches and treated like words. Same architecture, new data type.
  6. Mar 2022
    HardwareNVIDIA H100 with a Transformer EngineA chip designed around one model architecture.
  7. Nov 2022
    ProductChatGPT launchesAI becomes a board-level topic overnight.
  8. Jul 2024
    ModelLlama 3 405BOpen weights, trained on up to 16,384 H100s, with its infrastructure problems documented in public.
  9. Today
    IndustryTransformers everywhereText, code, images, speech and video. Enterprises now plan GPU platforms, not AI experiments.

Wider impactNot only infrastructure

Infrastructure is where this post focuses, but it is only one part of what changed. Once one architecture could do almost everything, and do it better with more compute, the effects spread well beyond the data centre.

Transformer 2017 Infrastructure Research Business model Products Skills and jobs Data Geopolitics Open source
One architecture, eight areas changed. This post goes deep on the dark one; the cards below cover the rest.
01Research

Translation, speech and images each had their own kind of model. Now one architecture handles nearly all of them.

02Business model

Most companies stopped building models. A few labs train very large ones; everyone else rents them or reuses open weights.

03Products

Chat assistants, coding copilots, AI search and agents exist because one pre-trained model can be reused for many tasks.

04Skills and jobs

Value moved from training a model to using one well: prompting, retrieval (RAG), building agents and testing their answers.

05Data

Text from across the internet became raw material, which led to copyright disputes and paid licensing deals with publishers.

06Geopolitics

Access to GPUs and power became a national question, behind US export controls on advanced AI chips and programmes such as the IndiaAI Mission.

07Open source

Open-weight models such as Llama, Mistral and Qwen let any organisation run a capable model on its own hardware, which made private AI practical.

Each of these deserves its own post. What connects them is the same chain this post follows: a parallel architecture, predictable returns from scale, and therefore a race for compute. Infrastructure is not the whole story, but it is the part every other card depends on.

For infrastructureSix things the Transformer changed in the data centre

Each change below follows from the same root: one parallel architecture that gets better with scale. For every one, here is how things worked before, how they work now, and the evidence.

Memory01

GPU memory decides the bill

BeforeCPU and RAM sized with generous overcommitNowWhole model plus every live conversation must sit in GPU memory at once
≈ 140 GB for a 70B model, before the first user connects

VMware view: a full memory reservation on every VM: no ballooning, no swap. Sizing walkthrough

Network02

Network is now part of the computer

BeforeNetwork connects servers to users and storageNowGPUs swap results constantly: NVLink inside a server, InfiniBand or RoCE between servers
One slow link holds back every GPU in the job

VMware view: your vSAN network, promoted to the most critical part of the design

Resilience03

Hardware failure is a daily event

BeforeA failed component is an incident ticketNowCheckpoints, automatic restart and spare nodes are built into the plan
419 unexpected interruptions in 54 days of Llama 3 training

VMware view: HA and admission control, applied to thousands of GPUs at once

Power04

Power and cooling set the ceiling

BeforeRack power was rarely the first questionNowLiquid cooling, denser racks and power contracts come before the purchase order
One 8-GPU server can draw what a whole rack of ordinary servers did

VMware view: check the PDUs before you check the price

Silicon05

Chips are now built for one design

BeforeGeneral-purpose processors for every workloadNowNVIDIA Transformer Engine (FP8), Google TPU, AWS Trainium, Groq LPU
Hardware followed the paper, not the other way round

VMware view: like CPUs gaining VT-x once virtualisation became the norm

Context06

Long prompts have a price tag

BeforeMore input meant a little more workNowDouble the context and the attention work goes up four times
“Read the whole contract” is a GPU sizing request

VMware view: size for context length the way you size for database growth

Failure figures: Meta, “The Llama 3 Herd of Models” (arXiv 2407.21783): 419 unexpected interruptions, about 78% hardware, 58.7% GPU-related.

For the enterpriseWhat this means for your AI strategy

Very few enterprises will train a frontier model. What they will do is run them, for chat over internal documents, coding assistants, agents and search. That is inference, and it is where most enterprise AI infrastructure decisions now sit.

Decision 01Build or rent

WhyModels are expensive to train but cheap to reuse

Do thisUse existing models; spend on data and integration, not training
Decision 02Cloud API or own GPUs

WhySame model gives the same answers anywhere

Do thisDecide on data residency, volume and control, not on quality
Decision 03Sizing

WhyMemory, not compute, is usually the constraint

Do thisSize from model size, precision, context and concurrent users
Decision 04Platform

WhyOne architecture serves text, code and images

Do thisBuild one shared GPU platform, not one stack per use case
Decision 05Model choice

WhyBigger is better, but with diminishing returns per dollar

Do thisUse the smallest model that meets the quality bar; RAG lowers that bar
Decision 06Operations

WhyGPUs fail, and models change every few months

Do thisTreat models like VM images: versioned, tested, swappable

For most organisations, the first real decision is where models should run. Two questions settle most cases:

Cloud API or your own GPUs? a quick first answer for running (not training) models Must the data stay inside your network? Yes Private AI on your own GPUs No Heavy use all day, or must it work offline? Yes Compare: own GPUs vs API cost No Start with a cloud API, pay only for what you use Training your own large model? Very few organisations ever need to.
Same model, same answers either way. The choice is about data, volume and control, not quality.

Five questions to ask before any AI infrastructure purchase

  1. Are we training, fine-tuning or only running models? (Almost always only running.)
  2. Which data may not leave our network, by regulation or contract?
  3. What is the largest model and longest context we genuinely need?
  4. How many users will be mid-question at the same moment at peak?
  5. Do we have power, cooling and people to run GPUs, or should we rent them first?
In VMware terms: this is the same journey as server virtualisation in 2005. First it was a lab curiosity, then a cost-saving project, then the default platform everything ran on. Private AI platforms such as VMware Private AI Foundation with NVIDIA, Red Hat OpenShift AI and Nutanix Enterprise AI are the vSphere moment for Transformers: a shared, governed platform instead of one GPU box per team.

LimitsWhere the Transformer struggles, and what may come next

Eight years on, Transformers still dominate, but their weak spots are well known. Each one already has a fix in progress, and each fix changes what the hardware needs to look like.

Limit 01Attention cost grows with the square of context length
Fix In use todayFlashAttention (2022) and similar methods reorganise the calculation so it moves far less data in and out of GPU memory
Infrastructure effectLonger context on the same GPU
Limit 02Every parameter works on every word
Fix In use todayMixture-of-Experts models, for example Mixtral (December 2023), switch on only a few “expert” sub-networks per word
Infrastructure effectMore total memory, less compute per answer
Limit 03Memory for long conversations keeps growing
Fix Early stageState-space models such as Mamba (December 2023) keep a fixed-size memory instead
Infrastructure effectPossible relief for long context; mostly hybrids so far
Limit 04Serving cost at scale
Fix Standard practiceQuantization, batching engines like vLLM, speculative decoding
Infrastructure effectMore users per GPU

Outlook for a five-year platform

  • None of these fixes replaces the core idea. They make attention cheaper, sparser or partly swapped out.
  • Plan for Transformer-style workloads as the main design target.
  • Large GPU memory and fast GPU-to-GPU links stay the two things worth paying for.

Read it yourselfHow to read the original paper in 30 minutes

You do not need to follow the maths to get value from it. Six stops, about 30 minutes, in this order:

  1. 1
    Abstract and Section 1Introduction

    One page. Problem and claim.

    6 min
  2. 2
    Figure 1Architecture diagram

    Famous diagram. Encoder on the left, decoder on the right.

    4 min
  3. 3
    Section 4 and Table 1Why Self-Attention

    Parallelism argument and the cost of the square.

    8 min
  4. 4
    Section 5.2Hardware and Schedule

    8 P100 GPUs, 12 hours, 3.5 days. Every infrastructure person should know this paragraph.

    3 min
  5. 5
    Table 2Quality vs training cost

    Compared with older models. Here the economic case is made.

    5 min
  6. 6
    Section 7Conclusion

    Last paragraph predicts images, audio and video.

    4 min
Free on arXivAttention Is All You NeedRead the paper (1706.03762)
Takeaway in one sentence

Transformers did not make AI smarter by themselves; they made AI scalable. Once intelligence could be bought with compute, AI strategy became a question of GPUs, memory, networks and power, which is exactly where infrastructure architects live.

About The Author


Discover more from Journal of Intelligent Infrastructure

Subscribe to get the latest posts sent to your email.

Leave a Reply

Your email address will not be published. Required fields are marked *

Architect’s Toolkit

About the Author

Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.

Discover more from Journal of Intelligent Infrastructure

Subscribe now to keep reading and get access to the full archive.

Continue reading