Generative AI · AI Infrastructure · Private AI · Architecture Strategy
In June 2017 eight researchers published a 15-page paper about translating English into German. It quietly removed the one design choice that stopped AI from using more hardware. Everything that followed, from ChatGPT to 16,000-GPU training clusters to the GPU shortage, comes from that change. This post reads the paper as an infrastructure architect would: what it changed, why it moved AI budgets from research to compute, and what that means for anyone designing a data centre today, and how far the effects reach beyond it.
Most articles about the Transformer explain how attention works. I have done that already in Attention, the Idea That Made Modern AI Work. This one asks a different question. Why did a change to how a model reads a sentence end up deciding what servers companies buy, how much power a rack draws, and why NVIDIA became one of the most valuable companies in the world?
Short answer: the Transformer made AI parallel. Once AI could be parallel, more hardware reliably meant a better model. From that point on, AI strategy became infrastructure strategy.
Before 2017AI that read one word at a time
Before the Transformer, the best language models were recurrent neural networks (RNNs) and their improved version, LSTMs. They read a sentence the way you read a news ticker: word one, then word two, then word three. That design had two problems, and both mattered.
Slow: one word at a time
Each step needs the result of the step before it. Word 50 cannot start until word 49 is finished, so a GPU’s thousands of cores mostly sit idle.
Forgetful: long text fades
Meaning is passed along word by word, like a message in a long queue. By the end of a long paragraph, the beginning has mostly faded.
“This inherently sequential nature precludes parallelization within training examples.”
Vaswani et al., Attention Is All You Need (2017), Section 1
So in 2017 the authors asked a bold question
Hence the title: attention is all you need.
2017What the paper changed, in plain words
In a Transformer, every word in the input looks at every other word at the same time and decides which ones matter to it. This is called self-attention. Nothing waits for anything else. A whole sentence, or a whole document, is processed in one parallel pass.
Table 1 of the paper puts this in one line. A recurrent layer needs a number of sequential steps that grows with sentence length. A self-attention layer needs one. That single row is the reason this post exists.
Step by step: what happens inside
You do not need the maths to follow the idea. Here is what a Transformer does with one short sentence, in six steps:
Reading the paperWhat it actually claimed, and how modest it was
It is worth reading the paper itself (link at the end). Its claims are small and specific. Nowhere does it say “this will change the world”. It says: we built a translation model that is better and cheaper to train.
Attention Is All You Need
Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser, Polosukhin
Google Brain, Google Research, University of Toronto
Look at the first number. One server, eight GPUs, under four days. A competent university lab could reproduce it. Seven years later, models built on exactly this architecture were trained on clusters of sixteen thousand GPUs. Nothing in the architecture had to change to make that jump. That is what “it scales” means.
Why it changed strategyMore hardware finally meant a better model
Parallelism alone would have been a nice speed-up. What turned it into a strategy was a second discovery three years later.
In January 2020 a team at OpenAI published Scaling Laws for Neural Language Models. They trained many Transformers of different sizes and found that model quality improved as a smooth, predictable power law in three things: model size, amount of data and amount of compute. The trend held across more than seven orders of magnitude.
Read that as a CFO would. If you can predict how much better a model gets for each extra unit of compute, then AI progress stops being a research gamble and becomes a capital investment with a forecastable return. Spend ten times more on GPUs, get a known improvement. Boards understand that sentence. Few of them understood “we have a promising new architecture”.
Seven years, measured in hardware
Here is the original paper next to Meta’s Llama 3 405B (July 2024), which uses the same basic architecture and published its training details openly:
Sources: the Transformer paper (Tables 2 and 3, Section 5) and Meta’s “The Llama 3 Herd of Models” (arXiv 2407.21783).
A 1.6-million-fold increase in compute in seven years did not come from cleverer algorithms. It came from buildings, power contracts, GPUs and networks. Here AI strategy and infrastructure strategy became the same conversation.
TimelineFrom a translation paper to the AI industry
- Jun 2017ResearchTransformer paperTranslation model, trained on 8 GPUs.
- 2018ResearchGPT-1 and BERTGPT-1 (June) uses the decoder half to generate text; BERT (October) uses the encoder half to understand it. One pre-trained model, many tasks.
- Jan 2020ResearchScaling lawsQuality is predictable from size, data and compute.
- May 2020ModelGPT-3, 175 billion parametersScale alone produces abilities nobody explicitly trained for.
- Oct 2020ResearchVision TransformerImages cut into patches and treated like words. Same architecture, new data type.
- Mar 2022HardwareNVIDIA H100 with a Transformer EngineA chip designed around one model architecture.
- Nov 2022ProductChatGPT launchesAI becomes a board-level topic overnight.
- Jul 2024ModelLlama 3 405BOpen weights, trained on up to 16,384 H100s, with its infrastructure problems documented in public.
- TodayIndustryTransformers everywhereText, code, images, speech and video. Enterprises now plan GPU platforms, not AI experiments.
Wider impactNot only infrastructure
Infrastructure is where this post focuses, but it is only one part of what changed. Once one architecture could do almost everything, and do it better with more compute, the effects spread well beyond the data centre.
Translation, speech and images each had their own kind of model. Now one architecture handles nearly all of them.
Most companies stopped building models. A few labs train very large ones; everyone else rents them or reuses open weights.
Chat assistants, coding copilots, AI search and agents exist because one pre-trained model can be reused for many tasks.
Value moved from training a model to using one well: prompting, retrieval (RAG), building agents and testing their answers.
Text from across the internet became raw material, which led to copyright disputes and paid licensing deals with publishers.
Access to GPUs and power became a national question, behind US export controls on advanced AI chips and programmes such as the IndiaAI Mission.
Open-weight models such as Llama, Mistral and Qwen let any organisation run a capable model on its own hardware, which made private AI practical.
Each of these deserves its own post. What connects them is the same chain this post follows: a parallel architecture, predictable returns from scale, and therefore a race for compute. Infrastructure is not the whole story, but it is the part every other card depends on.
For infrastructureSix things the Transformer changed in the data centre
Each change below follows from the same root: one parallel architecture that gets better with scale. For every one, here is how things worked before, how they work now, and the evidence.
GPU memory decides the bill
VMware view: a full memory reservation on every VM: no ballooning, no swap. Sizing walkthrough
Network is now part of the computer
VMware view: your vSAN network, promoted to the most critical part of the design
Hardware failure is a daily event
VMware view: HA and admission control, applied to thousands of GPUs at once
Power and cooling set the ceiling
VMware view: check the PDUs before you check the price
Chips are now built for one design
VMware view: like CPUs gaining VT-x once virtualisation became the norm
Long prompts have a price tag
VMware view: size for context length the way you size for database growth
Failure figures: Meta, “The Llama 3 Herd of Models” (arXiv 2407.21783): 419 unexpected interruptions, about 78% hardware, 58.7% GPU-related.
For the enterpriseWhat this means for your AI strategy
Very few enterprises will train a frontier model. What they will do is run them, for chat over internal documents, coding assistants, agents and search. That is inference, and it is where most enterprise AI infrastructure decisions now sit.
WhyModels are expensive to train but cheap to reuse
WhySame model gives the same answers anywhere
WhyMemory, not compute, is usually the constraint
WhyOne architecture serves text, code and images
WhyBigger is better, but with diminishing returns per dollar
WhyGPUs fail, and models change every few months
For most organisations, the first real decision is where models should run. Two questions settle most cases:
Five questions to ask before any AI infrastructure purchase
- Are we training, fine-tuning or only running models? (Almost always only running.)
- Which data may not leave our network, by regulation or contract?
- What is the largest model and longest context we genuinely need?
- How many users will be mid-question at the same moment at peak?
- Do we have power, cooling and people to run GPUs, or should we rent them first?
LimitsWhere the Transformer struggles, and what may come next
Eight years on, Transformers still dominate, but their weak spots are well known. Each one already has a fix in progress, and each fix changes what the hardware needs to look like.
Outlook for a five-year platform
- None of these fixes replaces the core idea. They make attention cheaper, sparser or partly swapped out.
- Plan for Transformer-style workloads as the main design target.
- Large GPU memory and fast GPU-to-GPU links stay the two things worth paying for.
Read it yourselfHow to read the original paper in 30 minutes
You do not need to follow the maths to get value from it. Six stops, about 30 minutes, in this order:
- 1Abstract and Section 1Introduction6 min
One page. Problem and claim.
- 2Figure 1Architecture diagram4 min
Famous diagram. Encoder on the left, decoder on the right.
- 3Section 4 and Table 1Why Self-Attention8 min
Parallelism argument and the cost of the square.
- 4Section 5.2Hardware and Schedule3 min
8 P100 GPUs, 12 hours, 3.5 days. Every infrastructure person should know this paragraph.
- 5Table 2Quality vs training cost5 min
Compared with older models. Here the economic case is made.
- 6Section 7Conclusion4 min
Last paragraph predicts images, audio and video.
Transformers did not make AI smarter by themselves; they made AI scalable. Once intelligence could be bought with compute, AI strategy became a question of GPUs, memory, networks and power, which is exactly where infrastructure architects live.
See how this lands in a real design








DrJha