WebDrift

LOADING DIGITAL SYSTEMS

BLOG · AI NEWS & MODELS

Why does AI need GPUs? Chips, data centres and what it means for prices

AI models consist of matrix maths that GPUs run in parallel. Why chips are scarce, what data centres consume and what that means for prices.

4 min read

By WebDrift RedaktionAuf Deutsch lesen

A single glowing processor chip with heat rising from it in front of blurred rows of server racks in a dark settingAI news & models

AI models consist at their core of huge numbers of similar calculation steps on tables of numbers (matrices), and graphics processors (GPUs) can carry out such steps thousands at a time, whereas ordinary processors work through them one after another. That is why language models can practically only be trained and run on GPUs or similar accelerators. As a business you notice the consequences in prices, availability and energy demand.

Why are GPUs better than ordinary processors for AI?

A CPU has a few very flexible cores and quickly works through tasks one after another. A GPU has thousands of simpler cores and carries out many similar calculations in parallel. It was originally developed for graphics, in which millions of pixels are computed at the same time. Neural networks consist largely of matrix multiplications, a similarly parallel kind of calculation.

ChipStrengthTypical use
CPUFlexible, fast at single tasksOperating system, office applications, control
GPUMany similar calculations in parallelGraphics, training and running AI models
AI acceleratorSpecialised for AI calculations, such as TPUsProviders' large data centres

For background and terms, see the overview of graphics processing units. How language models work is explained in What is an LLM?.

Why does graphics memory decide?

For a model to run, its parameters must sit in the GPU's fast memory. As a rough rule of thumb, you need memory of the order of parameter count times bytes per parameter. A model with seven billion parameters needs around 14 gigabytes at 16-bit precision and about a quarter of that in a reduced representation (quantisation). On top comes memory for the request's context. That explains why large models need server cards and smaller models run on a workstation. More on context in Tokens and context windows explained.

How does training differ from operation?

Training a large model often takes weeks to months on very many GPUs and is the biggest cost item for providers. Operation (inference) needs less per request but adds up across many users. For you, operation is what counts: every request occupies computing time, which is why providers bill by tokens. Smaller, more efficient models, caching of recurring requests and batched processing reduce the effort.

Why are chips and data centres a bottleneck?

Modern AI chips and the fast memory that goes with them are made by only a few manufacturers, and building production capacity takes years. Demand has risen at the same time. Data centres also need electricity, cooling and network connections. In its report Energy and AI, the International Energy Agency expects data centre electricity consumption to more than double by 2030. AI is a major driver, but not the only one.

What does this mean for prices in your business?

Prices for AI services depend on demand, hardware costs and the efficiency of models. More efficient models and hardware can lower prices; scarce computing power can support them. This cannot be predicted reliably. So plan with usage limits, watch costs in the first weeks and choose smaller models for simple tasks. Whether your own hardware is an alternative is covered in Running AI models locally: when your own server beats the cloud.

  • Tasks sorted by difficulty, simple tasks tried with smaller models
  • Monthly cost of typical use measured
  • Usage limits and budget alerts set up with the provider
  • Dependence on one provider checked, switch prepared
  • Own hardware considered only for high, steady use or data protection needs

Conclusion: computing power is the foundation, not your task

As a business you do not have to buy GPUs, but you should understand why computing power shapes prices and availability. Plan with limits, watch costs and keep switching open. If you want to build AI into processes economically, see our AI automation service or describe your project.

Sources

#GPU#AI chips#data centres#energy demand#AI costs#hardware

FREQUENTLY ASKED QUESTIONS

Answered briefly.

01What is the difference between a CPU and a GPU?
A CPU has a few powerful cores and handles tasks one after another very flexibly. A GPU has thousands of simpler cores and performs many similar calculations at the same time.
02Does my business need its own GPUs?
Usually not. For most applications you rent computing power from a provider. Your own hardware pays off with strict data protection requirements or very high, steady use.
03Why are AI chips so expensive and scarce?
Demand has risen sharply, while producing cutting-edge chips and memory is complex and done by only a few manufacturers.
04Will AI services get more expensive as a result?
That cannot be predicted reliably. More efficient models and hardware can lower prices, high demand and scarce computing power can support them. Watch your costs.
05How much electricity does AI use?
The demand of data centres is growing. In its report on energy and AI, the International Energy Agency expects it to more than double by 2030. Individual requests need comparatively little.

ABOUT THE EDITORS

WebDrift Redaktion

WebDrift Redaktion is the team behind WebDrift in Dresden for development, design, AI automation and visibility. We write about what we build every day for small and mid-sized businesses: honest, practical and without invented numbers.