AI models consist at their core of huge numbers of similar calculation steps on tables of numbers (matrices), and graphics processors (GPUs) can carry out such steps thousands at a time, whereas ordinary processors work through them one after another. That is why language models can practically only be trained and run on GPUs or similar accelerators. As a business you notice the consequences in prices, availability and energy demand.
Why are GPUs better than ordinary processors for AI?
A CPU has a few very flexible cores and quickly works through tasks one after another. A GPU has thousands of simpler cores and carries out many similar calculations in parallel. It was originally developed for graphics, in which millions of pixels are computed at the same time. Neural networks consist largely of matrix multiplications, a similarly parallel kind of calculation.
| Chip | Strength | Typical use |
|---|---|---|
| CPU | Flexible, fast at single tasks | Operating system, office applications, control |
| GPU | Many similar calculations in parallel | Graphics, training and running AI models |
| AI accelerator | Specialised for AI calculations, such as TPUs | Providers' large data centres |
For background and terms, see the overview of graphics processing units. How language models work is explained in What is an LLM?.
Why does graphics memory decide?
For a model to run, its parameters must sit in the GPU's fast memory. As a rough rule of thumb, you need memory of the order of parameter count times bytes per parameter. A model with seven billion parameters needs around 14 gigabytes at 16-bit precision and about a quarter of that in a reduced representation (quantisation). On top comes memory for the request's context. That explains why large models need server cards and smaller models run on a workstation. More on context in Tokens and context windows explained.
How does training differ from operation?
Training a large model often takes weeks to months on very many GPUs and is the biggest cost item for providers. Operation (inference) needs less per request but adds up across many users. For you, operation is what counts: every request occupies computing time, which is why providers bill by tokens. Smaller, more efficient models, caching of recurring requests and batched processing reduce the effort.
Why are chips and data centres a bottleneck?
Modern AI chips and the fast memory that goes with them are made by only a few manufacturers, and building production capacity takes years. Demand has risen at the same time. Data centres also need electricity, cooling and network connections. In its report Energy and AI, the International Energy Agency expects data centre electricity consumption to more than double by 2030. AI is a major driver, but not the only one.
What does this mean for prices in your business?
Prices for AI services depend on demand, hardware costs and the efficiency of models. More efficient models and hardware can lower prices; scarce computing power can support them. This cannot be predicted reliably. So plan with usage limits, watch costs in the first weeks and choose smaller models for simple tasks. Whether your own hardware is an alternative is covered in Running AI models locally: when your own server beats the cloud.
- Tasks sorted by difficulty, simple tasks tried with smaller models
- Monthly cost of typical use measured
- Usage limits and budget alerts set up with the provider
- Dependence on one provider checked, switch prepared
- Own hardware considered only for high, steady use or data protection needs
Conclusion: computing power is the foundation, not your task
As a business you do not have to buy GPUs, but you should understand why computing power shapes prices and availability. Plan with limits, watch costs and keep switching open. If you want to build AI into processes economically, see our AI automation service or describe your project.




