A local AI model beats the cloud when data must not leave the building or when very many similar requests arise; with little or fluctuating use and high quality demands, the cloud is usually cheaper and simpler. This article shows how to weigh the two, what hardware is needed for which model size and what operation involves.
When does a local model pay off?
Local pays off in three cases. First, when data is particularly sensitive and should not go to a model provider, such as client, patient or HR data. Second, when very many similar requests arise, because fixed costs then fall per request. Third, when you want to be independent of a provider's prices, usage limits and changes. We run a local language model on our own hardware for internal tasks and additionally use cloud models for difficult tasks.
| Criterion | Cloud (through the provider) | Local (own server) |
|---|---|---|
| Data control | Data goes to the provider, governed by contract | Data stays in-house |
| Cost profile | Usage-based | Fixed costs for hardware and operation |
| Quality | Often the most capable models | Good for defined tasks, smaller |
| Effort | Low | Set-up, updates, monitoring |
| Availability | Provider's commitments | Down to you |
| Adaptation | Limited | Fine-tuning possible |
How do you calculate the cost and the threshold?
Compare your monthly bill from the cloud provider with the monthly total cost of your own server. That is made up of the purchase (spread over a period of use), electricity and the effort of looking after it. An example with assumed figures: if a workstation costs €3,000 and you use it for 36 months, that is about €83 a month, plus electricity and upkeep. Only when your cloud usage is permanently above that, or data protection tips the balance, does running your own pay off. For the weighing between open and proprietary models, read Open-source vs proprietary AI models.
What hardware do you need?
What matters is the memory in which the model sits. As a rough guide, a model in a fourfold-reduced representation (quantisation) takes up about half a byte per parameter, plus room for the context. Why graphics processors are needed for this is explained in Why does AI need GPUs?.
| Class | Typical equipment | Model size (rough, quantised) | Suitable for |
|---|---|---|---|
| Workstation or mini PC | 16 to 32 GB of working memory or a small GPU | up to about 8 to 14 billion parameters | Summarising, sorting, drafts for a few users |
| Workstation with GPU | One GPU with 16 to 48 GB of graphics memory | about 14 to 32 billion parameters | Questions on your own documents, several users |
| Server | Several GPUs, from about 80 GB of memory | 70 billion parameters and more | Many users, higher quality, higher cost |
Software such as llama.cpp and Ollama makes running open models easier.
What do operation and upkeep involve?
A local server needs a responsible person. That covers security updates, access rights, monitoring, backups, power and cooling and regular checks on whether a new model is better. The GDPR still applies: Art. 32 requires appropriate technical and organisational measures. This is not legal advice.
- Data protection needs and data types named
- Monthly cloud costs at real usage determined
- Total cost of a server over three years calculated
- Model size and required memory set
- Ownership of updates, security and backups clarified
- Quality checked against a cloud model with your own test set
- Intermediate layer planned so that models stay swappable
Conclusion: test first, then invest
Local models are a good option with high data protection requirements or consistently high use, but they are no free ride. Begin small, compare with the cloud and plan for operation. If you want to know what makes sense for your data and tasks, see our AI automation service or describe your project. How to evaluate models is described in How to evaluate a new AI model.




