WebDrift

LOADING DIGITAL SYSTEMS

BLOG · AI NEWS & MODELS

Running AI models locally: when your own server beats the cloud

Running AI models on your own server: when it improves privacy and cost, what hardware you need and what operation and upkeep involve for a small company.

4 min read

By WebDrift RedaktionAuf Deutsch lesen

A compact server box glowing under a desk with a small planet hologram above it in a dark officeAI news & models

A local AI model beats the cloud when data must not leave the building or when very many similar requests arise; with little or fluctuating use and high quality demands, the cloud is usually cheaper and simpler. This article shows how to weigh the two, what hardware is needed for which model size and what operation involves.

When does a local model pay off?

Local pays off in three cases. First, when data is particularly sensitive and should not go to a model provider, such as client, patient or HR data. Second, when very many similar requests arise, because fixed costs then fall per request. Third, when you want to be independent of a provider's prices, usage limits and changes. We run a local language model on our own hardware for internal tasks and additionally use cloud models for difficult tasks.

CriterionCloud (through the provider)Local (own server)
Data controlData goes to the provider, governed by contractData stays in-house
Cost profileUsage-basedFixed costs for hardware and operation
QualityOften the most capable modelsGood for defined tasks, smaller
EffortLowSet-up, updates, monitoring
AvailabilityProvider's commitmentsDown to you
AdaptationLimitedFine-tuning possible

How do you calculate the cost and the threshold?

Compare your monthly bill from the cloud provider with the monthly total cost of your own server. That is made up of the purchase (spread over a period of use), electricity and the effort of looking after it. An example with assumed figures: if a workstation costs €3,000 and you use it for 36 months, that is about €83 a month, plus electricity and upkeep. Only when your cloud usage is permanently above that, or data protection tips the balance, does running your own pay off. For the weighing between open and proprietary models, read Open-source vs proprietary AI models.

What hardware do you need?

What matters is the memory in which the model sits. As a rough guide, a model in a fourfold-reduced representation (quantisation) takes up about half a byte per parameter, plus room for the context. Why graphics processors are needed for this is explained in Why does AI need GPUs?.

ClassTypical equipmentModel size (rough, quantised)Suitable for
Workstation or mini PC16 to 32 GB of working memory or a small GPUup to about 8 to 14 billion parametersSummarising, sorting, drafts for a few users
Workstation with GPUOne GPU with 16 to 48 GB of graphics memoryabout 14 to 32 billion parametersQuestions on your own documents, several users
ServerSeveral GPUs, from about 80 GB of memory70 billion parameters and moreMany users, higher quality, higher cost

Software such as llama.cpp and Ollama makes running open models easier.

What do operation and upkeep involve?

A local server needs a responsible person. That covers security updates, access rights, monitoring, backups, power and cooling and regular checks on whether a new model is better. The GDPR still applies: Art. 32 requires appropriate technical and organisational measures. This is not legal advice.

  • Data protection needs and data types named
  • Monthly cloud costs at real usage determined
  • Total cost of a server over three years calculated
  • Model size and required memory set
  • Ownership of updates, security and backups clarified
  • Quality checked against a cloud model with your own test set
  • Intermediate layer planned so that models stay swappable

Conclusion: test first, then invest

Local models are a good option with high data protection requirements or consistently high use, but they are no free ride. Begin small, compare with the cloud and plan for operation. If you want to know what makes sense for your data and tasks, see our AI automation service or describe your project. How to evaluate models is described in How to evaluate a new AI model.

Sources

#local AI#own server#open-weight models#data protection#hardware#operation

FREQUENTLY ASKED QUESTIONS

Answered briefly.

01Are local AI models automatically GDPR-compliant?
No. They avoid passing data to a model provider, but the GDPR still applies. You carry security, access rights, logs and deletion periods yourself.
02How good are local models compared with the cloud?
For clearly defined tasks such as summarising, sorting and drafting, often good enough. For demanding reasoning and very long texts, the largest cloud models are often ahead. Test with your own tasks.
03What is the minimum hardware I need?
Graphics or working memory is decisive. Small models run on a workstation with 16 to 32 gigabytes; larger ones need a workstation GPU or a server.
04Who looks after the solution?
You need someone who takes care of updates, security, monitoring and backups, in-house or through a contractor. Without care, a local server quickly becomes a risk.
05Can I combine local and cloud?
Yes, and it is often sensible. Sensitive routine tasks run locally, demanding uncritical tasks in the cloud. An intermediate layer makes switching easy.

ABOUT THE EDITORS

WebDrift Redaktion

WebDrift Redaktion is the team behind WebDrift in Dresden for development, design, AI automation and visibility. We write about what we build every day for small and mid-sized businesses: honest, practical and without invented numbers.