Self-hosting an LLM: what an open-weight model on Hetzner really costs — and when the API stays cheaper

Self-hosting a 70B model currently costs 1,197.30 € net per month plus 599 € setup at Hetzner. The same model class is available from German data centers via API from 0.65 € per million tokens: the break-even sits at around 1.84 billion tokens per month. I show you the calculation — and when your own server pays off anyway.
15 min readMatthias RadscheitMatthias Radscheit
Happycodingen-US

TL;DR

Self-hosting a 70B model currently costs 1,197.30 € net per month plus 599 € setup at Hetzner. The same model class is available from German data centers via API from 0.65 € per million tokens: the break-even sits at around 1.84 billion tokens per month. I show you the calculation — and when your own server pays off anyway.

  • On pure arithmetic, the API almost always wins: the break-even of the Hetzner GEX131 against Llama 3.3 70B at IONOS sits at around 1.84 billion tokens per month.
  • Hetzner raised its prices on June 15, 2026: the GEX131 now costs 1,197.30 € net per month plus 599 € setup, no longer the often-quoted 889 €.
  • VRAM decides the model class: 20 GB realistically carries 14B models, the 70B class needs 96 GB — and CPU servers are not an option for team operation.
  • Self-hosting is justified by confidentiality, fixed costs, true sustained load and version fidelity, not by the price per token.
  • Self-hosting does not automatically make you GDPR-compliant: the EU location only settles the third-country transfer of the LLM processing, the rest remains your job.
  • Check the license before going to production: Apache 2.0 (Mistral Small, Qwen3) is condition-free, Llama 3.3 and Gemma come with community terms.

Up front: the short answer for decision-makers

You are considering running an open language model on your own hardware, and you want to know the cost before you commit. Here is the calculation everything comes down to: at Hetzner, a GPU server that carries the 70-billion-parameter class currently costs 1,197.30 € net per month plus 599 € setup (retrieved August 18, 2026). The same model, Llama 3.3 70B, is available from a German data center via API for 0.65 € per million tokens.

For your own server to pay off on pure arithmetic, you need to process around 1.8 billion tokens per month: that is more than two million A4 pages of text. Month after month.

Takeaway: in 2026, self-hosting is not a cost decision, it is a data-sovereignty decision. On pure arithmetic, the API almost always wins — the EU ones included.

Why I wrote this article anyway: most texts on this topic cite either no numbers or wrong ones. I walk you through the full calculation: architecture, hardware requirements, up-to-date Hetzner prices, a fair API baseline and the break-even as a formula. At the end, you can judge for yourself whether you belong to the exceptions for which your own server pays off.

What «self-hosting an LLM» actually means: the architecture

Self-hosting means: you run the inference of a finished model on hardware you control. Four building blocks belong to that.

Open-weight model: the model's weights are freely downloadable, usually via Hugging Face. Examples are Mistral Small 3.2, Qwen3 or Llama 3.3. You train nothing yourself: you download a finished model and run it.

Inference server: software that loads the model into GPU memory and answers requests. For experiments, Ollama is enough. For production with several concurrent users, we rely on vLLM: it batches parallel requests efficiently and exposes an OpenAI-compatible interface.

GPU server: rented, not bought. A dedicated server with enough graphics memory, in this article specifically Hetzner's GEX series.

Integration: your applications talk to the server through the API interface, exactly as they would talk to an external provider. That keeps the switch between API and your own server technically small: usually a changed URL and a different key.

Part of the picture is what «production-ready» means beyond these four building blocks: access control in front of the endpoint, monitoring for GPU utilization and response times, orderly updates for drivers and the inference server. Getting a model running takes an afternoon. Running it reliably for two years is the real task, and that is exactly the difference many tutorials leave out.

One boundary matters: this article covers operations, not training. Whether you need your own model at all, or whether API plus RAG is enough, is settled by the decision tree in our overview on training your own AI model. And whether fine-tuning pays off in your case is covered in the sibling article LLM fine-tuning: when it pays off.

VRAM decides: which model class fits which hardware

The key metric in self-hosting is not the CPU and not the RAM: it is graphics memory (VRAM). As a rule of thumb, a model in full bf16 precision needs around two bytes per parameter. Mistral Small 3.2 with its 24 billion parameters occupies about 55 GB of VRAM according to its model card. Quantization to 4 bits pushes that down to roughly a quarter, but measurably costs quality and limits the usable context.

The two GPU classes at Hetzner: 20 and 96 GB VRAM

20 GB VRAM (Hetzner GEX44): the 14B class runs comfortably here, for example Qwen3-14B with 14.8 billion parameters. A 24B model only fits quantized and with a small context window. For a team that wants to process long documents, it gets tight.

96 GB VRAM (Hetzner GEX131): only this class carries Llama 3.3 70B in quantized form with usable context. Alternatively, you run a 24B model on it in full precision and serve many parallel users.

A detail that comparison tables like to omit: the weights are not the only memory consumer. Every running request occupies additional VRAM for its context (the so-called KV cache), and that grows with context length and user count. A model that arithmetically «just fits» leaves little room for parallel operation. So do not plan to the edge: plan one class larger than the pure weight arithmetic suggests.

Language quality in German and the CPU counter-example

And output quality in German? Mistral Small, Qwen3 and Llama 3.3 are all trained multilingually and usable for everyday tasks in German. But no benchmark gives you a dependable ranking for your use case: test with your own documents and your own prompts before you commit. That is exactly what the hourly rental is for, which we get to in a moment.

0 GB VRAM (CPU server): finally, the counter-example. Strato advertises LLM hosting on pure CPU servers from 16 € per month and quotes «8 to 15 tokens per second» as a guide value — measured, however, on a quantized 7-billion model, one class below everything covered in this article (retrieved August 18, 2026).

And even that figure describes a single, slow response stream. For one person experimenting, that may be enough. For team operation, CPU inference is not an option, and that is exactly why no CPU variant appears in my cost comparison.

What Hetzner really costs (retrieved August 18, 2026)

I checked all prices directly at Hetzner on August 18, 2026, net plus VAT.

GEX44-1: 232.30 € per month plus 114.00 € setup. For that you get an RTX 4000 SFF Ada with 20 GB GDDR6 ECC, an i5-13500, 64 GB RAM and two 1.92 TB NVMe SSDs. Traffic is unlimited, the location is Falkenstein in Germany. Gross, that is 276.44 € monthly plus 135.66 € setup.

GEX131-1: 1,197.30 € per month plus 599.00 € setup. At its heart is an RTX PRO 6000 Blackwell Max-Q with 96 GB GDDR7 ECC, plus a 24-core Xeon and 256 GB RAM. Locations are Falkenstein and Helsinki, depending on availability also Nuremberg — all EU. Gross: 1,424.79 € plus 712.81 € setup.

Two things you should know before reading on elsewhere. First: Hetzner raised its prices on June 15, 2026. The GEX131 launched in December 2025 at 889 € net with no setup fee; anyone calculating with that figure today lands around a quarter too low.

Second: some comparison articles circulate «Hetzner prices» for RTX 4090 or A100 servers. Hetzner does not offer such products at all. Check every number at the source, mine included.

Useful for testing: the GEX131 is also available with hourly billing, at 1.9188 € net per hour according to Hetzner's order API (retrieved August 18, 2026). You can spend a weekend measuring how your use case behaves before committing to a monthly contract. A full test day costs less than 50 €: that is the cheapest insurance against a 1,200 € misjudgment I know of.

What the rent already includes is often overlooked when comparing against purchased hardware: electricity, cooling, connectivity, hardware replacement on failure and the unlimited traffic are all in. A workstation of your own in the server room looks cheaper on paper but carries every one of these items itself.

The honest comparison baseline: the same model from an EU API

Many comparisons fail right here: they pit your own 14B server against a frontier API. That is unfair in both directions, because the models do not play in the same league. The clean baseline is: the same model, once self-hosted and once as an API call.

That is exactly what the IONOS AI Model Hub allows, offering open models from German data centers (prices from August 18, 2026, net): Llama 3.3 70B costs 0.65 € per million tokens there, input and output alike. Mistral Small 24B sits at 0.10 € for input and 0.30 € for output, Qwen 3.5-9B at 0.10 € and 0.15 €.

That is the same model class you would run yourself on the GEX131 or GEX44: same quality, same EU processing, just without your own server.

As an international price reference: DeepInfra, a US provider, lists Llama 3.3 70B at $0.10 input and $0.32 output, Mistral Small 24B at $0.05 and $0.08 (retrieved August 18, 2026). Cheaper still, in other words. Where DeepInfra operates its data centers, however, the provider does not document: I list it not as a GDPR option, only as evidence of where API prices for open models are heading worldwide.

The EU API is the right yardstick for a second reason: it neutralizes part of the privacy argument for self-hosting. Processing in a German data center with a data-processing agreement is available without your own server. Anyone recommending a dedicated GPU server has to argue against this bar, not against a US API.

A note on diligence: API prices for open models move fast, and all figures given here apply to the retrieval date. For your own decision, pull them fresh: the pricing pages are linked below.

Break-even as arithmetic, not gut feeling

The formula is simple: break-even token volume per month = monthly server costs net ÷ API price per million tokens. If your actual volume sits above that, your own server saves money. If it sits below, you pay extra. Let us run both Hetzner options through it.

Two worked examples: GEX131 and GEX44

Example 1, the 70B class: GEX131 at 1,197.30 € against Llama 3.3 70B at IONOS for 0.65 € per million tokens. 1,197.30 ÷ 0.65 ≈ 1,842 million: the break-even sits at around 1.84 billion tokens per month. The 599 € setup and the operating effort are not even counted yet.

Example 2, the mid class: GEX44 at 232.30 € against Mistral Small at IONOS. With a typical mix of three parts input to one part output, the blended price is 0.15 € per million tokens. 232.30 ÷ 0.15 ≈ 1,549 million: break-even at around 1.5 billion tokens per month.

The pattern stands out: the smaller server is five times cheaper, but the API prices of the smaller model class drop by the same measure. The break-even therefore stays at one and a half to two billion tokens in both classes. There is no trick that pushes it down to a reachable volume through clever hardware choice.

Reality check: typical volumes at mid-sized companies

And what does a mid-sized company actually consume? A model calculation: 50 employees each make 30 requests per day to an internal assistant, and each request moves around 2,000 tokens including RAG context. Over 21 working days, that comes to about 63 million tokens per month — a factor of 29 below the GEX131's break-even. As an API bill at IONOS: around 41 € per month. Your own server costs thirty times that.

A batch scenario also reaches the threshold more slowly than you would think. Suppose you tag 5,000 documents of 3,000 tokens each every night: that is 15 million tokens per night, around 450 million per month. Still only a quarter of the 70B class break-even. Only from about 20,000 documents per night, every night, does your own GEX131 become arithmetically cheaper than the API call.

What the formula still leaves out: operations and failover

Then there is operations: updates, monitoring, occasional model swaps and keeping the inference server current cost roughly one working day per month in our project practice at happycoding — an experience value, not a law of nature. Price that day in at internal or external hourly rates and the break-even shifts further in the API's favor.

This calculation even flatters the server. Not priced in is the missing failover: a single server has none, the API does. I deliberately let the result stand as it is: on pure cost, the API wins for practically every mid-sized company. If someone tells you flatly that self-hosting pays off «from 20 to 30 users», ask for the formula behind it.

When self-hosting wins anyway

There are five situations in which your own server is the right decision despite the numbers above.

Confidentiality: your data does not leave your infrastructure. No API provider logs prompts, no additional processor sits between you and the model. For M&A documents, HR data subject to works-council reservations or special categories under Art. 9 GDPR, that is often the deciding point.

Predictable fixed costs: 1,197.30 € per month can be budgeted, a volume-dependent API bill cannot. Some budget processes at mid-sized companies prefer the fixed number, even when it is higher.

True sustained load: if you classify inventory data at night, tag document archives or run ETL pipelines with LLM steps, you do in fact approach the billion-token mark. With a GPU that is busy around the clock, the calculation flips: then the fixed price becomes the advantage.

Contractual obligations: some clients and industries simply mandate on-premises or offline operation. Then the question is settled before anyone calculates.

Model stability: an API provider can swap models, change prices or discontinue products. On your server, exactly the model version you tested and signed off runs for as long as you want. For systems with documented validation runs, for example in regulated processes, that version fidelity is worth more than any price advantage.

Radically honest: these are exceptions, not the rule. In most projects we deliver at happycoding, the EU API is the default and the dedicated GPU server the justified exception. How to embed this decision in a company-wide rollout is described in our guide on adopting AI in your company.

GDPR and third-country transfers: a sober look

Up front: this is a technical assessment, not legal advice. Three points can be cleanly separated, though.

First, the location: the Hetzner locations Falkenstein, Nuremberg and Helsinki are in Germany and Finland, that is, in the EU. For the LLM processing itself, no third-country transfer takes place.

Second, the reach of that statement: it applies only to the LLM processing. Your overall system can still contain third-country transfers, for example through analytics, email services or the CRM. Your own model server heals none of that.

Third, the model download: obtaining the weights from Hugging Face, a US company, is not a transfer of personal data. You are downloading matrices of numbers; no customer data flows in the process.

Between a US API and your own server there is also the middle path shown above: the EU API. It moves processing to a German data center with a data-processing agreement, without you operating hardware. For many mid-sized companies, that is the pragmatic first step, and it answers the third-country-transfer question just as well.

So self-hosting does not automatically make you GDPR-compliant: legal basis, a data-processing agreement with the hoster, possibly a data protection impact assessment and data-subject rights remain your job — just like the transparency obligations under Art. 50 of the AI Act, which have applied since August 2026 regardless of the hosting model.

The practical questions around both are covered in two dedicated articles: the hands-on guide to GDPR-compliant AI on your website and the assessment of the EU AI Act for software buyers.

Licenses: Apache 2.0 is not the same as a community license

«Open source» is printed on many models but means different things. For commercial operation, you should distinguish three license families — and do it before the model goes into production, not after. A license change mid-operation otherwise means: swap the model, retest the prompts, repeat the sign-offs.

Apache 2.0: Mistral Small 3.2 and Qwen3-14B are released under the Apache License 2.0. Commercial use without special conditions: no naming obligations, no usage limits. That is the most straightforward choice for mid-sized companies.

Llama 3.3 Community License: usable commercially, but with conditions (license text). The well-known 700-million-user clause does not affect you as a mid-sized company. More relevant are the mandatory «Built with Llama» notice and the Acceptable Use Policy you contractually bind yourself to.

Gemma Terms of Use: Google's terms (as of April 1, 2026, license text) allow commercial use but tie it to a Prohibited Use Policy and pass-through obligations. Careful: Gemma 4 has separate terms, so always check the version you deploy.

The operational advantage of local weights: once downloaded, they remain usable under the license you obtained them under. Nobody can switch them off retroactively. How real that risk is with hosted services is currently on display at OpenAI: the company is winding down its self-serve fine-tuning completely by early 2027, with no successor product.

Next steps

If you want to make this concrete for your company, I recommend this order. First: estimate your token volume with the formula from this article, honestly and per use case. Second: start with an EU API like the IONOS AI Model Hub — you get EU processing without fixed costs and can switch later, because the interfaces are compatible. Third: rent a GPU server only once proven sustained load or a confidentiality obligation calls for it, and test by the hour first.

This order costs you nothing in options: because model and interface stay the same, the later move from API calls to your own server is an infrastructure change, not a rebuild.

Usually the self-hosted model is only one building block anyway: the real value emerges in the processes around it, from automated inbox handling to document classification. How we build such process automation with n8n and custom pipelines on your infrastructure, set up and maintained by us, is something we gladly walk through on your concrete case. The sibling article on workflow automation with n8n on EU infrastructure gives you a way in as well.

If you are unsure whether an API or your own server fits your case: book a free initial consultation. We will run your scenario through the numbers together in 30 minutes, with your figures instead of my example values.

Frequently asked questions

Is a cheap CPU server enough to get started?
For you alone, to experiment: yes. For team operation: no. Strato advertises LLM hosting on CPU servers from 16 € per month and quotes «8 to 15 tokens per second» — measured on a quantized 7-billion model, smaller than anything this article compares (retrieved August 18, 2026). That is a single, slow response stream with no parallel operation. As soon as several employees are meant to work at the same time, you need a GPU or an API.
From what usage volume does your own server pay off?
Calculate it with the formula: monthly server costs net divided by API price per million tokens. For the Hetzner GEX131 against Llama 3.3 70B at IONOS, the break-even sits at around 1.84 billion tokens per month. Typical internal assistants at mid-sized companies come in a factor of 20 to 30 below that — then you drive cheaper with the EU API.
Is a self-hosted LLM automatically GDPR-compliant?
No. The EU location only settles the third-country-transfer question for the LLM processing itself. Legal basis, a data-processing agreement with the hoster, possibly a data protection impact assessment and data-subject rights remain your job. The transparency obligations under Art. 50 of the AI Act also apply regardless of where your model runs.
How many concurrent users does a single GPU server carry?
Honest answer: that depends on model size, context length and the inference server, and without a load test every number is a guess. With vLLM and batched requests, a GEX131 carries a typical internal assistant for a mid-sized team. Rent the server by the hour (1.9188 € net per hour according to Hetzner's order API, retrieved August 18, 2026) and measure your own case before you commit.
Can I also fine-tune on the rented GPU server?
Technically yes: on the GEX131 with 96 GB VRAM, LoRA trainings of smaller open-weight models are feasible. Whether that pays off for you is a question of its own with its own pitfalls: I answer it in the sibling article «LLM fine-tuning: when it pays off».

Sources

Related articles

Open for select projects

Let's talk about your project

Book a no-obligation call, send us an email, or use the form – we'd love to hear from you.

150+
Completed projects
15
Years of experience
8
Senior‑level team members