Run open LLMs on your own dedicated EU GPU server — fully private, no data sharing. Every model is benchmarked on real hardware, so we show you the GPU with the best price-performance and its measured speed.
| Modèle | Contexte | Parallèle t/s | Débit simple (t/s) | À l'heure | Mensuel | Config |
|---|---|---|---|---|---|---|
| Llama 3.3 70B Perfect on sparbox.m2 | 14,745 | 72 | 9 | €0.73/h | €420.00/mo | sparbox.m2 |
| Qwen3 30B Perfect on ranger.s1 More context, higher speed on infinityai.s1 | 29,491 | 170 | 21 | €0.38/h | €215.00/mo | ranger.s1 |
| 131,072 +344% | 185 +9% | 24 | €0.59/h | €335.00/mo | infinityai.s1 | |
| Qwen3.6 27B (AWQ) Perfect on infinityai.s1 | 65,536 | 103 | 14 | €0.59/h | €335.00/mo | infinityai.s1 |
| Qwen3.6 35B Perfect on infinityai.s1 | 95,550 | 111 | 15 | €0.59/h | €335.00/mo | infinityai.s1 |
| DeepSeek R1 0528 Qwen3 8B Perfect on infinityai.s1 | 85,995 | 592 | 75 | €0.59/h | €335.00/mo | infinityai.s1 |
| Gemma 4 31B (QAT) Perfect on infinityai.s1 | 16,384 | 146 | 19 | €0.59/h | €335.00/mo | infinityai.s1 |
| Gemma 4 E2B Perfect on infinityai.s1 | 131,072 | 1,105 | 153 | €0.59/h | €335.00/mo | infinityai.s1 |
| Gemma 4 E4B Perfect on infinityai.s1 | 106,167 | 684 | 93 | €0.59/h | €335.00/mo | infinityai.s1 |
| Granite 4.1 8B Perfect on infinityai.s1 | 77,395 | 530 | 69 | €0.59/h | €335.00/mo | infinityai.s1 |
| Llama 3.1 8B Perfect on ranger.s1 More context, higher speed on sparbox.m2 | 26,541 | 378 | 51 | €0.38/h | €215.00/mo | ranger.s1 |
| 106,167 +300% | 601 +59% | 84 | €0.73/h | €420.00/mo | sparbox.m2 | |
| Phi 4 Perfect on infinityai.s1 | 16,384 | 337 | 44 | €0.59/h | €335.00/mo | infinityai.s1 |
| Phi 4 multimodal instruct Perfect on infinityai.s1 | 117,964 | 947 | 130 | €0.59/h | €335.00/mo | infinityai.s1 |
| Ministral 3 14B Perfect on infinityai.s1 | 21,497 | 605 | 79 | €0.59/h | €335.00/mo | infinityai.s1 |
| Ministral 3 3B Perfect on ranger.s1 More context, higher speed on infinityai.s1 | 19,347 | 1,221 | 165 | €0.38/h | €215.00/mo | ranger.s1 |
| 85,995 +344% | 1,605 +31% | 215 | €0.59/h | €335.00/mo | infinityai.s1 | |
| Mistral 7B Perfect on infinityai.s1 | 32,768 | 622 | 81 | €0.59/h | €335.00/mo | infinityai.s1 |
| GPT oss 20B Perfect on infinityai.s1 | 131,072 | 701 | 196 | €0.59/h | €335.00/mo | infinityai.s1 |
| Ornith 1.0 9B Perfect on ranger.s1 | 47,774 | 155 | 68 | €0.38/h | €215.00/mo | ranger.s1 |
| Qwen3 Coder 30B (AWQ) Perfect on ranger.s1 More context, higher speed on infinityai.s1 | 29,491 | 864 | 169 | €0.38/h | €215.00/mo | ranger.s1 |
| 131,072 +344% | 908 | 149 | €0.59/h | €335.00/mo | infinityai.s1 | |
| Qwen3 VL 32B (AWQ) Perfect on infinityai.s1 | 32,768 | 408 | 55 | €0.59/h | €335.00/mo | infinityai.s1 |
| Qwen2.5 32B Perfect on infinityai.s1 | 32,768 | 419 | 56 | €0.59/h | €335.00/mo | infinityai.s1 |
| Qwen2.5 7B Perfect on ranger.s1 | 32,768 | 393 | 51 | €0.38/h | €215.00/mo | ranger.s1 |
| Qwen3 14B Perfect on infinityai.s1 | 29,491 | 335 | 43 | €0.59/h | €335.00/mo | infinityai.s1 |
| Qwen3 32B Perfect on sparbox.m2 | 21,497 | 291 | 41 | €0.73/h | €420.00/mo | sparbox.m2 |
| Qwen3 4B Perfect on ranger.s1 | 40,960 | 623 | 84 | €0.38/h | €215.00/mo | ranger.s1 |
| Qwen3 4B (2507) Perfect on ranger.s1 More context, higher speed on infinityai.s1 | 65,536 | 622 | 83 | €0.38/h | €215.00/mo | ranger.s1 |
| 131,072 +100% | 868 +40% | 116 | €0.59/h | €335.00/mo | infinityai.s1 | |
| Qwen3 8B Perfect on infinityai.s1 | 40,960 | 569 | 74 | €0.59/h | €335.00/mo | infinityai.s1 |
| Qwen3 Coder 30B (FP8) Perfect on sparbox.m2 | 32,768 | 776 | 148 | €0.73/h | €420.00/mo | sparbox.m2 |
| Qwen3 VL 32B (FP8) Perfect on sparbox.m2 | 13,270 | 288 | 40 | €0.73/h | €420.00/mo | sparbox.m2 |
| Qwen3 VL 4B Perfect on ranger.s1 | 47,774 | 252 | 63 | €0.38/h | €215.00/mo | ranger.s1 |
| Qwen3.5 9B Perfect on ranger.s1 More context, higher speed on sparbox.m2 | 21,497 | 283 | 46 | €0.38/h | €215.00/mo | ranger.s1 |
| 95,550 +344% | 388 +37% | 79 | €0.73/h | €420.00/mo | sparbox.m2 | |
| Qwen3.6 27B (FP8) Perfect on sparbox.m2 | 32,768 | 242 | 46 | €0.73/h | €420.00/mo | sparbox.m2 |
| Gemma 4 31B (FP8) Perfect on sparbox.m2 | 8,192 | 117 | 42 | €0.73/h | €420.00/mo | sparbox.m2 |
| GLM 4.6V Flash Perfect on infinityai.s1 | 106,167 | 482 | 60 | €0.59/h | €335.00/mo | infinityai.s1 |
Context = max usable context window. Parallel/Single = tokens/sec. Pricing reflects the best price-performance GPU per model. Prices are prepaid; you keep full root access.
Trooper.AI gives you a managed GPU server preinstalled with vLLM, ready to serve any open large language model through a fast, OpenAI-compatible API. Instead of spending hours picking a GPU, installing CUDA drivers, compiling vLLM and tuning launch flags, you pick a model above and click Start. Within minutes you get a dedicated EU-hosted GPU server with vLLM already running the exact configuration we benchmarked for that model — context length, parallelism and command-line arguments included.
vLLM is the industry-standard high-throughput inference engine, but getting it production-ready is fiddly: matching driver and CUDA versions, choosing --max-num-batched-tokens, --max-num-seqs and the right context window for your GPU's VRAM, and validating that the model actually loads. Our managed GPU server preinstalled with vLLM removes that work. Every configuration on this page comes straight from automated benchmarks on the real hardware, so the server you deploy behaves exactly like the tested run — no trial and error.
Each server runs on bare-metal GPUs in ISO/IEC 27001-certified, GDPR-compliant German data centers. You get full root SSH access, a persistent machine, and a private endpoint — your prompts and data never leave your server. Because it is a full GPU server (not shared inference), you can also install additional AI software, fine-tune, or run image and audio models alongside your LLM.
For every model we show the GPU with the best price per performance, with honest hourly and monthly pricing. You can review real sample answers per model via Show response quality before deploying, so you know both the speed and the answer quality you are paying for. Pay hourly to experiment or monthly for production — a managed GPU server preinstalled with vLLM that scales with your needs.