GPU SERVERS

A whole NVIDIA GPU.
One flat monthly price.

A dedicated RTX PRO 6000 with 96 GB of VRAM, CUDA pre-installed and no per-hour meter.

✓1:1 passthrough, no slicingNo GPU slicing✓CUDA ready on Ubuntu 24.04CUDA ready✓Europe · US Central
See GPU plans→
GPU Cloud G96 from $1,922/mo · about $2.63/hr · VAT extra
gpu-g96-01NVIDIA RTX PRO 6000
Running

RegionEurope (DE)

StackCUDA · Docker

$ ollama run llama3.3:70b
✓ 70B model loaded · FP8 · 69.8 GB VRAM
NVIDIA RTX PRO graphics card
● Limited stock · 2 regions
96 GBGDDR7 ECC VRAM
24,064CUDA cores
120FP32 TFLOPS
1,597 GB/smemory bandwidth
WHY DEDICATED

Why rent a whole GPU?

A sliced or shared GPU is fine for experiments. When the model is your product, you want the full card, the full memory and a bill you can predict.

The whole card is yours

1:1 PCI passthrough maps the full RTX PRO 6000 to your server. No vGPU slicing and no noisy neighbours.

One flat monthly price

No per-hour meter and no surprise egress bill. Traffic is unmetered under fair use.

CUDA from the first login

Ubuntu 24.04 with NVIDIA drivers, the CUDA toolkit and the container toolkit already installed.

PRICING

GPU plans

Prices exclude VAT. Pay monthly, or lock in a lower rate for 12 or 24 months.

GPU CLOUD · LIMITED STOCK

GPU Cloud G96

NVIDIA RTX PRO 6000 Blackwell Server Edition · 96 GB GDDR7 ECC

18 vCPU96 GB RAM900 GB NVMeUnmetered trafficEurope · US CentralRoot access
$1,922/mo
$23,060 every 12 months · about $2.63 an hourMonthly $2,233 · 24 months $1,805/mo
Deploy G96→
ADA LOVELACE

Dedicated L40S

48 GB VRAM$1,552/mo

32 CPU cores · 234 GB RAM · 1.75 TB · 15 TB traffic · EU data centre

Best for: Inference, graphics and video

Configure
BLACKWELL

Dedicated RTX PRO 5000

48 GB VRAM$1,552/mo

32 CPU cores · 234 GB RAM · 1.75 TB · 15 TB traffic · EU data centre

Best for: Rendering, design and mid-size LLMs

Configure
BLACKWELL

Dedicated RTX PRO 6000

96 GB VRAM$3,010/mo

32 CPU cores · 234 GB RAM · 1.92 TB · 15 TB traffic · EU data centre

Best for: 70B at FP8 and fine-tuning

Configure
HOPPER

Dedicated H200

141 GB VRAM$3,884/mo

32 CPU cores · 234 GB RAM · 1.9 TB · 15 TB traffic · EU data centre

Best for: Long context and training

Configure

Multi-GPU clusters

8× H200 HGX (1,128 GB) from £28,855/mo, or 8× B300 with 400G InfiniBand. Sized with you on a call.

Full specifications

Every GPU server includes root access, DDoS protection and a free firewall.

PLANGPUVRAMCPURAMSTORAGETRAFFICLOCATIONFROM /MO
GPU Cloud G96RTX PRO 600096 GB18 vCPU96 GB900 GB NVMeUnmetered*Europe · US Central£1,434.05
Dedicated L40SL40S48 GB32 cores234 GB1.75 TB15 TBEU£1,158.55
Dedicated RTX PRO 5000RTX PRO 500048 GB32 cores234 GB1.75 TB15 TBEU£1,158.55
Dedicated RTX PRO 6000RTX PRO 600096 GB32 cores234 GB1.92 TB15 TBEU£2,246.05
Dedicated H200H200141 GB32 cores234 GB1.9 TB15 TBEU£2,898.55
Cluster 8× H200 HGX8× H2001,128 GB192 coresOn request7.68 TBOn requestEU£28,855
Cluster 8× B3008× B3002,304 GBOn requestOn request13.2 TB400G IBEUSales
PLANVRAMCPU / RAMFROM
G96 Cloud96 GB18 / 96 GB£1,434.05
L40S48 GB32 / 234 GB£1,158.55
PRO 500048 GB32 / 234 GB£1,158.55
PRO 600096 GB32 / 234 GB£2,246.05
H200141 GB32 / 234 GB£2,898.55
8× H2001,128 GB192 cores£28,855
*Fair-use policy applies. Prices exclude VAT. 12-month rate shown for GPU Cloud, monthly for dedicated.
SIZING

Will my model fit?

Rough VRAM for model weights alone. Leave 10–20% headroom for context and batching.

48 GB · L40S · PRO 500096 GB · RTX PRO 6000 · G96141 GB · H200
8B model at FP16≈16 GB
32B model at FP16≈64 GB
70B model at FP8≈70 GB
120B MoE at 4-bit≈65 GB
70B model at FP16≈140 GB
405B model at FP8≈405 GB · cluster
Fits 48 GBFits 96 GBNeeds H200Needs a cluster
USE CASES

Built for real GPU work.

Self-host your own LLM

Serve Llama, Qwen or Mistral behind your own API. Customer data stays on your server.

Fine-tuning and LoRA

Adapt open models to your documents and tone on a single, predictable node.

Image and video generation

Run diffusion and video models without queueing for shared capacity.

3D and batch rendering

ISV-certified for Maya, Houdini and Cinema 4D. Render overnight at a fixed cost.

Scientific simulation

Single-node CFD, molecular dynamics and FEM workloads with ECC memory.

Encoding and transcoding

Hardware video encoders for streaming, archives and social cut-downs.

SETUP

From checkout to CUDA in minutes.

No driver installs and no provisioning queue. The basics are done before you log in.

01

Choose a region

Europe or US Central, close to your users or your data.

02

We deploy it

Ubuntu 24.04 with NVIDIA drivers, CUDA toolkit and container toolkit.

03

SSH in and run

Pull your model or container and start working straight away.

PRE-INSTALLED AND 1-CLICK
Ubuntu 24.04 LTSNVIDIA driverCUDA toolkitnvidia-container-toolkitDockerOllaman8nOpenClawHermes AgentDokployCoolify
COST

Flat monthly beats the meter at 24/7.

One RTX PRO 6000-class GPU running all month. On-demand cloud GPUs bill every hour, even when you forget to switch them off.

OneNet GPU Cloud G96 · 12 months£1,434.05
OneNet GPU Cloud G96 · monthly£1,666.05
Comparable on-demand capacity (low)≈ £1,943
Comparable on-demand capacity (high)≈ £2,610
Illustrative comparison using the same 45% commercial uplift across the referenced September 2026 baseline. Reconfirm current stock and supplier pricing before order acceptance.
FAQ

GPU questions, answered.

Sizing a workload? An engineer will check it with you before you buy.

Is the GPU dedicated or shared?−

Dedicated. The full card is passed through 1:1 to your server. There is no vGPU slicing and no one else’s jobs on it.

Can a 70B model run on GPU Cloud G96?+

Yes, many 70B deployments fit at FP8 or lower precision. Context length, batching and runtime overhead still need sizing.

Is CUDA pre-installed?+

Yes. Ubuntu 24.04 arrives with the NVIDIA driver, CUDA toolkit and NVIDIA container toolkit installed.

Which regions can I choose?+

GPU Cloud is available in Europe and US Central, subject to live stock.

Can I move my server to another region later?+

Yes, through a planned migration. Our engineers will confirm downtime and data-transfer steps first.

Can I run Windows?+

Linux is the standard GPU image. Ask an engineer about Windows licensing and availability for dedicated configurations.

Put a whole GPU to work this week.

Tell us the model and the workload. We will size it with you and have CUDA running the same day.

See GPU plans→
GPU Servers | OneNet Servers