GPU rental · on request
Systems with four NVIDIA RTX 3090 cards — 24 GB GDDR6X each — for workloads that need dedicated EU-hosted GPU capacity without the newest silicon. On the way now; reserve a system and an engineer confirms timing.
Specifications
| System | 4× NVIDIA RTX 3090 in one dedicated machine — the budget tier |
|---|---|
| Memory | 24 GB GDDR6X per card |
| Availability | On request — reservable now, not rentable today |
| Tenancy | Dedicated machine, your traffic only |
| Pricing | €0.17 / GPU-hr · billed monthly (launch pricing) |
The honest pitch: proven cards, real capacity, dedicated to you. If your workload needs newer silicon or more memory per card, look at the RTX 5090 systems, the RTX 6000 Pro or the DGX Spark — the last one is rentable today.
Community data
On the same-workload llama.cpp CUDA scoreboard the RTX 3090 decodes at 158.2 t/s with 5,175 t/s prompt processing — about 2.1× an RTX 3060, and still ahead of several newer 16 GB cards. Its 24 GB of VRAM at ~936 GB/s bandwidth is the combination that keeps it interesting: capacity that newer consumer cards at this price point don't match.
Same-workload numbers from the llama.cpp CUDA scoreboard (Llama 2 7B Q4_0, tg128, full GPU offload) — community measurements, not ours. We have not measured this card on our own nodes yet; when it joins the fleet, we publish our own numbers. Full ladder on the GPU catalogue.
Community data
A recent community benchmark ran Qwen3.6-35B-A3B — the same MoE family we serve in production on GB10 — on a single RTX 3090:
| Configuration | Decode |
|---|---|
| Model setup fully in VRAM | ~140 t/s |
| Full long context via partial CPU offload | ~89 t/s |
| Less-optimized offload configuration | ~65 t/s |
The spread is the lesson: on 24 GB, the offload configuration decides the result. Figures above are single-card — rent one card or a multi-GPU system, as your workload needs. Community measurements, not ours — when we measure on our own nodes, we publish it.
Data & privacy
Like every AxForge machine: your model, your traffic, our hardware — prompts never persisted. Only request metadata (token counts, timestamps, status) is kept for billing and operations. Full policy at axforge.ai/privacy.
FAQ
Not yet — the RTX 3090 systems are on the way to the fleet. You can reserve one now; an engineer confirms timing with you before you commit.
It is our budget tier: 24 GB GDDR6X per card is real capacity for smaller models, dev and test nodes, and workloads that do not need the newest silicon — on a dedicated EU-hosted machine.
Four NVIDIA RTX 3090 cards with 24 GB GDDR6X each, rented as one dedicated machine serving only your traffic.
At €0.17 per GPU-hour it is the cheapest way onto dedicated EU hardware we offer. We watch the market and price under it — that rate is 90% of the tracked market median. Compare freely.
No. Your model, your traffic, our hardware — prompts never persisted. Only request metadata is kept for billing and operations — see the privacy policy.