Ada Lovelace · PCIe card

# Rent a dedicated RTX 4090 GPU server.

24 GB of GDDR6X at 1,008 GB/s, in a machine that is yours alone. Monthly, paid in crypto, opened without an identity check and delivered in under 5 minutes.

- $193/mo one card, everything included
- 24 GB GDDR6X
- 1,008 GB/s memory bandwidth
- 1×, 2×, 4×, 8× node sizes available

## Price, by node size and term

One number per line, and it is the whole number: no setup fee, no reinstall fee, no bandwidth charge, no support plan. Longer terms are paid up front.

**Monthly price of an RTX 4090 server, by node size and term length**

| Node | CPU and memory | Monthly | 3 months | 12 months |  |
|---|---|---|---|---|---|
| 1 × RTX 4090 24 GB total · PCIe 5.0 | AMD Ryzen 9 7950X16 c / 32 t · 128 GB · 2 × 2 TB NVMe | $193/mo cancel any time | $183/mo save $30/yr | $170/mo save $276/yr | [Configure](https://gpuserver.io/configure?c=rtx4090-x1) |
| 2 × RTX 4090 48 GB total · PCIe 5.0 ×16 | AMD Threadripper 7960X24 c / 48 t · 256 GB · 2 × 3.84 TB NVMe | $416/mo cancel any time | $395/mo save $63/yr | $366/mo save $600/yr | [Configure](https://gpuserver.io/configure?c=rtx4090-x2) |
| 4 × RTX 4090 96 GB total · PCIe 5.0 ×16 | AMD Threadripper 7970X32 c / 64 t · 512 GB · 4 × 3.84 TB NVMe | $814/mo cancel any time | $773/mo save $123/yr | $716/mo save $1,176/yr | [Configure](https://gpuserver.io/configure?c=rtx4090-x4) |
| 8 × RTX 4090 192 GB total · PCIe 5.0 ×16 | 2 × Intel Xeon Gold 6438Y+64 c · 1 TB · 8 × 3.84 TB NVMe | $1,579/mo cancel any time | $1,500/mo save $237/yr | $1,390/mo save $2,268/yr | [Configure](https://gpuserver.io/configure?c=rtx4090-x8) |

Every configuration above is available in all 6 data centres, racked and burned in, with an image ready to write — which is why delivery is under 5 minutes rather than a working day.

## Full specification

GPU NVIDIA RTX 4090 Ada Lovelace architecture, PCIe add-in card.

Memory 24 GB GDDR6X 1,008 GB/s of bandwidth — the number that predicts generation speed, far better than any teraflop figure.

Interconnect PCIe 5.0 ×16 Peer-to-peer over PCIe. Fine for one model per card and for pipeline parallelism; slower than NVLink for tensor parallelism. [Which one your job needs](https://gpuserver.io/guides/nvlink-vs-pcie).

Node sizes 1 × card, 2 × card, 4 × card, 8 × card PCIe cards go up to eight in a standard chassis, and ten only on low-power cards.

Storage 2 × 2 TB NVMe Handed over raw beyond the system volume — we impose no RAID level. More NVMe and archive HDD are monthly options.

Network 1 Gbit/s, unmetered Both directions, no included volume and no overage at any traffic level. One IPv4, a routed IPv6 /64, DDoS filtering on from delivery.

## What runs on an RTX 4090

Every model below fits on a *single* card at 8k context — no sharding, no interconnect to think about. The precision shown is the best one that fits; the arithmetic is in the [VRAM sizing guide](https://gpuserver.io/guides/vram-sizing), so you can check it.

**Language models that fit on a single RTX 4090 at 8k context**

| Model | Best precision that fits | Memory needed | Estimated tokens/s |
|---|---|---|---|
| Llama 3.1 8B 8B parameters | BF16 / FP16Reference quality | 20 GBof 24 GB | ~28 |
| Qwen 3 32B 32B parameters | 4-bit (AWQ, GPTQ)Smallest weights, cache stays FP16 | 21 GBof 24 GB | ~28 |
| Mistral Small 24B 24B parameters | 4-bit (AWQ, GPTQ)Smallest weights, cache stays FP16 | 15 GBof 24 GB | ~38 |
| Gemma 3 27B 27B parameters | 4-bit (AWQ, GPTQ)Smallest weights, cache stays FP16 | 20 GBof 24 GB | ~34 |
| Phi-4 14B 14B parameters | FP8Near-reference, Hopper and Blackwell | 17 GBof 24 GB | ~32 |

On the 8-card node the same models are held across 192 GB in total, which changes what is possible entirely — up to Qwen 3 235B-A22B (MoE). Sharding costs synchronisation at every layer, so a model that fits on one card should stay on one card.

## Throughput to expect

Generation is bounded by memory bandwidth: each token requires re-reading the active weights. At 1,008 GB/s, that sets a ceiling no amount of tuning gets past.

**Our figures are a floor, not a promise.** The estimates on this page are calibrated against published measurements and are deliberately conservative — they model single-stream generation. With continuous batching under vLLM the *aggregate* across concurrent requests is several times higher. If a number here looks low against a benchmark you have seen, that is usually the difference.

## Compared with nearby cards

The three cards closest to this one in price. Memory decides what loads; bandwidth decides how fast it answers.

**The RTX 4090 compared with the three closest cards by price**

| Card | Memory | Bandwidth | Interconnect | From |
|---|---|---|---|---|
| RTX 4090 This page Ada Lovelace | 24 GB | 1,008 GB/s | PCIe | $193/mo |
| [L4](https://gpuserver.io/gpu/nvidia-l4) Ada Lovelace | 24 GB | 300 GB/s | PCIe | $125/mo |
| [RTX A6000](https://gpuserver.io/gpu/rtx-a6000) Ampere | 48 GB | 768 GB/s | PCIe | $286/mo |
| [RTX 5090](https://gpuserver.io/gpu/rtx-5090) Blackwell | 32 GB | 1,792 GB/s | PCIe | $335/mo |

[See all 12 cards side by side](https://gpuserver.io/gpu), or let the [configurator](https://gpuserver.io/configure) pick from your model and context length instead of from a price.

## Compared with other cards

The arbitrations people actually make against an RTX 4090, each worked out on price, memory, bandwidth and the models that fit.

- [RTX 4090 vs RTX 5090  The consumer flagship, one generation apart See the comparison](https://gpuserver.io/compare/rtx-5090-vs-rtx-4090)
- [RTX 4090 vs L4  Raw speed against low power See the comparison](https://gpuserver.io/compare/rtx-4090-vs-nvidia-l4)
- [RTX 4090 vs A100 PCIe  The cheap fast card against the data-centre one See the comparison](https://gpuserver.io/compare/rtx-4090-vs-a100-40gb)

## What is included

- ### No identity check

  No document, no selfie, no phone number, no company registration.

  Nothing verified · at any spend
- ### The whole card

  The GPU is passed through to one machine, and that machine is yours.

  No MIG · no vGPU · no hypervisor
- ### Root and IPMI

  Full root from the first minute, plus remote power and virtual media.

  Out-of-band on its own VLAN
- ### Unmetered bandwidth

  No included volume, no overage tier, nothing to watch on a graph.

  1 to 25 Gbit/s · both directions
- ### No setup or reinstall fee

  Reimage as often as you like, to any operating system we offer.

  $0 · reinstalls unlimited
- ### Hardware swapped, fast

  From spares held in the same suite, at any hour, disks left in place.

  Four-hour target · 24/7

## Where you can have it

Every RTX 4090 configuration is available in all 6 data centres. There is no site where a card costs more, and none where it is unavailable.

AMS Available

### Amsterdam

Netherlands

Round trip 7 ms to Frankfurt, 12 ms to London

Facility Tier III · 100% wind

STO Available

### Stockholm

Sweden

Round trip 22 ms to Frankfurt, 31 ms to London

Facility Tier III · 100% hydro

ZRH Available

### Zürich

Switzerland

Round trip 9 ms to Milan, 15 ms to Frankfurt

Facility Tier IV · 96% carbon-free

REY Available

### Reykjavík

Iceland

Round trip 18 ms to London, 40 ms to New York

Facility Tier III · 100% geothermal + hydro

YUL Available

### Montréal

Canada

Round trip 12 ms to New York, 19 ms to Toronto

Facility Tier III · 99% hydro

SIN Available

### Singapore

Singapore

Round trip 38 ms to Tokyo, 45 ms to Sydney

Facility Tier III · Grid + REC offset

## RTX 4090 server questions

How much does an RTX 4090 server cost per month?

A single-card RTX 4090 server is $193 a month, with everything included: 128 GB of system RAM, 2 × 2 TB NVMe, unmetered 1 Gbit/s, one IPv4 address and a routed IPv6 /64. There is no setup fee and no bandwidth charge. Multi-GPU nodes run from $416 for 2 cards up to $1,579 for 8.

How much VRAM does the RTX 4090 have, and what fits in it?

24 GB of GDDR6X per card, at 1,008 GB/s. The largest model it holds on one card at 8k context is Qwen 3 32B (4-bit (AWQ, GPTQ)). A 8-card node has 192 GB in total, enough for Qwen 3 235B-A22B (MoE).

Is the RTX 4090 dedicated, or shared with other customers?

Dedicated. The card is passed through to a single physical machine that only you have an account on — no MIG partitioning, no vGPU layer and no other tenant on the same silicon. On a multi-GPU node you get the whole node, including the PCIe 5.0 ×16 fabric between the cards.

Do I need to verify my identity to rent an RTX 4090?

No. There is no document upload, no selfie, no phone number and no company registration, at any level of spend — an eight-card node opens on exactly the same terms as the cheapest single card. Opening an account takes an email address, a password and a billing address, and none of it is verified.

How do I pay, and how quickly is the server delivered?

In cryptocurrency only: Bitcoin, Monero, Ethereum (ERC-20), Tether (ERC-20) and others. Each invoice gets its own address, the amount is fixed at the rate shown when it is issued, and the machine is provisioned on the first confirmation — root access in under 5 minutes, at any hour.

Can I rent an RTX 4090 by the hour instead?

Not here. Every price on this site is a monthly price, with no meter and no per-second billing. Against a typical $2.99/hour cloud, a monthly term is cheaper past roughly 65 hours of the machine simply existing — about 3 days out of thirty.

## An RTX 4090 of your own, from $193 a month.

Racked, burned in and waiting in all 6 data centres. Pay in crypto and it is yours in under 5 minutes.

[Configure this server](https://gpuserver.io/configure?c=rtx4090-x1) [Read the guides](https://gpuserver.io/guides)

## Related pages

- [Hardware policy The twelve NVIDIA GPUs we operate, the 72-hour burn-in before a node is sold, the four-hour replacement target, and how disks are erased between tenants.](https://gpuserver.io/hardware)
- [Guides Eight practical guides: VRAM sizing, NVLink versus PCIe, image and video models, monthly versus hourly, self-hosting versus an API, serving Llama 70B, and QLoRA.](https://gpuserver.io/guides)
- [Documentation First SSH, verifying the hardware, CUDA containers, serving a model with vLLM, RAID, firewalling, IPMI and reinstalls — the commands, on a real machine.](https://gpuserver.io/docs)
- [Network Unmetered ports up to 25 Gbit/s, two carriers and an IX per site, always-on DDoS filtering, routed IPv6 — and the four things we do not offer, stated up front.](https://gpuserver.io/network)

---

Source: https://gpuserver.io/gpu/rtx-4090/. This file is generated from the same data as the website; if a figure here differs from a page, the page is authoritative and this file is stale — the canonical source is https://gpuserver.io/.
