Ada Lovelace · PCIe card

# Rent a dedicated L4 GPU server.

24 GB of GDDR6 at 300 GB/s, in a machine that is yours alone. Monthly, paid in crypto, opened without an identity check and delivered in under 5 minutes.

- $125/mo one card, everything included
- 24 GB GDDR6
- 300 GB/s memory bandwidth
- 1×, 2×, 4×, 8×, 10× node sizes available

## Price, by node size and term

One number per line, and it is the whole number: no setup fee, no reinstall fee, no bandwidth charge, no support plan. Longer terms are paid up front.

**Monthly price of an L4 server, by node size and term length**

| Node | CPU and memory | Monthly | 3 months | 12 months |  |
|---|---|---|---|---|---|
| 1 × L4 24 GB total · PCIe 5.0 | Intel Xeon Silver 4410Y12 c / 24 t · 128 GB · 2 × 2 TB NVMe | $125/mo cancel any time | $119/mo save $18/yr | $110/mo save $180/yr | [Configure](https://gpuserver.io/configure?c=l4-x1) |
| 2 × L4 48 GB total · PCIe 5.0 ×16 | Intel Xeon Gold 5416S16 c / 32 t · 256 GB · 2 × 3.84 TB NVMe | $285/mo cancel any time | $271/mo save $42/yr | $251/mo save $408/yr | [Configure](https://gpuserver.io/configure?c=l4-x2) |
| 4 × L4 96 GB total · PCIe 5.0 ×16 | 2 × Intel Xeon Gold 6438Y+64 c · 512 GB · 4 × 3.84 TB NVMe | $564/mo cancel any time | $536/mo save $84/yr | $496/mo save $816/yr | [Configure](https://gpuserver.io/configure?c=l4-x4) |
| 8 × L4 192 GB total · PCIe 5.0 ×16 | 2 × Intel Xeon Platinum 8462Y+64 c · 1 TB · 8 × 3.84 TB NVMe | $1,106/mo cancel any time | $1,051/mo save $165/yr | $973/mo save $1,596/yr | [Configure](https://gpuserver.io/configure?c=l4-x8) |
| 10 × L4 240 GB total · PCIe 5.0 ×16 | 2 × Intel Xeon Platinum 8462Y+64 c · 1 TB · 8 × 7.68 TB NVMe | $1,371/mo cancel any time | $1,302/mo save $207/yr | $1,206/mo save $1,980/yr | [Configure](https://gpuserver.io/configure?c=l4-x10) |

Every configuration above is available in all 6 data centres, racked and burned in, with an image ready to write — which is why delivery is under 5 minutes rather than a working day.

## Full specification

GPU NVIDIA L4 Ada Lovelace architecture, PCIe add-in card.

Memory 24 GB GDDR6 300 GB/s of bandwidth — the number that predicts generation speed, far better than any teraflop figure.

Interconnect PCIe 5.0 ×16 Peer-to-peer over PCIe. Fine for one model per card and for pipeline parallelism; slower than NVLink for tensor parallelism. [Which one your job needs](https://gpuserver.io/guides/nvlink-vs-pcie).

Node sizes 1 × card, 2 × card, 4 × card, 8 × card, 10 × card PCIe cards go up to eight in a standard chassis, and ten only on low-power cards.

Storage 2 × 2 TB NVMe Handed over raw beyond the system volume — we impose no RAID level. More NVMe and archive HDD are monthly options.

Network 1 Gbit/s, unmetered Both directions, no included volume and no overage at any traffic level. One IPv4, a routed IPv6 /64, DDoS filtering on from delivery.

## What runs on an L4

Every model below fits on a *single* card at 8k context — no sharding, no interconnect to think about. The precision shown is the best one that fits; the arithmetic is in the [VRAM sizing guide](https://gpuserver.io/guides/vram-sizing), so you can check it.

**Language models that fit on a single L4 at 8k context**

| Model | Best precision that fits | Memory needed | Estimated tokens/s |
|---|---|---|---|
| Llama 3.1 8B 8B parameters | BF16 / FP16Reference quality | 20 GBof 24 GB | ~8 |
| Qwen 3 32B 32B parameters | 4-bit (AWQ, GPTQ)Smallest weights, cache stays FP16 | 21 GBof 24 GB | ~8 |
| Mistral Small 24B 24B parameters | 4-bit (AWQ, GPTQ)Smallest weights, cache stays FP16 | 15 GBof 24 GB | ~11 |
| Gemma 3 27B 27B parameters | 4-bit (AWQ, GPTQ)Smallest weights, cache stays FP16 | 20 GBof 24 GB | ~10 |
| Phi-4 14B 14B parameters | FP8Near-reference, Hopper and Blackwell | 17 GBof 24 GB | ~10 |

On the 10-card node the same models are held across 240 GB in total, which changes what is possible entirely — up to Llama 3.1 405B. Sharding costs synchronisation at every layer, so a model that fits on one card should stay on one card.

## Throughput to expect

Generation is bounded by memory bandwidth: each token requires re-reading the active weights. At 300 GB/s, that sets a ceiling no amount of tuning gets past.

**Our figures are a floor, not a promise.** The estimates on this page are calibrated against published measurements and are deliberately conservative — they model single-stream generation. With continuous batching under vLLM the *aggregate* across concurrent requests is several times higher. If a number here looks low against a benchmark you have seen, that is usually the difference.

## Compared with nearby cards

The three cards closest to this one in price. Memory decides what loads; bandwidth decides how fast it answers.

**The L4 compared with the three closest cards by price**

| Card | Memory | Bandwidth | Interconnect | From |
|---|---|---|---|---|
| L4 This page Ada Lovelace | 24 GB | 300 GB/s | PCIe | $125/mo |
| [RTX 4090](https://gpuserver.io/gpu/rtx-4090) Ada Lovelace | 24 GB | 1,008 GB/s | PCIe | $193/mo |
| [RTX A6000](https://gpuserver.io/gpu/rtx-a6000) Ampere | 48 GB | 768 GB/s | PCIe | $286/mo |
| [RTX 5090](https://gpuserver.io/gpu/rtx-5090) Blackwell | 32 GB | 1,792 GB/s | PCIe | $335/mo |

[See all 12 cards side by side](https://gpuserver.io/gpu), or let the [configurator](https://gpuserver.io/configure) pick from your model and context length instead of from a price.

## Compared with other cards

The arbitrations people actually make against an L4, each worked out on price, memory, bandwidth and the models that fit.

- [L4 vs RTX 4090  Raw speed against low power See the comparison](https://gpuserver.io/compare/rtx-4090-vs-nvidia-l4)

## What is included

- ### No identity check

  No document, no selfie, no phone number, no company registration.

  Nothing verified · at any spend
- ### The whole card

  The GPU is passed through to one machine, and that machine is yours.

  No MIG · no vGPU · no hypervisor
- ### Root and IPMI

  Full root from the first minute, plus remote power and virtual media.

  Out-of-band on its own VLAN
- ### Unmetered bandwidth

  No included volume, no overage tier, nothing to watch on a graph.

  1 to 25 Gbit/s · both directions
- ### No setup or reinstall fee

  Reimage as often as you like, to any operating system we offer.

  $0 · reinstalls unlimited
- ### Hardware swapped, fast

  From spares held in the same suite, at any hour, disks left in place.

  Four-hour target · 24/7

## Where you can have it

Every L4 configuration is available in all 6 data centres. There is no site where a card costs more, and none where it is unavailable.

AMS Available

### Amsterdam

Netherlands

Round trip 7 ms to Frankfurt, 12 ms to London

Facility Tier III · 100% wind

STO Available

### Stockholm

Sweden

Round trip 22 ms to Frankfurt, 31 ms to London

Facility Tier III · 100% hydro

ZRH Available

### Zürich

Switzerland

Round trip 9 ms to Milan, 15 ms to Frankfurt

Facility Tier IV · 96% carbon-free

REY Available

### Reykjavík

Iceland

Round trip 18 ms to London, 40 ms to New York

Facility Tier III · 100% geothermal + hydro

YUL Available

### Montréal

Canada

Round trip 12 ms to New York, 19 ms to Toronto

Facility Tier III · 99% hydro

SIN Available

### Singapore

Singapore

Round trip 38 ms to Tokyo, 45 ms to Sydney

Facility Tier III · Grid + REC offset

## L4 server questions

How much does an L4 server cost per month?

A single-card L4 server is $125 a month, with everything included: 128 GB of system RAM, 2 × 2 TB NVMe, unmetered 1 Gbit/s, one IPv4 address and a routed IPv6 /64. There is no setup fee and no bandwidth charge. Multi-GPU nodes run from $285 for 2 cards up to $1,371 for 10.

How much VRAM does the L4 have, and what fits in it?

24 GB of GDDR6 per card, at 300 GB/s. The largest model it holds on one card at 8k context is Qwen 3 32B (4-bit (AWQ, GPTQ)). A 10-card node has 240 GB in total, enough for Llama 3.1 405B.

Is the L4 dedicated, or shared with other customers?

Dedicated. The card is passed through to a single physical machine that only you have an account on — no MIG partitioning, no vGPU layer and no other tenant on the same silicon. On a multi-GPU node you get the whole node, including the PCIe 5.0 ×16 fabric between the cards.

Do I need to verify my identity to rent an L4?

No. There is no document upload, no selfie, no phone number and no company registration, at any level of spend — an eight-card node opens on exactly the same terms as the cheapest single card. Opening an account takes an email address, a password and a billing address, and none of it is verified.

How do I pay, and how quickly is the server delivered?

In cryptocurrency only: Bitcoin, Monero, Ethereum (ERC-20), Tether (ERC-20) and others. Each invoice gets its own address, the amount is fixed at the rate shown when it is issued, and the machine is provisioned on the first confirmation — root access in under 5 minutes, at any hour.

Can I rent an L4 by the hour instead?

Not here. Every price on this site is a monthly price, with no meter and no per-second billing. Against a typical $2.99/hour cloud, a monthly term is cheaper past roughly 42 hours of the machine simply existing — about 2 days out of thirty.

## An L4 of your own, from $125 a month.

Racked, burned in and waiting in all 6 data centres. Pay in crypto and it is yours in under 5 minutes.

[Configure this server](https://gpuserver.io/configure?c=l4-x1) [Read the guides](https://gpuserver.io/guides)

## Related pages

- [Hardware policy The twelve NVIDIA GPUs we operate, the 72-hour burn-in before a node is sold, the four-hour replacement target, and how disks are erased between tenants.](https://gpuserver.io/hardware)
- [Guides Eight practical guides: VRAM sizing, NVLink versus PCIe, image and video models, monthly versus hourly, self-hosting versus an API, serving Llama 70B, and QLoRA.](https://gpuserver.io/guides)
- [Documentation First SSH, verifying the hardware, CUDA containers, serving a model with vLLM, RAID, firewalling, IPMI and reinstalls — the commands, on a real machine.](https://gpuserver.io/docs)
- [Network Unmetered ports up to 25 Gbit/s, two carriers and an IX per site, always-on DDoS filtering, routed IPv6 — and the four things we do not offer, stated up front.](https://gpuserver.io/network)

---

Source: https://gpuserver.io/gpu/nvidia-l4/. This file is generated from the same data as the website; if a figure here differs from a page, the page is authoritative and this file is stale — the canonical source is https://gpuserver.io/.
