Ampere · PCIe card

# Rent a dedicated A100 PCIe GPU server.

40 GB of HBM2 at 1,555 GB/s, in a machine that is yours alone. Monthly, paid in crypto, opened without an identity check and delivered in under 5 minutes.

- $392/mo one card, everything included
- 40 GB HBM2
- 1,555 GB/s memory bandwidth
- 1×, 2×, 4×, 8× node sizes available

## Price, by node size and term

One number per line, and it is the whole number: no setup fee, no reinstall fee, no bandwidth charge, no support plan. Longer terms are paid up front.

**Monthly price of an A100 PCIe server, by node size and term length**

| Node | CPU and memory | Monthly | 3 months | 12 months |  |
|---|---|---|---|---|---|
| 1 × A100 PCIe 40 GB total · PCIe 5.0 | Intel Xeon Silver 4410Y12 c / 24 t · 128 GB · 2 × 2 TB NVMe | $392/mo cancel any time | $372/mo save $60/yr | $345/mo save $564/yr | [Configure](https://gpuserver.io/configure?c=a100-40-x1) |
| 2 × A100 PCIe 80 GB total · PCIe 5.0 ×16 | Intel Xeon Gold 5416S16 c / 32 t · 256 GB · 2 × 3.84 TB NVMe | $802/mo cancel any time | $762/mo save $120/yr | $706/mo save $1,152/yr | [Configure](https://gpuserver.io/configure?c=a100-40-x2) |
| 4 × A100 PCIe 160 GB total · PCIe 5.0 ×16 | 2 × Intel Xeon Gold 6438Y+64 c · 512 GB · 4 × 3.84 TB NVMe | $1,556/mo cancel any time | $1,478/mo save $234/yr | $1,369/mo save $2,244/yr | [Configure](https://gpuserver.io/configure?c=a100-40-x4) |
| 8 × A100 PCIe 320 GB total · PCIe 5.0 ×16 | 2 × Intel Xeon Platinum 8462Y+64 c · 1 TB · 8 × 3.84 TB NVMe | $2,983/mo cancel any time | $2,834/mo save $447/yr | $2,625/mo save $4,296/yr | [Configure](https://gpuserver.io/configure?c=a100-40-x8) |

Every configuration above is available in all 6 data centres, racked and burned in, with an image ready to write — which is why delivery is under 5 minutes rather than a working day.

## Full specification

GPU NVIDIA A100 PCIe Ampere architecture, PCIe add-in card.

Memory 40 GB HBM2 1,555 GB/s of bandwidth — the number that predicts generation speed, far better than any teraflop figure.

Interconnect PCIe 5.0 ×16 Peer-to-peer over PCIe. Fine for one model per card and for pipeline parallelism; slower than NVLink for tensor parallelism. [Which one your job needs](https://gpuserver.io/guides/nvlink-vs-pcie).

Node sizes 1 × card, 2 × card, 4 × card, 8 × card PCIe cards go up to eight in a standard chassis, and ten only on low-power cards.

Storage 2 × 2 TB NVMe Handed over raw beyond the system volume — we impose no RAID level. More NVMe and archive HDD are monthly options.

Network 1 Gbit/s, unmetered Both directions, no included volume and no overage at any traffic level. One IPv4, a routed IPv6 /64, DDoS filtering on from delivery.

## What runs on an A100 PCIe

Every model below fits on a *single* card at 8k context — no sharding, no interconnect to think about. The precision shown is the best one that fits; the arithmetic is in the [VRAM sizing guide](https://gpuserver.io/guides/vram-sizing), so you can check it.

**Language models that fit on a single A100 PCIe at 8k context**

| Model | Best precision that fits | Memory needed | Estimated tokens/s |
|---|---|---|---|
| Llama 3.1 8B 8B parameters | BF16 / FP16Reference quality | 20 GBof 40 GB | ~44 |
| Qwen 3 32B 32B parameters | FP8Near-reference, Hopper and Blackwell | 38 GBof 40 GB | ~22 |
| Mistral Small 24B 24B parameters | FP8Near-reference, Hopper and Blackwell | 28 GBof 40 GB | ~29 |
| Gemma 3 27B 27B parameters | FP8Near-reference, Hopper and Blackwell | 33 GBof 40 GB | ~26 |
| Phi-4 14B 14B parameters | BF16 / FP16Reference quality | 34 GBof 40 GB | ~25 |

On the 8-card node the same models are held across 320 GB in total, which changes what is possible entirely — up to Llama 3.1 405B. Sharding costs synchronisation at every layer, so a model that fits on one card should stay on one card.

## Throughput to expect

Generation is bounded by memory bandwidth: each token requires re-reading the active weights. At 1,555 GB/s, that sets a ceiling no amount of tuning gets past.

**Our figures are a floor, not a promise.** The estimates on this page are calibrated against published measurements and are deliberately conservative — they model single-stream generation. With continuous batching under vLLM the *aggregate* across concurrent requests is several times higher. If a number here looks low against a benchmark you have seen, that is usually the difference.

## Compared with nearby cards

The three cards closest to this one in price. Memory decides what loads; bandwidth decides how fast it answers.

**The A100 PCIe compared with the three closest cards by price**

| Card | Memory | Bandwidth | Interconnect | From |
|---|---|---|---|---|
| A100 PCIe This page Ampere | 40 GB | 1,555 GB/s | PCIe | $392/mo |
| [RTX 5090](https://gpuserver.io/gpu/rtx-5090) Blackwell | 32 GB | 1,792 GB/s | PCIe | $335/mo |
| [RTX A6000](https://gpuserver.io/gpu/rtx-a6000) Ampere | 48 GB | 768 GB/s | PCIe | $286/mo |
| [RTX 4090](https://gpuserver.io/gpu/rtx-4090) Ada Lovelace | 24 GB | 1,008 GB/s | PCIe | $193/mo |

[See all 12 cards side by side](https://gpuserver.io/gpu), or let the [configurator](https://gpuserver.io/configure) pick from your model and context length instead of from a price.

## Compared with other cards

The arbitrations people actually make against an A100 PCIe, each worked out on price, memory, bandwidth and the models that fit.

- [A100 PCIe vs A100 PCIe  The same GPU with twice the memory See the comparison](https://gpuserver.io/compare/a100-80gb-vs-a100-40gb)
- [A100 PCIe vs RTX 4090  The cheap fast card against the data-centre one See the comparison](https://gpuserver.io/compare/rtx-4090-vs-a100-40gb)

## What is included

- ### No identity check

  No document, no selfie, no phone number, no company registration.

  Nothing verified · at any spend
- ### The whole card

  The GPU is passed through to one machine, and that machine is yours.

  No MIG · no vGPU · no hypervisor
- ### Root and IPMI

  Full root from the first minute, plus remote power and virtual media.

  Out-of-band on its own VLAN
- ### Unmetered bandwidth

  No included volume, no overage tier, nothing to watch on a graph.

  1 to 25 Gbit/s · both directions
- ### No setup or reinstall fee

  Reimage as often as you like, to any operating system we offer.

  $0 · reinstalls unlimited
- ### Hardware swapped, fast

  From spares held in the same suite, at any hour, disks left in place.

  Four-hour target · 24/7

## Where you can have it

Every A100 PCIe configuration is available in all 6 data centres. There is no site where a card costs more, and none where it is unavailable.

AMS Available

### Amsterdam

Netherlands

Round trip 7 ms to Frankfurt, 12 ms to London

Facility Tier III · 100% wind

STO Available

### Stockholm

Sweden

Round trip 22 ms to Frankfurt, 31 ms to London

Facility Tier III · 100% hydro

ZRH Available

### Zürich

Switzerland

Round trip 9 ms to Milan, 15 ms to Frankfurt

Facility Tier IV · 96% carbon-free

REY Available

### Reykjavík

Iceland

Round trip 18 ms to London, 40 ms to New York

Facility Tier III · 100% geothermal + hydro

YUL Available

### Montréal

Canada

Round trip 12 ms to New York, 19 ms to Toronto

Facility Tier III · 99% hydro

SIN Available

### Singapore

Singapore

Round trip 38 ms to Tokyo, 45 ms to Sydney

Facility Tier III · Grid + REC offset

## A100 PCIe server questions

How much does an A100 PCIe server cost per month?

A single-card A100 PCIe server is $392 a month, with everything included: 128 GB of system RAM, 2 × 2 TB NVMe, unmetered 1 Gbit/s, one IPv4 address and a routed IPv6 /64. There is no setup fee and no bandwidth charge. Multi-GPU nodes run from $802 for 2 cards up to $2,983 for 8.

How much VRAM does the A100 PCIe have, and what fits in it?

40 GB of HBM2 per card, at 1,555 GB/s. The largest model it holds on one card at 8k context is Qwen 3 32B (FP8). A 8-card node has 320 GB in total, enough for Llama 3.1 405B.

Is the A100 PCIe dedicated, or shared with other customers?

Dedicated. The card is passed through to a single physical machine that only you have an account on — no MIG partitioning, no vGPU layer and no other tenant on the same silicon. On a multi-GPU node you get the whole node, including the PCIe 5.0 ×16 fabric between the cards.

Do I need to verify my identity to rent an A100 PCIe?

No. There is no document upload, no selfie, no phone number and no company registration, at any level of spend — an eight-card node opens on exactly the same terms as the cheapest single card. Opening an account takes an email address, a password and a billing address, and none of it is verified.

How do I pay, and how quickly is the server delivered?

In cryptocurrency only: Bitcoin, Monero, Ethereum (ERC-20), Tether (ERC-20) and others. Each invoice gets its own address, the amount is fixed at the rate shown when it is issued, and the machine is provisioned on the first confirmation — root access in under 5 minutes, at any hour.

Can I rent an A100 PCIe by the hour instead?

Not here. Every price on this site is a monthly price, with no meter and no per-second billing. Against a typical $2.99/hour cloud, a monthly term is cheaper past roughly 132 hours of the machine simply existing — about 5 days out of thirty.

## An A100 PCIe of your own, from $392 a month.

Racked, burned in and waiting in all 6 data centres. Pay in crypto and it is yours in under 5 minutes.

[Configure this server](https://gpuserver.io/configure?c=a100-40-x1) [Read the guides](https://gpuserver.io/guides)

## Related pages

- [Hardware policy The twelve NVIDIA GPUs we operate, the 72-hour burn-in before a node is sold, the four-hour replacement target, and how disks are erased between tenants.](https://gpuserver.io/hardware)
- [Guides Eight practical guides: VRAM sizing, NVLink versus PCIe, image and video models, monthly versus hourly, self-hosting versus an API, serving Llama 70B, and QLoRA.](https://gpuserver.io/guides)
- [Documentation First SSH, verifying the hardware, CUDA containers, serving a model with vLLM, RAID, firewalling, IPMI and reinstalls — the commands, on a real machine.](https://gpuserver.io/docs)
- [Network Unmetered ports up to 25 Gbit/s, two carriers and an IX per site, always-on DDoS filtering, routed IPv6 — and the four things we do not offer, stated up front.](https://gpuserver.io/network)

---

Source: https://gpuserver.io/gpu/a100-40gb/. This file is generated from the same data as the website; if a figure here differs from a page, the page is authoritative and this file is stale — the canonical source is https://gpuserver.io/.
