Ampere · SXM module

# Rent a dedicated A100 SXM4 GPU server.

80 GB of HBM2e at 2,039 GB/s, in a machine that is yours alone. Monthly, paid in crypto, opened without an identity check and delivered in under 5 minutes.

- $4,395/mo one card, everything included
- 80 GB HBM2e
- 2,039 GB/s memory bandwidth
- 4×, 8× node sizes available

## Price, by node size and term

One number per line, and it is the whole number: no setup fee, no reinstall fee, no bandwidth charge, no support plan. Longer terms are paid up front.

**Monthly price of an A100 SXM4 server, by node size and term length**

| Node | CPU and memory | Monthly | 3 months | 12 months |  |
|---|---|---|---|---|---|
| 4 × A100 SXM4 320 GB total · NVLink 3 · 600 GB/s | 2 × Intel Xeon Gold 6438Y+64 c · 512 GB · 4 × 3.84 TB NVMe | $4,395/mo cancel any time | $4,175/mo save $660/yr | $3,868/mo save $6,324/yr | [Configure](https://gpuserver.io/configure?c=a100-sxm-x4) |
| 8 × A100 SXM4 640 GB total · NVLink 3 · 600 GB/s | 2 × Intel Xeon Platinum 8462Y+64 c · 1 TB · 8 × 3.84 TB NVMe | $8,355/mo cancel any time | $7,937/mo save $1,254/yr | $7,352/mo save $12,036/yr | [Configure](https://gpuserver.io/configure?c=a100-sxm-x8) |

Every configuration above is available in all 6 data centres, racked and burned in, with an image ready to write — which is why delivery is under 5 minutes rather than a working day.

## Full specification

GPU NVIDIA A100 SXM4 Ampere architecture, SXM module on an HGX board.

Memory 80 GB HBM2e 2,039 GB/s of bandwidth — the number that predicts generation speed, far better than any teraflop figure.

Interconnect NVLink 3 · 600 GB/s All-to-all between every card on the board. Tensor parallelism without the link becoming the bottleneck. [Which one your job needs](https://gpuserver.io/guides/nvlink-vs-pcie).

Node sizes 4 × card, 8 × card SXM cards live on an HGX board of four or eight — never two, never ten.

Storage 4 × 3.84 TB NVMe Handed over raw beyond the system volume — we impose no RAID level. More NVMe and archive HDD are monthly options.

Network 10 Gbit/s, unmetered Both directions, no included volume and no overage at any traffic level. One IPv4, a routed IPv6 /64, DDoS filtering on from delivery.

## What runs on an A100 SXM4

Every model below fits on a *single* card at 8k context — no sharding, no interconnect to think about. The precision shown is the best one that fits; the arithmetic is in the [VRAM sizing guide](https://gpuserver.io/guides/vram-sizing), so you can check it.

**Language models that fit on a single A100 SXM4 at 8k context**

| Model | Best precision that fits | Memory needed | Estimated tokens/s |
|---|---|---|---|
| Llama 3.1 8B 8B parameters | BF16 / FP16Reference quality | 20 GBof 80 GB | ~57 |
| Llama 3.3 70B 70B parameters | 4-bit (AWQ, GPTQ)Smallest weights, cache stays FP16 | 43 GBof 80 GB | ~26 |
| Qwen 3 32B 32B parameters | BF16 / FP16Reference quality | 76 GBof 80 GB | ~14 |
| Mistral Small 24B 24B parameters | BF16 / FP16Reference quality | 57 GBof 80 GB | ~19 |
| Gemma 3 27B 27B parameters | BF16 / FP16Reference quality | 67 GBof 80 GB | ~17 |
| Phi-4 14B 14B parameters | BF16 / FP16Reference quality | 34 GBof 80 GB | ~33 |

On the 8-card node the same models are held across 640 GB in total, which changes what is possible entirely — up to DeepSeek V3 671B-A37B (MoE). Sharding costs synchronisation at every layer, so a model that fits on one card should stay on one card.

## Throughput to expect

Generation is bounded by memory bandwidth: each token requires re-reading the active weights. At 2,039 GB/s, that sets a ceiling no amount of tuning gets past.

**Our figures are a floor, not a promise.** The estimates on this page are calibrated against published measurements and are deliberately conservative — they model single-stream generation. With continuous batching under vLLM the *aggregate* across concurrent requests is several times higher. If a number here looks low against a benchmark you have seen, that is usually the difference.

## Compared with nearby cards

The three cards closest to this one in price. Memory decides what loads; bandwidth decides how fast it answers.

**The A100 SXM4 compared with the three closest cards by price**

| Card | Memory | Bandwidth | Interconnect | From |
|---|---|---|---|---|
| A100 SXM4 This page Ampere | 80 GB | 2,039 GB/s | NVLink | $4,395/mo |
| [H100 SXM5](https://gpuserver.io/gpu/h100-sxm5) Hopper | 80 GB | 3,350 GB/s | NVLink | $6,061/mo |
| [H100 PCIe](https://gpuserver.io/gpu/h100-pcie) Hopper | 80 GB | 2,000 GB/s | PCIe | $1,469/mo |
| [A100 PCIe](https://gpuserver.io/gpu/a100-80gb) Ampere | 80 GB | 1,935 GB/s | PCIe | $1,091/mo |

[See all 12 cards side by side](https://gpuserver.io/gpu), or let the [configurator](https://gpuserver.io/configure) pick from your model and context length instead of from a price.

## Compared with other cards

The arbitrations people actually make against an A100 SXM4, each worked out on price, memory, bandwidth and the models that fit.

- [A100 SXM4 vs H100 SXM5  The generational jump, on the same board See the comparison](https://gpuserver.io/compare/h100-sxm5-vs-a100-sxm4)

## What is included

- ### No identity check

  No document, no selfie, no phone number, no company registration.

  Nothing verified · at any spend
- ### The whole card

  The GPU is passed through to one machine, and that machine is yours.

  No MIG · no vGPU · no hypervisor
- ### Root and IPMI

  Full root from the first minute, plus remote power and virtual media.

  Out-of-band on its own VLAN
- ### Unmetered bandwidth

  No included volume, no overage tier, nothing to watch on a graph.

  1 to 25 Gbit/s · both directions
- ### No setup or reinstall fee

  Reimage as often as you like, to any operating system we offer.

  $0 · reinstalls unlimited
- ### Hardware swapped, fast

  From spares held in the same suite, at any hour, disks left in place.

  Four-hour target · 24/7

## Where you can have it

Every A100 SXM4 configuration is available in all 6 data centres. There is no site where a card costs more, and none where it is unavailable.

AMS Available

### Amsterdam

Netherlands

Round trip 7 ms to Frankfurt, 12 ms to London

Facility Tier III · 100% wind

STO Available

### Stockholm

Sweden

Round trip 22 ms to Frankfurt, 31 ms to London

Facility Tier III · 100% hydro

ZRH Available

### Zürich

Switzerland

Round trip 9 ms to Milan, 15 ms to Frankfurt

Facility Tier IV · 96% carbon-free

REY Available

### Reykjavík

Iceland

Round trip 18 ms to London, 40 ms to New York

Facility Tier III · 100% geothermal + hydro

YUL Available

### Montréal

Canada

Round trip 12 ms to New York, 19 ms to Toronto

Facility Tier III · 99% hydro

SIN Available

### Singapore

Singapore

Round trip 38 ms to Tokyo, 45 ms to Sydney

Facility Tier III · Grid + REC offset

## A100 SXM4 server questions

How much does an A100 SXM4 server cost per month?

A single-card A100 SXM4 server is $4,395 a month, with everything included: 512 GB of system RAM, 4 × 3.84 TB NVMe, unmetered 10 Gbit/s, one IPv4 address and a routed IPv6 /64. There is no setup fee and no bandwidth charge. Multi-GPU nodes run from $8,355 for 8 cards up to $8,355 for 8.

How much VRAM does the A100 SXM4 have, and what fits in it?

80 GB of HBM2e per card, at 2,039 GB/s. The largest model it holds on one card at 8k context is Llama 3.1 405B (4-bit (AWQ, GPTQ)). A 8-card node has 640 GB in total, enough for DeepSeek V3 671B-A37B (MoE).

Is the A100 SXM4 dedicated, or shared with other customers?

Dedicated. The card is passed through to a single physical machine that only you have an account on — no MIG partitioning, no vGPU layer and no other tenant on the same silicon. On a multi-GPU node you get the whole node, including the NVLink 3 · 600 GB/s fabric between the cards.

Do I need to verify my identity to rent an A100 SXM4?

No. There is no document upload, no selfie, no phone number and no company registration, at any level of spend — an eight-card node opens on exactly the same terms as the cheapest single card. Opening an account takes an email address, a password and a billing address, and none of it is verified.

How do I pay, and how quickly is the server delivered?

In cryptocurrency only: Bitcoin, Monero, Ethereum (ERC-20), Tether (ERC-20) and others. Each invoice gets its own address, the amount is fixed at the rate shown when it is issued, and the machine is provisioned on the first confirmation — root access in under 5 minutes, at any hour.

Can I rent an A100 SXM4 by the hour instead?

Not here. Every price on this site is a monthly price, with no meter and no per-second billing. Against a typical $2.99/hour cloud, a monthly term is cheaper past roughly 1470 hours of the machine simply existing — about 61 days out of thirty.

## An A100 SXM4 of your own, from $4,395 a month.

Racked, burned in and waiting in all 6 data centres. Pay in crypto and it is yours in under 5 minutes.

[Configure this server](https://gpuserver.io/configure?c=a100-sxm-x4) [Read the guides](https://gpuserver.io/guides)

## Related pages

- [Hardware policy The twelve NVIDIA GPUs we operate, the 72-hour burn-in before a node is sold, the four-hour replacement target, and how disks are erased between tenants.](https://gpuserver.io/hardware)
- [Guides Eight practical guides: VRAM sizing, NVLink versus PCIe, image and video models, monthly versus hourly, self-hosting versus an API, serving Llama 70B, and QLoRA.](https://gpuserver.io/guides)
- [Documentation First SSH, verifying the hardware, CUDA containers, serving a model with vLLM, RAID, firewalling, IPMI and reinstalls — the commands, on a real machine.](https://gpuserver.io/docs)
- [Network Unmetered ports up to 25 Gbit/s, two carriers and an IX per site, always-on DDoS filtering, routed IPv6 — and the four things we do not offer, stated up front.](https://gpuserver.io/network)

---

Source: https://gpuserver.io/gpu/a100-sxm4/. This file is generated from the same data as the website; if a figure here differs from a page, the page is authoritative and this file is stale — the canonical source is https://gpuserver.io/.
