Guides.
Longer pieces on the questions that decide which machine you should rent — written against the hardware we actually operate, with the arithmetic shown.
- Written against the hardware we operate
- Commands you can paste, not pseudocode
- Arithmetic shown, so you can check it
All 10 guides
5 on sizing · 2 on cost · 2 on practice · 1 on payment
-
Sizing · 9 min
How much VRAM a model actually needs
The arithmetic behind “will it fit”: weights, KV cache, and the two places a rule of thumb goes wrong by a factor of three.
Read the guide -
Sizing · 7 min
NVLink or PCIe: which one your job needs
When the link between cards decides your throughput, when it changes nothing, and how to tell which case you are in before you pay for the wrong node.
Read the guide -
Sizing · 9 min
Choosing a card for image and video models
What FLUX, SDXL, SD 3.5 and Wan 2.1 need in VRAM, which card in our catalogue holds each, and why the cheapest one that fits is rarely the one to rent.
Read the guide -
Sizing · 10 min
Choosing a quantisation format
What AWQ, GPTQ, GGUF and FP8 each cost in memory, speed and quality — and why the format with the smallest weights rarely gives the smallest model.
Read the guide -
Sizing · 10 min
Running a mixture-of-experts model
Total parameters decide the machine, active parameters decide the speed. What DeepSeek V3 and Qwen 3 235B need in VRAM, and which node to rent for them.
Read the guide -
Cost · 6 min
When monthly rental beats per-hour
The break-even worked out on real numbers, the three costs an hourly price hides until the invoice, and the idle-time trap that triples an estimate.
Read the guide -
Cost · 9 min
Self-hosting against a per-token API
The break-even between paying an API per token and renting a GPU, worked out on our own prices and throughput — and the four things the arithmetic leaves out.
Read the guide -
Practice · 11 min
Serving Llama 3.3 70B on one node
From a delivered machine to an OpenAI-compatible endpoint, with the flags that matter and the two that silently halve your throughput.
Read the guide -
Practice · 10 min
Fine-tuning a 70B model on one card
QLoRA on a single 48 GB GPU: what fits, what it costs for a month, and why full fine-tuning is a different order of machine.
Read the guide -
Payment · 8 min
Paying for a server in crypto
What actually happens between clicking pay and getting root, which coin to choose, and the four mistakes that lose money on a first payment.
Read the guide
Ten questions decide which machine you should rent.
We would rather write ten pieces that answer a real question than sixty that rank for a search term. Each guide above exists because it is a question we are asked before an order, and each one is written against the cards in our own catalogue — the numbers in them are numbers from machines we operate.
Where a guide shows a formula, it is the same formula the configurator uses to tell you whether a model fits. That is deliberate: you should be able to check our arithmetic rather than take the green tick on trust.
Straight to the point
If you already know what you need
Every guide ends at the same place.
A machine you can rent by the month, pay for in crypto, and open without proving who you are.