Latitude.sh prices a bare-metal NVIDIA H100 80 GB at $1.68 an hour or $1,230 a month. Multiply the hourly rate by the 730 hours in an average month and you get $1,226.40 — the monthly plan costs $3.60 more than running the same server by the hour all month. Cherry Servers, meanwhile, takes about 19% off an A40 if you commit to the month.
Both are honest rate cards, and the difference between them is the most useful thing to understand before you sign for GPU capacity: the higher you go up the GPU range, the less a commitment buys you in price — and the more it buys you in access. This piece works through the published numbers and what they mean in practice.
What a GPU dedicated server costs, from published rate cards
Every figure below comes from the provider’s own public page, read on 5 October 2026; the Latitude.sh, Hetzner and OVHcloud figures were re-checked on 7 October 2026. Currencies are as published, Hetzner prices exclude VAT, and no FX has been applied. “Break-even hours” is the monthly price divided by the hourly price — the number of hours a month above which the monthly plan becomes cheaper than paying by the hour.
| Provider | Configuration | Hourly | Monthly | Break-even hours | Monthly saving | Included traffic |
|---|---|---|---|---|---|---|
| Hetzner | GEX131 — 1x RTX PRO 6000 Blackwell Max-Q, 96 GB | EUR 1.4247 | EUR 889 | 624 (85%) | 14.5% | Not stated in announcement |
| Cherry Servers | 1x A40, 48 GB (pre-order) | $0.801 | $472.66 | 590 (81%) | 19.2% | 30–100 TB |
| Cherry Servers | 1x A100, 80 GB (pre-order) | $2.362 | $1,657.68 | 702 (96%) | 3.9% | 30–100 TB |
| Latitude.sh | 1x H100 80 GB, bare metal | $1.68 | $1,230 | 732 (over 100%) | none (−0.3%) | 20 TB out, inbound unmetered |
| Latitude.sh | 8x RTX PRO 6000, bare metal | $24.00 | $17,520 | 730 (100%) | none | 20 TB out |
| Latitude.sh | 8x HGX B300, bare metal | $64.00 | $46,720 | 730 (100%) | none | 20 TB out |
| OVHcloud | HGR-AI-2 — L40S | not published | from $4,355 + $4,355 installation | — | — | not stated |
| Atlantic.Net | 1x H100 NVL, managed, 12-month term | not published | $4,729.44 | — | — | 10 TB |
Sources: Hetzner pressroom (GEX131), Cherry Servers dedicated GPU servers page, Latitude.sh pricing, OVHcloud GPU dedicated server page, Atlantic.Net dedicated GPU hosting. Only the Hetzner announcement carries a date.
Two things follow from the table.
The commitment discount narrows as the hardware gets more in demand. About 19% on an A40, 14.5% on Hetzner’s workstation-class RTX PRO 6000 Max-Q, about 4% on an A100, and nothing on Latitude.sh’s H100, RTX PRO 6000 and B300 nodes. It is not a clean rule — Hetzner still discounts a current-generation Blackwell card — but the direction is consistent.
Headline prices for the “same” GPU are not comparable until you normalise the service. $1,230 and $4,729.44 are both “H100-class” servers, but one is a self-managed H100 80 GB billed by the month and the other is a managed H100 NVL on a 12-month term with a security stack and support included. The gap is management, support and contract terms, not silicon.
Why the monthly discount shrinks on top-tier GPUs
On Cherry Servers’ own card, the mid-range accelerators (P4, A2, A10, A16) and the A40 all break even at 583–590 hours — a consistent 19–20% discount for a monthly commitment. The A100 breaks even at 702 hours, 96% of the month, for a 3.9% discount. On Latitude.sh’s 8x RTX PRO 6000 and 8x B300 nodes the monthly price is the hourly rate multiplied by exactly 730, to the dollar: there is no monthly discount, only a monthly-sized invoice.
A monthly discount is the provider paying you to take idle-capacity risk off its hands. Where cards can sit unrented, that is worth paying for. Where every unit is allocated as soon as it comes online, the provider carries no idle risk and has no reason to discount. That matches what we see in our own inventory: H200 and B300 capacity is taken as fast as it is installed, while workstation-class cards such as the RTX A4000, RTX A5000 and the RTX PRO series are where availability — and room on price — still exist.
One practical check before you treat any monthly figure as a discount: ask how it is applied. On some offers the monthly price is a separate term with a minimum commitment; on others it works more like a ceiling on hourly billing. The break-even arithmetic is the same either way; the cost of leaving early is not.
Is it cheaper to rent a GPU server monthly or hourly? It depends on the GPU and on how many hours a month the server is allocated to you. On published 2026 rate cards, a monthly commitment saves roughly 19–20% on mid-range cards such as the A40, about 4% on an A100, and nothing on some providers’ H100, RTX PRO 6000 and B300 nodes, where the monthly price equals the hourly rate times 730.
Long-term contracts: read what is being reserved
Multi-year contracts follow the same logic as monthly ones, only more visibly.
Vultr publishes per-GPU on-demand rates next to prepaid contract rates. On the NVIDIA L40S, on-demand is $1.671/GPU/hr and 36-month prepaid “starts at $0.848/GPU/hr” — a 49.3% discount. On the AMD MI300X, $1.850 against a 24-month prepaid rate from $1.750 — 5.4%. On HGX H100, 36-month prepaid starts at $1.490/GPU/hr, so a long enough term still buys a lower rate there. But on AMD’s newest parts the published numbers point the other way: MI325X is $2.000 on demand and from $2.100 on 24-month prepaid; MI355X is $2.590 on demand and from $2.650 on 48-month prepaid. Vultr’s page quotes prepaid rates as “starts at” and does not say the configurations are identical, so read this as a signal — long-term pricing on the newest accelerators is not built around discounting — rather than proof that prepaying costs more for the same server.
Lambda shows the same shape on a different product line. Its on-demand NVIDIA B200 SXM6 instance is $6.69/GPU/hr. Its 1-Click Clusters on HGX B200, sold on commitments of two weeks to one year, are $9.86/GPU/hr at 16 GPUs, $9.36 at 64 and $8.87 at 256+ — 33–47% above the on-demand rate (Lambda pricing page, checked 7 October 2026). These are different products: a cluster is interconnected multi-node capacity, an instance is a single server. Lambda reserves its “lowest prices” for terms over a year, which go through sales. The published commitment price, in other words, buys guaranteed interconnected capacity — not a discount.
The operational consequence: if you are quoted a long-term contract on top-tier GPUs at or near the on-demand rate, do not evaluate it as a discount. Evaluate it as the price of having the hardware at all, for a stated term. That is a different decision from a volume commitment on dedicated server pricing, where committing usually lowers the unit rate.
The line items the GPU price card does not show
Setup fees. On all four of OVHcloud’s GPU dedicated server configurations, the public page lists an installation fee equal to one month’s price: $1,145 on Scale-GPU-1, $1,180 on Scale-GPU-2, $1,216 on Scale-GPU-3 and $4,355 on HGR-AI-2. Without any adjustment, a twelve-month first year is a thirteen-month invoice — $56,615 rather than $52,260 on the HGR-AI-2, 8.3% above the headline. The page does not say whether a 12- or 24-month term changes the fee, so ask: the setup fee is one of the most negotiable lines in a bare-metal quote. Hetzner, for comparison, removed the GEX131 setup fee from 16 December 2025.
Egress. GPU workloads are usually described as ingest-heavy, which makes traffic allowances look generous. Inference is the opposite. Latitude.sh includes 20 TB out with inbound unmetered; Atlantic.Net includes 10 TB a month; Cherry Servers includes 30–100 TB depending on plan; Hetzner’s GEX45 page states traffic is “unlimited and free of charge”, though the GEX131 announcement does not state terms — confirm per model. A serving endpoint that returns images or audio will hit a 10 TB allowance much sooner than a training job. The rules are the same as for any bare-metal fleet — see metered vs unmetered bandwidth.
Power, if you own. The RTX PRO 6000 Blackwell Max-Q has a maximum power draw of 300 W (NVIDIA). At full load for 730 hours that is 219 kWh. At the EU average non-household electricity price of EUR 18.37 per 100 kWh in the second half of 2025 (Eurostat, 8 May 2026), the card alone costs EUR 40.23 a month — from EUR 16.38 in Finland (EUR 7.48/100 kWh) to EUR 55.89 in Ireland (EUR 25.52/100 kWh). That is the card only. A whole GPU server — CPUs, hundreds of GB of RAM, NVMe, fans, network — typically draws two to three times the card’s figure, and the facility adds its PUE on top. At 2.5 times the card and a PUE of 1.5, the bill is roughly EUR 150 a month, about 17% of the GEX131 rental price. Meaningful, but not what decides rent versus own. The capital does — and so does memory.
Memory, if you buy in 2026. TrendForce expects server DRAM contract prices to rise 13–18% quarter-over-quarter in 3Q26, noting that several US cloud providers have multi-year long-term agreements “which restrict suppliers from raising prices for these clients”, and that “the primary source of server DRAM price increases will shift toward customers without LTAs” (TrendForce, 9 July 2026). A GPU server carries hundreds of gigabytes of system DRAM. If you are buying one this year, you are most likely a customer without an LTA — the part of the market absorbing the increase. The provider renting you the same server may not be.
How to run the rent-or-own calculation
- Measure allocated hours, not busy hours. Hourly billing on dedicated hardware runs from the moment a server is provisioned until it is released — not only while the GPU is computing. Busy hours lower your bill only if you actually release the server between jobs, and on top-tier cards you may not get it back. Track both over 30 days: allocated hours drive the bill; GPU-busy hours tell you whether you are oversized. Leaseweb’s guidance on sizing is blunt: “Below 50% means you are paying for capacity that sits idle” (Leaseweb, 2026).
- Compare allocated hours with the break-even for your own quoted rates. Below it, hourly billing wins.
- Price the full first year, not the monthly rate: setup fees, IP addresses, remote hands and any support tier that is mandatory for production.
- Price egress at your real inference volume, not your training ingest.
- Compare ownership over the hardware’s service life, typically three to five years: hardware, rack space and power, remote hands, spares and the engineering time to keep drivers and CUDA current. Twelve months of rent is a useful sanity check — EUR 10,668 for a GEX131, $14,760 for Latitude.sh’s single H100. If the hardware costs several times that, owning only pays off if you will run it hard for years — the same test behind most cloud repatriation decisions.
- Check whether the GPU is available on demand at all. If it is not, you are in the capacity market and the comparison is access versus waiting, not price versus price.
- Re-price at renewal. Every pattern above depends on current supply, and supply moves.
How many hours a month justify a monthly GPU server? Divide the quoted monthly price by the quoted hourly price. On the 2026 rate cards checked, break-even lands between 583 and 732 hours of a 730-hour month — roughly 80% to 100%, counted in hours the server is allocated to you, not hours the GPU is busy.
What breaks in production
The same spec sheet, two different GPUs. The RTX PRO 6000 Blackwell family ships in three variants with identical 96 GB GDDR7 ECC memory and maximum power of 600 W (Workstation Edition), up to 600 W configurable (Server Edition) and 300 W (Max-Q) (NVIDIA). Two providers can both list “96 GB RTX PRO 6000” and deliver materially different sustained throughput. Ask which variant, in writing, before benchmarking.
Reserved capacity left idle. A server reserved to guarantee availability costs the same whether or not the job is running. Reserved scarce capacity punishes idling harder than on-demand does.
Training-data egress discovered at cutover. Moving a training corpus out of object storage to a new GPU provider is an egress event priced by the storage side. Budget it from the cloud egress fees comparison before scheduling the move, and keep the corpus close to the accelerators — the argument for co-locating S3-compatible object storage with compute.
Driver and CUDA pinning on bare metal. A dedicated GPU server is yours to maintain, including the kernel-module, driver and framework chain. A working image on one hardware generation is not portable to the next without a rebuild.
Provisioning time treated as zero. Cloud GPUs appear in minutes. Dedicated servers with in-demand cards are often pre-order or waitlisted — Cherry Servers lists its A40 and A100 as pre-order and several other cards as waiting-list only. A plan that assumes same-day delivery of eight current-generation accelerators will slip.
One utilisation number for a mixed fleet. Averaging a training cluster at 95% with an inference fleet at 20% hides the fact that they need opposite billing models.
GPU servers at INXY.hosting
We split our GPU line the same way the market does. Every configuration below is a single-tenant bare-metal server — no virtualisation, no shared card.
Netherlands
| GPU | GPU memory | Tier | Price |
|---|---|---|---|
| NVIDIA RTX A4000 | 16 GB | Entry-level | from $525/month |
| NVIDIA RTX PRO 4000 | 24 GB | Entry-level | from $600/month |
| NVIDIA RTX A5000 | 24 GB | Mid-range | from $660/month |
| NVIDIA RTX PRO 6000 | 96 GB | High-end | from $1,750/month |
| NVIDIA H100 | 80 GB | High-end | from $1,995/month |
Monthly starting prices for a single-GPU configuration in our Netherlands location.
NVIDIA H200 and B300. All GPUs are currently in use. Reserve your spot to get access as soon as one frees up. For large projects, please contact us directly.
North America
Bare-metal GPU servers are also available in North America. Contact our team for current configurations and pricing.
Order a GPU server, reserve an H200 / B300 or talk to our team
The decision, framed
Three questions, in this order.
Is the GPU you need available on demand, in the quantity you need? If not, you are buying capacity. A price at or above the on-demand rate is the cost of certainty; evaluate it as insurance with a stated term, and get the renewal rate in writing.
Are your allocated hours above the break-even on your own quote? If not, stay hourly. If they are, a monthly or longer commitment is defensible — worth around 19–20% on mid-range cards, often little or nothing on top-tier accelerators, and the difference is visible on the provider’s own page in two minutes of arithmetic.
Does owning beat renting over the hardware’s service life? If yes — and you have somewhere to rack it and someone to maintain it — owning can win over several years. But you are buying system memory in a market where, on TrendForce’s reading, buyers without long-term agreements absorb the price rises, so use this quarter’s component prices, not last year’s.
Most teams end up with a split: committed capacity for the steady baseline, hourly for the spikes. That works best when baseline and burst do not have to come from the same rate card. If you have 30 days of allocated-hours data and a target configuration, that is enough to run the break-even against current quotes — see our dedicated and bare-metal hosting.
FAQ
Is a monthly GPU dedicated server cheaper than paying by the hour?
Only above the break-even hours. On published 2026 rate cards, a monthly plan saves about 19–20% on mid-range cards such as the A40, 14.5% on Hetzner’s RTX PRO 6000 Max-Q server, about 4% on an A100, and nothing on some providers’ H100, RTX PRO 6000 and B300 nodes. Below roughly 80% of the month in allocated hours, hourly billing is usually cheaper.
What is the minimum utilisation that justifies a dedicated GPU server?
Divide the quoted monthly price by the quoted hourly price. Across the providers checked, break-even is 583 to 732 hours of a 730-hour month — roughly 80% to 100%. Count hours the server is allocated to you, not hours the GPU is busy: hourly billing runs until you release the server.
Why don’t long-term contracts on H100, H200 or B300 come with a big discount?
Because the hardware is in short supply. Providers discount commitments when they carry idle-capacity risk; on top-tier accelerators they rarely do. Some long terms still lower the rate — Vultr’s 36-month prepaid HGX H100 starts at $1.490/GPU/hr — but on the newest parts published commitment prices often sit at or above on-demand, as with Lambda’s B200 clusters at $8.87–9.86/GPU/hr against $6.69 for an on-demand instance. What the commitment buys is guaranteed access.
What is included in the price of a GPU dedicated server?
It varies sharply. Among the plans compared, included traffic ranged from 10 TB to unlimited, setup fees from zero to a full month’s price, and management from none to a full managed security stack — which is most of why one H100-class server costs $1,230 a month and another $4,729.44.
Does owning GPUs beat renting them?
Only if you will run the hardware hard over its three-to-five-year life and have somewhere to host and maintain it. Electricity matters but does not decide it — roughly EUR 150 a month for a single-GPU server at EU average prices. The capital cost and today’s memory prices are what decide it.
How much bandwidth does a GPU server need?
Training is ingest-heavy and usually fits within included allowances; inference that returns images, audio or video is egress-heavy and often does not. Size the allowance against serving volume, not training volume, and check whether the provider meters inbound traffic as well as outbound.

