BlazingCDN racks dedicated GPU nodes in the same data centres as our delivery network. You install your own AI stack and keep root. We supply, patch and monitor the hardware, and the frameworks, the models and the weights stay yours.
Built for teams whose cards are busy most of the month. If your work comes in bursts, an hourly cloud is cheaper, and you will hear that from us on the call rather than after the contract.
Reply within 3 business hours, from an engineer. No sales sequence.
Worth saying plainly, because the two are often sold as one thing and they are not.
The machine and everything under it: GPUs, CPU, memory, NVMe, transit and the data centre. We rack it, patch it, monitor it and replace it when hardware fails. On a self-managed build you get root and a node that behaves like the rest of your infrastructure.
We do not train, fine-tune or serve your models for you. There is no managed notebook, no model catalogue and no per-token price list. Your data, your weights and your choice of framework stay inside your own environment. We operate the hardware underneath it and have no reason to touch what runs on top.
This is custom infrastructure ordered per build. It is not a feature of the CDN and it is not bundled with delivery.
When AI work ends in a file somebody has to move, an image, a rendered clip, a training set or a model artefact, every vendor boundary that file crosses gets billed and the clock keeps running.
Generate where the output already needs to live. An image, a rendered clip or a model artefact does not have to travel from a compute vendor to a delivery vendor before anyone can use it.
Datasets are large and they have to reach the card before anything trains. Pulling them across a vendor boundary is billed twice, once on the way in and again when the results come back out.
When a launch is late, the compute and the delivery are on one contract with one support channel, so nobody spends the morning deciding whose incident it is.
Whatever you already use. A self-managed node with root access does not care which framework you picked, and neither do we.
PyTorch, JAX or TensorFlow, built the way your team builds it, with the CUDA version your code was tested against rather than the one a platform decided to standardise on.
vLLM, TensorRT-LLM, Triton or a service you wrote yourself. The node is a machine on your own network, so the model sits behind your gateway and your auth rather than somebody else’s.
The unglamorous half of the work. Embedding a corpus, cleaning and converting a dataset, generating synthetic samples: all of it on the same hardware instead of a second bill.
The same class of card handles rendering and video encoding. If that is the larger half of your workload, the same node does it, and Video Content Processing Servers is the page written for it.
BlazingCDN sells dedicated GPU servers, not GPU-hours. Each node is a single-tenant machine racked in the same data centres as our CDN delivery nodes and billed as a fixed monthly figure. Self-managed by default: you hold root and run your own stack. A managed build is also possible, and there the access model is written into the agreement rather than assumed. Most AI builds land on the NVIDIA RTX PRO 6000 Blackwell Server Edition, because its 96 GB per GPU is usually what decides whether a model fits on one card or has to be split across several.
192 GB of GPU memory in one box: a production inference service, or fine-tuning runs that do not need a cluster.
384 GB of GPU memory. For serving several models at once, or for training that fits on a single node and would sit in a queue on most shared clouds.
Large model training and high-concurrency inference, where the fabric between the cards matters as much as the cards do.
Real-time inference next to your audience rather than in a distant region. Deploys into the same locations as your cache nodes.
A fifth build sits beside these: 5U, two RTX PRO 6000 Blackwell cards, 64 CPU cores across two sockets, 1.5 TB of memory and 960 GB of flash, for inference where host memory matters as much as GPU memory. All of them ship pre-configured and pre-tested rather than assembled to order, which takes the build queue out of the lead time. Configurations outside these five are specified per build. GPU pricing moves with the AI market month to month, so we quote the current rate rather than print one that goes stale. The same hardware line, alongside our cache and storage builds, sits on Custom Enterprise CDN Infrastructure.
Whatever your models produce has to live somewhere and then reach users. Both of those steps are priced openly, so you can add them up before you talk to us.
Your job finishes a batch on local NVMe. Nothing has left the building yet, and nothing has been billed for movement.
Cloud Storage runs at $0.015 per GB per month. If the artefact should stay resident in the network instead, permanent cache is $0.02 per GB per month with the first 500 GB included.
Delivery runs on the standard tiers, $5.00 down to $2.50 per TB. A month of 100 TB comes to $415. Put your own volume through the cost calculator, or read the full tiers on pricing.
Only the node itself is quoted per build. Everything downstream of it is on the published price list, which is the line that usually surprises people on a hyperscaler invoice. It is also one contract and one support channel, so when a launch is late nobody spends the morning deciding whose incident it is.
Two ways of buying the same silicon. Which one is cheaper depends almost entirely on how many hours a month the card is actually busy.
The card is busy most of the month, the data is large enough that moving it hurts, or the output has to reach users from the same network that produced it. Then a fixed monthly figure beats a meter, and the hardware stops being somebody else’s queue.
Your work comes in bursts. A week of experiments, then nothing for a fortnight. Paying by the hour is the right answer there, and a dedicated node standing idle is the wrong one. You will hear that from us on the call rather than after the contract.
Not sure which of the two you are? Tell us what the workload looks like and we will say so plainly.
If what you need is one dedicated machine for delivery rather than compute, that is Individual CDN Servers. If the workload is video encoding, it is Video Content Processing Servers.
The questions that come up on every scoping call about GPU hardware.
On a self-managed build, you do. We hand over a node with root access and a working network, and the drivers, the container runtime and the model stack are yours, which is also why we do not charge you for them. A managed build is also possible, and there the split of work and the access model go into the agreement rather than being assumed. If you would rather receive the node with drivers already in place, say so during scoping.
That depends on the model, the quantisation and the batch size, and any number quoted without those three is marketing. Bring the model and the latency you need to the call, and we size against them.
A fixed monthly figure, quoted per build, following the GPU count, the CPU, the memory and the term. GPU pricing moves with the AI market, so we quote the current rate instead of publishing a stale one.
Fine-tuning and single-node training, yes, up to an eight-GPU HGX B300 baseboard with 400GbE or XDR800 fabric. Beyond that, a run spread across hundreds of interconnected GPUs in one fabric is a different kind of build and a different kind of vendor. We will tell you that on the call rather than sell you a node and hope.
Set per build, because it follows the hardware and the data centre commitment behind it. Whatever it turns out to be, it is in the quote before anything is ordered.
In the data centres where our delivery nodes already are, which is the point of buying compute here rather than anywhere else. Our network runs across Europe, Asia and North America, and which of those data centres can take a GPU build depends on the card. Tell us the region you have to be in and the residency rule you have to meet, and we confirm availability before the quote goes out.
No. One tenant, bare metal, no hypervisor and no neighbours. The GPUs, the CPU, the memory and the NVMe are yours for the term, and nobody outside your team holds an account on the machine.
Not the way it covers delivery. What we publish is written for the CDN service. What we commit to on a dedicated node, monitoring, response time and hardware replacement, goes into your order, so you can read it before you sign rather than infer it from a badge on a page.
It depends on the card and the location. The four configurations above ship pre-assembled and pre-tested instead of waiting on parts, so what is left is transit, racking and networking rather than a build queue. A data centre where we already hold capacity moves faster than one waiting on a specific baseboard. You get a date in the quote, before you commit, and if that date moves we tell you rather than let it slide.
Memory and topology. The RTX PRO 6000 Blackwell Server Edition carries 96 GB of VRAM per card, against the 80 GB an H100 usually carries, so a model that spills over on one card may fit on the other. What it does not have is NVLink: the cards talk over PCIe Gen 5, which suits serving several models side by side and training that fits on one node, and does not suit a single model sharded across many cards. That is what the HGX B300 baseboard is for.
If the output is served through the CDN, it is the standard delivery tiers: $5.00 down to $2.50 per TB, with 100 TB working out at $415 a month, and no per-request fee. Transit for traffic the node serves directly is part of what the node is quoted on, so it sits in the monthly figure rather than on a meter. Either way there is no separate compute egress rate.
The hardware is ours, so the replacement is ours. We monitor the node, and a failed GPU, drive or power supply is swapped by us rather than by you or by a queue you have to chase. The replacement window is set per build and written into the quote, because it follows the data centre and the part.
Tell us what the model is and how busy it will be. You get a reply within 3 business hours, from an engineer rather than a sequence.
He reads every request sent from this page and replies within 3 business hours, including the reply that says a general cloud instance would serve you just as well. No sales sequence in between.