GPU CLOUD · ONE TENANT PER MACHINE

Dedicated GPU Cloud Servers

NVIDIA hardware for inference, rendering and video work, racked inside the network that already delivers your traffic. One tenant per machine, on a monthly term, quoted per build. No hourly meter.

Reply within 3 business hours. No sales sequence.

4 configurations
2-GPU and 4-GPU nodes, an 8-GPU baseboard, a fanless edge module
Inside the network
Nodes go into our delivery footprint, named by site in your quote
Root access
Self-managed builds hand over the machine, not a console
Monthly term
Quoted per build, price on request, no hourly meter
Pdfforge
Sony logo
Tapmind
Cityspark
xTool logo
Avanquest logo

When a Dedicated GPU Node Is the Right Call

The trigger is where the work sits and how often it runs, not how many GPUs you think you need. If one of these describes your case, dedicated hardware usually beats renting capacity by the hour.

Your output has to be delivered

Renders, renditions and processed images end up in front of an audience. Compute inside the delivery network means the file does not cross a vendor boundary before anyone can watch it.

The machine runs most of the month

Per-hour pricing is honest until utilisation climbs. Past a certain point a fixed monthly figure for hardware you keep is the cheaper of the two, and it stops moving with the market.

Queueing for capacity costs you deadlines

Capacity that appears when somebody else releases it is fine for experiments and expensive on release week. A dedicated node is there at 3am because nobody else is on it.

Data must stay in one jurisdiction

GDPR, sector regulation or a contract that names the country your data may be processed in. The data centre is named in the quote instead of described as a region.

You want the machine, not a console

Root access, your drivers, your containers, your scheduler. Self-managed builds hand over hardware and a network, not a platform your team has to learn first.

Inference belongs near the audience

A fanless edge module deploys into cache locations, so inference answers inside the region instead of travelling to a distant one and back, which is the whole point of putting it there.

None of these describe you? Then a general GPU cloud is the better tool and we will say so on the call. We are not the cheapest place to run a notebook for an afternoon.

Which Page You Actually Need

THE ACTUAL DIFFERENCE

Where the Machine Sits Is the Whole Decision

Same job, same output, same audience. What changes is where the GPU lives, and that is what the bill and the waiting are made of.

Hourly GPU cloud

Your data
Rented capacity

metered by the hour

Your audience

Cheapest to start and the right answer for experiments. You also get the queue, the neighbours, and a meter that keeps running while the job waits for a card.

Dedicated node here

Your data
Your GPU node

inside the delivery network

Your audience

The machine is yours for the term and it sits inside the network that serves the output, so finished work does not leave one vendor to reach another.

Node plus burst

Your data
Rented
Yours
Your audience

A dedicated node carries the steady load, hourly capacity absorbs the spikes. Teams with a real production workload and an unpredictable peak usually end up here.

You get a recommendation back, including the recommendation to stay on an hourly cloud.

Four Configurations, Built to Order

NVIDIA hardware, from Blackwell-class server cards down to a fanless edge module. GPU count, CPU, memory and NVMe are specified per build, because the right machine follows the workload rather than a plan tier.

Entry point

2U, 2 GPU

Agent-flow inference, video encoding and graphics workloads. The entry point for teams putting a first model into production.

Price on request
PaaS scale

4U, 4 GPU

Multi-tenant inference at PaaS scale, plus rendering and scientific computing on the same node.

Price on request
Training

8U, 8 GPU

Large model training and high-concurrency inference. The storage gap settles the choice more often than the CPU does: 960 GB against 61.4 TB.

Price on request
Edge

Edge AI node

Real-time inference next to your audience rather than in a distant region. Deploys into the same locations as your cache nodes.

Price on request

Why we do not publish GPU prices. The AI market moves hardware pricing month to month. A number printed here would be wrong by the time you read it, so we quote the current rate when you ask, and that rate is what you pay.

A fifth build sits beside these: 5U, two Blackwell cards, 64 CPU cores across two sockets and 1.5 TB of memory, with 960 GB of flash. It is for inference where host memory matters more than local storage, and that 960 GB is worth noticing next to the 7.6 TB on the 2U and 4U. Larger multi-GPU builds and mixed CPU and GPU racks are specified per order. Only the machine itself is quoted: storage and delivery for whatever it produces run on the published CDN tiers, which you can add up yourself on the pricing page.

Two Ways to Run the Same Hardware

Pick the configuration first, then decide who operates it. Both modes run on the same network, in the same data centres, on the same hardware.

Managed by BlazingCDN

We design the build, deploy it, then monitor it, patch it and replace hardware when it fails. You get capacity, dashboards and an operations team that already runs this network at scale. Support hours are set in your agreement.
Choose this when your team should own the model, not the racks.

Self-managed

We deliver the machine, the network and the transit. Your team keeps root access and full control over drivers, containers, schedulers and release process. We operate the network underneath and stay out of the box. Failed hardware is still replaced by us in both modes, and the replacement window is written into the quote.
Choose this when you already run infrastructure and need capacity in the right places.

Inference and training

Model serving, fine-tuning and full training runs. Frameworks, images and schedulers are yours: the node arrives without an opinion about them. Training or inference? Those builds, the frameworks they run and what the output costs after the GPU are covered on dedicated GPU servers for AI inference and fine-tuning.

Video and image processing

Encoding, transcoding, packaging and image pipelines that feed the delivery layer. If encoding is the job and the rest is secondary, start from Video Content Processing Servers.

Rendering and simulation

3D rendering, VFX passes, scientific and engineering simulation. Long jobs that need the whole machine rather than a share of one, and that nobody wants to restart because a shared instance was reclaimed.

From First Call to a Live Node

01

Scoping call, 30 minutes

An engineer maps the workload, the frameworks, the regions and the peak pattern. Bring the job you actually run, not a GPU count. No slide deck.

02

Design and quote

Hardware list, locations, term and a fixed monthly figure, in writing, before anything is ordered. The rate is the one live that week, not one printed on a web page months ago.

03

Build and hand over

The node is racked, networked and delivered with the access model you asked for. You get a date with the quote, because it follows hardware availability and the data centre.

04

Operate or step back

Managed mode: monitoring, patching, capacity planning and hardware replacement. Self-managed mode: root access, documentation and a handover call, then we stay out of the box.

BEFORE YOU ASK

Questions We Get on the First Call

GPU builds are quoted individually, so the honest answer to several of these starts with “it depends”. Here is what it depends on.

What does a GPU node cost?

A fixed monthly figure, quoted per build, following the GPU count, the CPU, the memory, the NVMe and the term. GPU pricing moves with the AI market month to month, so we quote the rate that is live when you ask rather than print one that goes stale.

Can I rent a GPU by the hour?

No. These are dedicated machines on a monthly term, which is exactly what makes the figure predictable. If your workload is a few hours a week, an hourly GPU cloud is the better tool and we will tell you so on the call.

Which GPUs do you actually supply?

NVIDIA. RTX PRO 6000 Blackwell Server Edition in the 2-GPU and 4-GPU builds, an HGX B300 baseboard in the 8-GPU build, and Jetson Orin NX for fanless edge nodes. Other cards are sourced per build when the workload calls for them.

Do I get root access?

In self-managed builds, yes. The machine, the network and the transit are delivered and everything above them is yours. Managed builds are operated by us, and the access model is written into the agreement rather than assumed.

Where can the node physically sit?

Inside our delivery footprint, which is the whole reason to put compute here. Not every site takes every build, because an 8-GPU baseboard needs power and cooling that a small cache location does not have, so the exact data centre is named in the quote rather than promised in advance. If your constraint is that data must not leave a jurisdiction, say so on the call and the answer comes back as a named location.

Does your CDN coverage include this node?

Not automatically. What we publish covers CDN delivery, not a dedicated machine. What is covered for a GPU build, and how a claim is raised, is written into that build’s agreement. Ask for it in the quote and read it before you sign.

What is the minimum term?

Set per build, because it follows the hardware and the data centre commitment behind it. Whatever it turns out to be, it is in the quote before anything is ordered.

Would a general cloud instance do?

Often, yes, and we will tell you when it would. The case for a node here is adjacency: the output lands inside the network that serves it. If nothing you produce is delivered to an audience, that adjacency is not worth paying for.

Tell Us What the Workload Actually Is

Bring the job, the frameworks it runs and the regions it has to run in. You get a build specified against those, or an honest answer that renting by the hour is the cheaper way to do what you described.

Ivan Vovk, COO at BlazingCDN
Who answers
Ivan Vovk
COO, BlazingCDN

He reads GPU enquiries personally and replies within 3 business hours, including the reply that says a dedicated node is not what your workload needs. No sales sequence in between.