NVIDIA hardware for inference, rendering and video work, racked inside the network that already delivers your traffic. One tenant per machine, on a monthly term, quoted per build. No hourly meter.
Reply within 3 business hours. No sales sequence.
The trigger is where the work sits and how often it runs, not how many GPUs you think you need. If one of these describes your case, dedicated hardware usually beats renting capacity by the hour.
Renders, renditions and processed images end up in front of an audience. Compute inside the delivery network means the file does not cross a vendor boundary before anyone can watch it.
Per-hour pricing is honest until utilisation climbs. Past a certain point a fixed monthly figure for hardware you keep is the cheaper of the two, and it stops moving with the market.
Capacity that appears when somebody else releases it is fine for experiments and expensive on release week. A dedicated node is there at 3am because nobody else is on it.
GDPR, sector regulation or a contract that names the country your data may be processed in. The data centre is named in the quote instead of described as a region.
Root access, your drivers, your containers, your scheduler. Self-managed builds hand over hardware and a network, not a platform your team has to learn first.
A fanless edge module deploys into cache locations, so inference answers inside the region instead of travelling to a distant one and back, which is the whole point of putting it there.
None of these describe you? Then a general GPU cloud is the better tool and we will say so on the call. We are not the cheapest place to run a notebook for an afternoon.
Same job, same output, same audience. What changes is where the GPU lives, and that is what the bill and the waiting are made of.
metered by the hour
Cheapest to start and the right answer for experiments. You also get the queue, the neighbours, and a meter that keeps running while the job waits for a card.
inside the delivery network
The machine is yours for the term and it sits inside the network that serves the output, so finished work does not leave one vendor to reach another.
A dedicated node carries the steady load, hourly capacity absorbs the spikes. Teams with a real production workload and an unpredictable peak usually end up here.
You get a recommendation back, including the recommendation to stay on an hourly cloud.
NVIDIA hardware, from Blackwell-class server cards down to a fanless edge module. GPU count, CPU, memory and NVMe are specified per build, because the right machine follows the workload rather than a plan tier.
Agent-flow inference, video encoding and graphics workloads. The entry point for teams putting a first model into production.
Multi-tenant inference at PaaS scale, plus rendering and scientific computing on the same node.
Large model training and high-concurrency inference. The storage gap settles the choice more often than the CPU does: 960 GB against 61.4 TB.
Real-time inference next to your audience rather than in a distant region. Deploys into the same locations as your cache nodes.
Why we do not publish GPU prices. The AI market moves hardware pricing month to month. A number printed here would be wrong by the time you read it, so we quote the current rate when you ask, and that rate is what you pay.
A fifth build sits beside these: 5U, two Blackwell cards, 64 CPU cores across two sockets and 1.5 TB of memory, with 960 GB of flash. It is for inference where host memory matters more than local storage, and that 960 GB is worth noticing next to the 7.6 TB on the 2U and 4U. Larger multi-GPU builds and mixed CPU and GPU racks are specified per order. Only the machine itself is quoted: storage and delivery for whatever it produces run on the published CDN tiers, which you can add up yourself on the pricing page.
Pick the configuration first, then decide who operates it. Both modes run on the same network, in the same data centres, on the same hardware.
We design the build, deploy it, then monitor it, patch it and replace hardware when it fails. You get capacity, dashboards and an operations team that already runs this network at scale. Support hours are set in your agreement.
Choose this when your team should own the model, not the racks.
We deliver the machine, the network and the transit. Your team keeps root access and full control over drivers, containers, schedulers and release process. We operate the network underneath and stay out of the box. Failed hardware is still replaced by us in both modes, and the replacement window is written into the quote.
Choose this when you already run infrastructure and need capacity in the right places.
Model serving, fine-tuning and full training runs. Frameworks, images and schedulers are yours: the node arrives without an opinion about them. Training or inference? Those builds, the frameworks they run and what the output costs after the GPU are covered on dedicated GPU servers for AI inference and fine-tuning.
Encoding, transcoding, packaging and image pipelines that feed the delivery layer. If encoding is the job and the rest is secondary, start from Video Content Processing Servers.
3D rendering, VFX passes, scientific and engineering simulation. Long jobs that need the whole machine rather than a share of one, and that nobody wants to restart because a shared instance was reclaimed.
An engineer maps the workload, the frameworks, the regions and the peak pattern. Bring the job you actually run, not a GPU count. No slide deck.
Hardware list, locations, term and a fixed monthly figure, in writing, before anything is ordered. The rate is the one live that week, not one printed on a web page months ago.
The node is racked, networked and delivered with the access model you asked for. You get a date with the quote, because it follows hardware availability and the data centre.
Managed mode: monitoring, patching, capacity planning and hardware replacement. Self-managed mode: root access, documentation and a handover call, then we stay out of the box.
GPU builds are quoted individually, so the honest answer to several of these starts with “it depends”. Here is what it depends on.
A fixed monthly figure, quoted per build, following the GPU count, the CPU, the memory, the NVMe and the term. GPU pricing moves with the AI market month to month, so we quote the rate that is live when you ask rather than print one that goes stale.
No. These are dedicated machines on a monthly term, which is exactly what makes the figure predictable. If your workload is a few hours a week, an hourly GPU cloud is the better tool and we will tell you so on the call.
NVIDIA. RTX PRO 6000 Blackwell Server Edition in the 2-GPU and 4-GPU builds, an HGX B300 baseboard in the 8-GPU build, and Jetson Orin NX for fanless edge nodes. Other cards are sourced per build when the workload calls for them.
In self-managed builds, yes. The machine, the network and the transit are delivered and everything above them is yours. Managed builds are operated by us, and the access model is written into the agreement rather than assumed.
Inside our delivery footprint, which is the whole reason to put compute here. Not every site takes every build, because an 8-GPU baseboard needs power and cooling that a small cache location does not have, so the exact data centre is named in the quote rather than promised in advance. If your constraint is that data must not leave a jurisdiction, say so on the call and the answer comes back as a named location.
Not automatically. What we publish covers CDN delivery, not a dedicated machine. What is covered for a GPU build, and how a claim is raised, is written into that build’s agreement. Ask for it in the quote and read it before you sign.
Set per build, because it follows the hardware and the data centre commitment behind it. Whatever it turns out to be, it is in the quote before anything is ordered.
Often, yes, and we will tell you when it would. The case for a node here is adjacency: the output lands inside the network that serves it. If nothing you produce is delivered to an audience, that adjacency is not worth paying for.
Bring the job, the frameworks it runs and the regions it has to run in. You get a build specified against those, or an honest answer that renting by the hour is the cheaper way to do what you described.
He reads GPU enquiries personally and replies within 3 business hours, including the reply that says a dedicated node is not what your workload needs. No sales sequence in between.