A neocloud is a cloud provider built around renting GPU compute for AI, rather than a general-purpose cloud that sells GPUs as one service among many. What it sells is GPU servers and virtual machines, and its pitch against the hyperscalers is focus and price. The trade-off is a much shorter list of everything else: fewer products, fewer capabilities and fewer regions.
How the term is defined#
Three research sources define the term in almost the same words.
| Source | Published | Definition |
|---|---|---|
| A semiconductor and AI research newsletter | October 2024 | "a new breed of cloud compute provider focused on offering GPU compute rental" |
| A data-center research institute | February 2025 | Providers that "focus on offering GPU-backed servers and virtual machines, often at prices more affordable than those of the hyperscalers" |
| A cloud market research firm | October 2025 | "new or emerging specialized cloud computing platforms that provide high-performance, GPU-centric infrastructure, primarily to support artificial intelligence (AI) workloads" |
They agree on the core: a neocloud's main product is GPU compute for AI, sold as servers or virtual machines, rather than a broad catalog of cloud services. The market research firm names three main offerings: GPU as a service, generative AI platform services and high-capacity data centers. GPU as a service explains the first and how it is priced.
The category is growing fast. The same firm counted more than $25 billion of neocloud revenue in 2025, $9 billion of it in the fourth quarter alone, up 223% on a year earlier, and it forecasts the market to approach $400 billion by 2031.
What a neocloud sells#
Under the product sits a GPU cluster, and the October 2024 analysis describes its parts. A back-end compute fabric carries GPU-to-GPU traffic "from tens of racks to thousands of racks". A front-end Ethernet network connects the servers to the internet, to the scheduler, such as Slurm or Kubernetes, and to networked storage. An out-of-band management network re-images the servers and watches their fans, temperatures and power draw. The same analysis notes that each 8-GPU H100 server has one customer at a time, so bare metal is possible and common, while virtual machines recover faster when a server breaks.
Customers see one of three shapes: a virtual machine with one or more GPUs, a whole bare-metal server, or a cluster of servers joined by that back-end fabric. QuantaCloud sells the first on demand and builds the other two to order.
How neoclouds differ from hyperscalers#
The difference is focus. A hyperscaler sells GPUs as one service in a broad catalog of compute, storage, databases and software, and a neocloud sells GPU compute as its main product. The market research firm puts it as a tight focus on GPUs "rather than offering a broad portfolio of cloud services".
| Neocloud | Hyperscaler | |
|---|---|---|
| Main product | GPU compute for AI: GPU servers and virtual machines | A broad catalog of cloud services, with GPUs as one of them |
| Product lines | "only a handful", in the February 2025 report's words | Broad |
| Average on-demand price of an 8-GPU H100 system, February 2025 | $34 an hour | $98 an hour |
| Regions | Limited, per the same report | Broad |
| Where the report expects enterprises to use it | Large-scale AI work such as training | Core infrastructure, inference included |
The price gap is the headline. In the February 2025 comparison, an 8-GPU H100 system cost $98 an hour on demand at the hyperscalers on average and $34 at the neoclouds, a 66% saving. The same report is clear about the trade: because neoclouds "have limited products, capabilities and regions", it expects most enterprises to keep their core infrastructure, inference included, on hyperscalers and their own hardware, and to use neoclouds for large-scale AI work such as training.
My read is that the split follows where the rest of your stack lives. A job that needs GPUs, a network port and little else, such as training a model or serving one behind an API, fits a GPU cloud. Anything tied to a hyperscaler's databases, identity and networking stays where it is.
How to evaluate a neocloud#
Judge a neocloud on what you can launch today, what it costs in full, and what happens when something goes wrong. These are the points I would check with any GPU cloud, with QuantaCloud's own answers in the last column, including the ones that are a no.
| What to ask about | Why it matters | QuantaCloud's answer |
|---|---|---|
| GPUs you can launch now | A listing is not the same as capacity you can start today | A live catalog from the 48 GB RTX A6000 to the 141 GB H200 NVL, with 1, 2, 4 or 8 GPUs per VM |
| The next size up | Whether the provider can grow with you | B200, B300, 8-GPU H100 or H200 servers and InfiniBand clusters, built to order with the lead time and terms in writing |
| Billing | Granularity and minimums change the real cost | Prepaid credit from $5. The first hour is charged at launch and unused seconds are refunded on stop |
| Data when an instance stops | Checkpoints, models and datasets | Stop terminates the instance and deletes its disk. There are no volumes or snapshots |
| Data transfer | Moving datasets and checkpoints out | No egress or ingress charges |
| Interconnect | Multi-GPU and multi-node jobs | On demand, single VMs: check NVLink with nvidia-smi topo -m. Reserved clusters use InfiniBand or RoCE between servers |
| Regions | Latency and data residency | US only: us-east-1 in Virginia, and us-midwest-1, us-midwest-2 and us-midwest-4 |
| Access | Your drivers, containers and scripts | Ubuntu 22.04 VMs with SSH as ubuntu and Docker, so your containers and scripts move with you |
| Support and service commitment | Whether production traffic belongs there | Email and in-console tickets. No SLA |
| Certifications | Procurement and compliance | Not SOC 2 or HIPAA certified |
Two rows deserve the most weight for AI work. Storage decides how you work day to day: on a cloud where stopping deletes the disk, as on QuantaCloud, checkpoints and models have to live somewhere else and every session starts with a download. The service commitment decides whether production traffic belongs there at all. Pricing covers the billing rules in full, and InfiniBand vs NVLink explains the interconnect row.
Is QuantaCloud a neocloud#
By these definitions, yes. QuantaCloud is a standalone GPU cloud whose product is NVIDIA GPU compute and not much else. On demand it sells GPU virtual machines with 1, 2, 4 or 8 GPUs, from the 48 GB RTX A6000 to the 141 GB H200 NVL, in US regions, paid from prepaid credit, with templates for ComfyUI, Jupyter and Open WebUI. For committed work, we build reserved capacity to order, from one dedicated server to an InfiniBand cluster. It does not sell databases, object storage, managed Kubernetes or model APIs, and it publishes no SLA.
The two jobs it is built for are serving open models and fine-tuning them, on GPUs you pick yourself.
Neocloud questions#
What does neocloud mean?
A cloud provider whose main product is GPU compute for AI, sold as servers or virtual machines. The word separates GPU-focused providers from the hyperscalers, whose catalogs cover every kind of cloud service.
Are neoclouds cheaper than hyperscalers?
For on-demand GPU prices, the published comparison says so: $34 an hour against $98 for an 8-GPU H100 system, a 66% saving. Those prices are from February 2025, so compare today's hourly rates together with the billing rules and what each price includes.
Is a neocloud the same as GPU as a service?
GPU as a service is the main thing a neocloud sells. The term describes the product, renting GPU compute by the hour or for a term, and neocloud describes the kind of provider that specializes in it.
Are neoclouds only for training?
No. The February 2025 report expected enterprises to use neoclouds mainly for large-scale work such as training, but serving a model needs the same GPUs, and an inference server needs little besides a GPU and a network port.
My rule: if a job is mostly GPU hours, price it at a GPU cloud first, and keep the services that need a broad catalog, such as databases and identity, where they already run. Check the ten points in the evaluation table with every provider before you commit, then see the GPUs QuantaCloud offers with live prices.