Saturn Cloud, an AI token factory platform, has integrated NVIDIA Run:ai, the workload and GPU orchestration software, into its core infrastructure. This integration allows neocloud and AI factory operators to convert GPU hardware historically rented on an hourly basis into multi-tenant, per-token AI inference services under their own custom branding.
Building on Saturn Cloud’s earlier integration of the NVIDIA DSX AI Factory Platform, this collaboration helps infrastructure providers maximize revenue density per megawatt without requiring additional hardware investments.
Unlocking Higher Revenue Streams from Existing Infrastructure
By pairing NVIDIA Run:ai’s orchestration capabilities with Saturn Cloud’s commercial management layer, operators can offer a diverse portfolio of AI services from a single compute fleet.
The platform equips cloud providers with turnkey solutions for:
- Per-Token Model-as-a-Service (MaaS): Delivering accessible API endpoints for clients who require pay-per-token consumption.
- Dedicated GPU Capacity: Serving enterprise customers needing isolated infrastructure with custom software stacks.
- Managed Fine-Tuning Environments: Offering self-service development spaces for teams building, training, and fine-tuning proprietary models.
The platform handles operational overhead, including model onboarding, tenant isolation tiers for regulated industries, and identity and access governance required by enterprise buyers.
Also Read: Parallel Works and CoreWeave Partner to Deploy Fully Managed AI Cloud Environment for DARPA Biological Research
“A neocloud can raise revenue per megawatt without adding a single GPU, just by changing how the capacity is sold,” said Sebastian Metti, Founder of Saturn Cloud. “Saturn Cloud gives operators the full commercial platform to do it, handling serving, multi-tenancy, and billing, all running under their own brand on NVIDIA Run:ai.”
Technical Architecture and Fleet Orchestration
Underneath the commercial layer, the joint solution relies on an enterprise-grade technology stack designed for high throughput and hardware reliability:
- Inference Serving: Utilizing NVIDIA Dynamo, the framework performs distribution for inference using disaggregated prefill and decode stages in conjunction with execution engines such as vLLM, SGLang, or NVIDIA TensorRT-LLM.
- Scheduling & Allocation: Multi-node orchestration is performed by NVIDIA Grove, whereas NVIDIA KAI Scheduler offers GPU-aware scheduling, GPU fractional allocation, gang scheduling, and cross-tenant quota management.
- Remediation: Health of hardware is managed by NVSentinel and NVIDIA Fleet Intelligence that detects and pulls the degraded hardware from the placement automatically to preserve service quality.
“Cloud providers are looking for new ways to turn GPU infrastructure into differentiated AI services,” said Omri Geller, NVIDIA VP DSX OS Platform Software. “Together, NVIDIA Run:ai and Saturn Cloud help operators improve GPU utilization while delivering scalable, multi-tenant inference services under their own brand.”


