Saturn Cloud Token Factory

Turn GPU capacity into tokens you can bill for.

Saturn Cloud is the token factory platform for AI clouds and enterprises. It adds the tenant, metering, and billing layer on top of your GPUs, so you can serve models, meter every token, and invoice customers, on infrastructure you own and under your own brand.

Deployed into your infrastructure, public cloud, private cloud, or on-premises.

Model endpoints in Token Factory, each a live token-metered endpoint

Model endpoints in Token Factory, each one a live, token-metered endpoint.

The economics

Renting GPU hours caps what your fleet can earn.

Rent a GPU by the hour and every improvement in the stack, faster hardware, better runtimes, lower cost per token, turns into pressure to drop your hourly rate. A token factory inverts that.

Rent GPU hours

Revenue capped by the clock

You bill for time, not output. Idle capacity earns nothing, and better utilization just lets customers ask for a lower rate. Every efficiency gain works against your margin.

Sell tokens

Revenue scales with throughput

As throughput rises and cost per token falls, the same GPU serves more billable tokens. Better utilization becomes margin, not a discount, and idle capacity is exactly what the platform monetizes.

Token revenue climbing above raw GPU cost on a live endpoint, with request rate, latency, and throughput

Token revenue pulling away from raw GPU cost on a live endpoint, with throughput and latency alongside.

How it works

Serving a model is the easy part. A token factory is everything around it.

We think about the stack in three planes. The bottom two are commodity, you should build on them, not reinvent them. The top one is where a serving system becomes a product you can sell.

Tenant plane Saturn Cloud
Meters and bills. Tenant isolation, per-token and per-GPU-hour metering, quotas, rate cards, billing, scoped API keys.
Control plane commodity
Decides what runs where. Request routing, GPU-aware scheduling, autoscaling to a latency target.
Data plane commodity
Moves tokens. Serving engines and GPUs producing the output.

The data and control planes have no concept of your customers, your billing, or your quotas. They will serve any model to anyone who can reach them.

The tenant plane knows who the caller is, counts every token they use, holds them to a quota, and turns that usage into an invoice you can defend. It is the layer no open-source component ships, and it is where a serving product is actually differentiated.

Saturn Cloud is that layer, plus the platform to run it.

What's inside

Everything between an endpoint and a customer you've never met.

The full platform layer, deployed as one system on your GPUs.

01

Tenant isolation

Each tenant isolated at the fabric, not just the namespace. Holds up to the first enterprise security review.

02

Fine-tuning

Full-weight and LoRA fine-tuning on the operator's own GPUs, served alongside the base models.

03

Model serving

Inference behind the API your customers already build against, with the serving engine you choose underneath.

04

Per-token & per-GPU-hour metering

Usage counted per tenant, captured even when a client disconnects mid-stream.

05

Rate cards & billing

Versioned pricing that reproduces the price live at time of use, so invoices hold up in a dispute.

06

Quotas & rate limiting

Enforced at the edge, degrading open rather than failing closed on your largest customer.

07

Scoped API keys

Per-tenant, per-project keys, the kind enterprise buyers ask for and reject account-wide keys over.

08

SSO, RBAC & audit

The identity, access, and audit controls that get you past enterprise security reviews.

09

Usage & chargeback

Per-user and per-project visibility into GPU usage, so you bill accurately and see where capacity goes.

Fine-tuning jobs in Token Factory, full-weight and LoRA runs across statuses

Fine-tuning jobs, full-weight and LoRA, on the operator's own GPUs.

Your infrastructure, your brand

Your infrastructure. Your brand. Your customers.

Saturn Cloud installs into infrastructure you own and runs white-label, in days, not quarters. Your customers see your product, on your domain. You keep full control of your GPUs, your pricing, and the customer relationship.

  • Runs on your hardware. Public cloud, private cloud, or on-premises, wherever your GPUs are.
  • Your engines stay yours. Serving runtime, scheduling, and tuning are yours to control.
  • Data stays inside your boundary. Inference and metering run under your own security and governance.
  • The console is one client of the API. Use the white-label portal, or build your own on top.

Turn the GPUs you already
run into recurring revenue.

See how operators turn metered inference into margin,
on their own infrastructure.