How it works

Choose a model

Select any open-weight model — a popular base model from the open ecosystem, or one you bring yourself — from the Spacebus model catalog.

Auto-sized, automatically

Spacebus reads the model's architecture and automatically determines the right GPU allocation, parallelism, and precision to run it efficiently — no manual cluster math required.

Deployed on dedicated capacity

The model is provisioned onto dedicated GB300 NVL72 capacity at the Spacebus site closest to your workload — never shared, spot-priced infrastructure.

Live and ready to call

A production-ready inference endpoint is live and callable, sized to the latency and throughput your application needs.

One system, two workloads

Train & fine-tune

Launch a training or fine-tuning job on the same one-click flow — Spacebus provisions the GPU allocation your job needs and manages it through to checkpoint.

Deploy & serve

Move straight from a fine-tuned checkpoint to a live, production-grade inference endpoint, without a separate infrastructure request or manual handoff.

What makes it possible

One-click deployment isn't a thin wrapper over raw GPUs — it's a purpose-built platform layer engineered for open-weight models specifically.

Auto-configuration engine

Automatically calculates GPU count, parallelism strategy, and precision from the model's own architecture — the manual sizing work is done for you.

Production-grade serving

Built on proven, open-source inference engines, coordinated across nodes for high-throughput, low-latency serving at scale.

Dedicated, isolated capacity

Every deployment runs on infrastructure isolated at the hardware level — not a shared multi-tenant pool — consistent with Spacebus's sovereign capabilities.

Built for the open ecosystem

Works with any open-weight model family, and stays current as the open-source model ecosystem evolves.

Without Spacebus vs. with Spacebus

Standing up an open-weight model, the old wayOn Spacebus
Manually provision and configure GPU clustersAuto-configured GPU allocation, calculated for you
Hand-tune parallelism and precision settingsDetermined automatically from the model's architecture
Separate infrastructure requests for training vs. servingOne system, one flow, for both
Share capacity with other tenantsDedicated, isolated capacity at every site
Weeks from decision to productionLive in a single deployment action

Deploy your first model.

Bring an open-weight model — or your own fine-tuned version — and see it live on dedicated capacity.

Get started