SpacebusSparkplane

GPU control plane

Open weight models, one click from serving.

Sparkplane provisions vLLM on Kubernetes across your GPU cluster, hands you an OpenAI-compatible endpoint, and keeps every token on hardware you control.

Serving stack

vLLM

Orchestration

Kubernetes

API surface

OpenAI

Data path

Private

01 — Flight sequence

01

Pick a model

Llama, Qwen, Mistral, Gemma and DeepSeek weights, sized against your available VRAM.

02

We schedule it

A vLLM pod and service land on a GPU node with tensor parallelism set for you.

03

Call the API

An OpenAI-compatible endpoint plus scoped keys, ready the moment the pod reports healthy.

04

Or just chat

A built-in context window with system prompt, temperature and token controls.

02 — In the catalog

Qwen3 4B Instruct ·4BLlama 3.1 8B Instruct ·8BMistral 7B Instruct v0.3 ·7BGemma 3 12B IT ·12BQwen3 32B ·32BDeepSeek R1 Distill 14B ·14BLlama 3.3 70B Instruct ·70BMixtral 8x7B Instruct ·47B MoE

Cluster credentials are encrypted at rest and never leave the backend. Your models run on private infrastructure reached over an internal network path — the platform itself is public, the GPUs are not.