Design your private AI deployment
Build a topology for a pilot, a team, or a resilient production farm. See how AI Gateway, workers, Kubernetes and licensing fit together before you create any cloud resources.
Deployment inputs
Gateway with a managed worker pool
Planning estimate, not a quote. Rates are a 21 August 2026 public-price snapshot. The estimate excludes data transfer, public IPv4, load balancers, monitoring ingestion, backups, support plans, discounts, reserved or Spot capacity, tax, and model-specific performance requirements.
Validate the design with your security and platform teams, benchmark your chosen models, confirm GPU quota and regional SKU availability, then price the final architecture in the provider calculator.
A practical starting point
Pilot
Use one CPU worker to validate networking, keys, client compatibility and operations. It proves the deployment, not production inference speed.
Team
Start with two GPU workers behind AI Gateway. You can maintain one worker while the other serves traffic, and add capacity without changing clients.
Production
Use three or more workers across failure domains, a production control-plane SLA, autoscaling guardrails, persistent model storage, TLS ingress and your standard observability stack.