AI Server

The private AI platform for teams. One server runs the models — every app and every device just uses them. AI Server serves chat, embeddings, image generation, speech and vision from hardware you own, through OpenAI-compatible APIs, to the whole AI Suite and to any compatible tool.

Start free on your own machine. When your team grows, serve your network; when one machine isn't enough, AI Gateway mode turns several servers into a single load-balanced endpoint — on your LAN, in Docker, or on Kubernetes. No vendor cloud in the data path at any step.

Get it on Microsoft Store See what it does ↓ Licensing & Pricing

WHAT AI SERVER DOES

One deployment powers everything

Without a server, every PC downloads its own models, needs its own GPU, and is configured and updated separately. With AI Server, models download once, run on the best hardware you have, and serve everyone.

Universal AI APIs

OpenAI-compatible endpoints for chat, embeddings, image generation, speech-to-text (including live streaming), text-to-speech, voice cloning, and vision — all served from your hardware. Point any compatible tool or your own code at it.

Powers the whole AI Suite

Every AI Suite app discovers the server and offloads its AI work to it — zero configuration. Notes, PDFs, translation, dictation, images, email and more, all sharing one set of models on one machine.

Governance built in

API keys with per-key plans, rate limits, quotas and budgets, usage analytics by app and model, and content-free audit logs with signed export. The controls IT expects, without a cloud vendor in the loop.

Scales with AI Gateway

Flip one switch and multiple servers become a single endpoint: health-checked load balancing, model-aware routing that prefers machines with the model already warm, canary rollouts, and zero-downtime upgrades.

Runs anywhere

A desktop app for the machine under the desk; containers, Kubernetes manifests, and a Helm chart for the rack or the cloud. Same product, same APIs, same license — from one laptop to a GPU farm.

Watchable by design

Every server hosts its own live dashboard and native Prometheus metrics: fleet health, latency, usage and capacity at a glance — no extra monitoring stack required to get started.

AI GATEWAY

From one server to a farm — without changing a single client

AI Gateway is AI Server running in its farm front-end mode. Clients connect to it exactly as if it were a single AI Server; behind it, a pool of worker servers does the actual inference. Add a machine and capacity grows; take one down for maintenance and traffic drains gracefully — nobody's request is dropped.

On Kubernetes, workers are discovered automatically as they scale, so an autoscaling farm needs no manual pool edits. The whole estate stays observable from the built-in dashboard and Prometheus metrics.

What the Gateway handles for you

  • Load balancing — round-robin, least-latency, least-connections, weighted, or sticky per client.
  • Model-aware routing — requests go to a machine that already has the requested model warm.
  • Failover — a failing worker is retried on a healthy one before the client ever notices.
  • Canary rollouts — send a consistent slice of traffic to new machines before trusting them fully.
  • Backpressure — when every worker is saturated, callers get an honest "retry shortly" instead of a timeout.
  • Zero-downtime upgrades — drain a worker, upgrade it, return it to the pool.
EDITIONS & TIERS

Start free. Grow when your team does.

The free tier is a real product, not a trial: a full local AI server for all AI Suite apps on your own machine, forever. Paid tiers begin exactly where value crosses machine boundaries.

Free

A complete local AI server on your own machine. Unlimited use by every AI Suite app on that machine; generic OpenAI-compatible clients are rate-limited.

Pro Personal

Serve your own devices across your network: LAN serving, run as an OS service, API keys for your laptop and phone, unlimited generic clients, and a small gateway pool.

Pro Commercial

Serve other people: commercial use, unlimited API keys, full governance (quotas, budgets, audit signing), and a production gateway farm with canary rollouts and automatic worker discovery.

Enterprise

Unlimited estates, organisation enrollment with AI Admin Console, offline and air-gapped operation, and priority support. Licensed per estate — talk to us.

WHERE IT FITS

One private AI platform, several deployment patterns

Centralize expensive models and governance while each team keeps using familiar apps and OpenAI-compatible tools.

Office and knowledge teams

Share approved chat, translation, rewriting, document RAG, image and speech models from one managed machine instead of downloading a copy to every workstation.

Read the deployment pattern →

Regulated and sensitive workloads

Give legal, healthcare, finance, government and research teams useful AI while prompts, files, model access and audit controls remain inside the organisation's chosen boundary.

Read the deployment pattern →

Developer and internal API platform

Point OpenAI-compatible applications, agents and CI evaluation tools at an internal endpoint. Platform teams control models and capacity without embedding a vendor key in every project.

Read the deployment pattern →

Branches, failover and GPU farms

Put AI Gateway in front of workers across machines or failure domains for health-aware routing, warm-model affinity, rolling maintenance and capacity growth without client reconfiguration.

Read the deployment pattern →
DEPLOY IN YOUR CLOUD

Certified delivery for Azure and AWS

AI Server is available as a free/BYOL Kubernetes application in Microsoft Marketplace and AWS Marketplace. Cloud infrastructure is billed by your provider; the same AI Server Commercial or Enterprise license activates the gateway and workers.

Microsoft Azure

Install the certified Kubernetes application into an existing AKS cluster. Start with a CPU node for deployment validation, then add NVIDIA GPU workers for real inference workloads.

Find it in Microsoft Marketplace

Amazon Web Services

Deploy the Helm delivery on Amazon EKS from AWS Marketplace. A small system node pool plus separate GPU worker nodes keeps cluster services stable and inference capacity easy to scale.

Find it in AWS Marketplace
PRIVATE BY ARCHITECTURE

Your prompts never leave your network — because there is nowhere for them to go

AI Server is not a proxy to a cloud API. The models run on your hardware; requests travel from your apps to your server and back, inside your perimeter. Offline and air-gapped operation are supported deployment modes, not special exceptions.

  • No vendor cloud in the data path. Prompts, documents and generated content stay between your clients and your server.
  • Content-free records. Usage analytics and audit logs record who called what, when — never the text of a prompt or a response.
  • Licensing that fails safe. If our licence service is ever unreachable, your licensed fleet keeps serving. Your uptime does not depend on ours.
  • Your keys, your rules. API keys are issued and revoked by you, on your server — not accounts in someone else's cloud.

Questions about AI Server or AI Gateway?

For deployment help, sizing advice, Kubernetes questions, or licensing, write to [email protected] — or use the contact form and pick "Product".

Get release updates

New free AI products, major updates, and a few releases available only via this site. No spam.