Private AI Platform for Software Teams

Your models. Your cloud. Your knowledge. One private AI platform.

Production AI inference on infrastructure you control — with your organisation's governed knowledge built in. OpenAI-compatible. Region-locked. Ready in minutes.

Find the right deployment for your team

Three questions. No email required.

What does your team need AI for?

Step 1 of 3

How many people will use it?

Step 2 of 3

Where should your data be hosted?

Step 3 of 3
Your recommendation
Get started
Process

From question to endpoint in minutes

No Kubernetes expertise required. No cloud accounts to configure. CaseDesk handles infrastructure so your team can focus on the product.

1

Answer five business questions

Use case, team size, region, and compliance needs. No cloud account required at this stage.

2

Review your recommended deployment

CaseDesk selects the model class, GPU tier, and region. You see what you will pay before committing.

3

Your endpoint is live

A dedicated inference endpoint, OpenAI-compatible. Drop the URL into Cursor, Continue.dev, or your own app.

CaseDesk Dedicated

Your model. Your endpoint. Your region.

Private AI deployment with your organisation's governed knowledge built in — running on infrastructure you control.

You retain full control while CaseDesk manages deployment, upgrades, monitoring, and operations.

🔒

Data stays in your region

Your prompts and completions never leave the UK, EU, or US region you choose. CaseDesk does not log or store inference traffic.

🎯

One endpoint per team

No shared inference queues. Your deployment runs in its own namespace on a dedicated GPU cluster. Standard tier uses a shared node pool. Enterprise tier adds hardware isolation.

💸

Scale to zero when idle

Deployments scale to zero after a configurable idle period. Your endpoint stays available on a flat monthly subscription - no per-request charges, no usage spikes.

🔧

OpenAI-compatible API

Every endpoint implements the OpenAI REST API. Swap the base URL in your existing tools - no code changes required.

Integrations

Works with the tools you already use

Any OpenAI-compatible client connects without modification.

Open WebUI
Chat interface with file upload, RAG, and tool support
Cursor
AI-native editor, drop in as a custom model
Continue.dev
VS Code and JetBrains AI coding assistant
LangChain
Chain-based AI applications and agents
OpenAI SDK
Python and TypeScript SDKs, zero changes needed
n8n
Workflow automation with AI nodes
Flowise
Low-code LLM app builder with visual flows
Ollama API
Compatible with Ollama-format API clients
Pricing

Simple monthly pricing

One flat monthly fee. Web search, model routing, and your organisation's knowledge included. UK, EU, and US regions available.

Prices shown in GBP. See all regions and currencies

Starter
£249
per month
1-8B parameter models
14-day free trial included
  • Dedicated private endpoint
  • Up to 5 concurrent users (vLLM)
  • Web search and web reader
  • Organisation knowledge layer
  • UK, EU, or US data residency
  • OpenAI-compatible API
Start free trial
Team
£499
per month
9-20B parameter models
14-day free trial included
  • Dedicated private endpoint
  • Up to 20 concurrent users (vLLM)
  • Web search and web reader
  • Model routing included
  • Organisation knowledge layer
  • UK, EU, or US data residency
  • OpenAI-compatible API
Start free trial
Advanced
£3,999
per month
21-70B parameter models
14-day free trial included
  • Dedicated private endpoint
  • Up to 50 concurrent users (vLLM)
  • Web search and web reader
  • Model routing included
  • Organisation knowledge layer
  • UK, EU, or US data residency
  • OpenAI-compatible API
  • Priority support
Start free trial

Enterprise Isolation

Dedicated hardware node, no other customer workloads on the same GPU. Required for NHS Trusts, financial services, and regulated organisations. Quoted after Architecture Review.

Contact Enterprise
14-day Free Trial

Evaluate before you commit.

Every new account gets a sandbox deployment automatically. Shared GPU, UK region, scale to zero after 1 hour. Use it to try a model before moving to a dedicated deployment.

Create free account
Enterprise

Built for regulated industries

NHS Trusts, financial services, and government.

CaseDesk Enterprise provides dedicated node pools, hardware isolation, and infrastructure designed to support NHS DSPT and GDPR compliance. Your data and inference traffic never leave the UK region.

The sales process includes Discovery, Technical Workshop, Architecture Review, and Security Review before any data touches production infrastructure.

  • Dedicated GPU node - no shared hardware
  • UK data residency, DSPT-ready
  • SSO and RBAC
  • Private networking options
  • SLA and support contract
  • Architecture Review included
Enterprise team reviewing AI deployment options

We start with a 30-minute discovery call. No slide decks - just the technical questions that matter for your compliance requirements.

Contact Enterprise Sales
FAQ

Common questions

Your deployment runs in its own namespace on a regional GPU cluster. You get a dedicated inference endpoint - no other team's traffic shares your model. Standard tier uses a shared node pool. Enterprise tier adds hardware isolation: your model runs on a physical GPU that no other customer touches.
Your inference requests and completions never leave your chosen region, and CaseDesk does not use them to train models. CaseDesk runs open-source models on infrastructure it controls in UK or EU data centres, with OpenAI-compatible APIs. Optional web search and page-reading tools may contact external providers when enabled.
Answer five questions in the Deploy Wizard and your endpoint is live in minutes. No Kubernetes knowledge required. CaseDesk handles provisioning, scaling, and maintenance.
Your inference requests and outputs stay within the region you choose: UK (AWS eu-west-2), EU (Azure westeurope), or US (GCP us-east1). CaseDesk does not log or store your prompts or completions.
Yes. Every CaseDesk endpoint exposes the OpenAI REST API. Swap the base URL and your tools - Cursor, Continue.dev, LangChain, n8n, or any OpenAI SDK - work without code changes.
Yes. CaseDesk Enterprise tier provides dedicated hardware isolation, UK data residency, and infrastructure designed to support NHS DSPT compliance. Patient data and clinical information never leave the UK region.
Deployments scale to zero after a configurable idle period. Your monthly subscription continues regardless - you are paying for the private endpoint, data residency, and tooling, not for individual requests. Scale-up from zero takes 60-120 seconds.
Yes. You can redeploy to a different model at any time from the deployment detail page. Your endpoint URL stays the same.

Give your team a trusted AI platform.

Start with the sandbox. Move to a dedicated deployment when you are ready.

Find My AI Platform
No credit card required. DSPT-ready infrastructure.