One endpoint. Three intelligences.

Private AI inference with web search, model routing, and your organisation's knowledge built in. UK data residency on AWS London (eu-west-2).

14-day trial — no credit card required. Start now

UK DSPT EU GDPR — Coming soon US Data Residency — Coming soon

Running on AWS EKS London (eu-west-2) — inference requests and stored organisational knowledge remain in this region

Starter
£249 / month
1-8B parameter models
  • Dedicated private inference endpoint
  • Up to 5 users at a time
  • Organisation knowledge layer (OKF)
  • Web search and page reader built in
  • Automatic model routing per query
  • OpenAI-compatible API
  • Scale to zero when idle
Start free trial
Most popular
Team
£499 / month
9-20B parameter models
  • Dedicated private inference endpoint
  • Up to 20 users at a time
  • Organisation knowledge layer (OKF)
  • Web search and page reader built in
  • Automatic model routing per query
  • OpenAI-compatible API
  • Scale to zero when idle
Start free trial
Advanced
£3,999 / month
21-70B parameter models
  • Dedicated private inference endpoint
  • Up to 50 users at a time
  • Organisation knowledge layer (OKF)
  • Web search and page reader built in
  • Automatic model routing per query
  • OpenAI-compatible API
  • Scale to zero when idle
  • Priority support
Start free trial

Enterprise Isolation

Dedicated hardware node. No other customer's workload runs on the same GPU. Required for NHS Trusts, financial services, and regulated organisations. All features included. Pricing after Architecture Review.

Thank you. We will be in touch within one business day.
Open email client
Works with your existing automation stack
n8n Zapier Make Open WebUI MCP tools Any OpenAI-compatible client

What is inside every endpoint

These three features run server-side in the inference proxy. Any client that calls your endpoint gets them automatically.

Smart endpoint

Web search and web reader

Your endpoint can search the web and read URLs server-side. Toggle on at deployment. Every caller - curl, Open WebUI, your own app - gets tool-augmented responses. No tool loop code to write.

Model routing

Best model per query

One URL. The endpoint classifies each incoming query (code, reasoning, general) then picks the best specialist model within your tier. Complex queries go to the larger model automatically. The caller sees a normal completion.

OKF knowledge

Organisation knowledge layer

Upload your policies, runbooks, and internal documents as a knowledge bundle. Queries that touch your organisation's data are answered from your content, grounded in facts, not guesswork. Runs entirely within your region.

Common questions

Any OpenAI-compatible client. Point your API base URL to your CaseDesk endpoint and set your API key. Works with curl, Python openai library, LangChain, Open WebUI, n8n, Zapier, Make, or any app you build yourself. Web search, model routing, and OKF knowledge are available to all clients automatically.
Standard (Starter, Team, Advanced) runs in a dedicated namespace on a shared node pool - your GPU may be co-located with other customer workloads at the hardware level. Enterprise Isolation provisions a dedicated physical GPU node: no other customer's workload runs on the same hardware. Required for NHS, financial services, and regulated sectors.
Your endpoint runs on a managed cluster in the region you select. UK endpoints run on AWS EKS in London (eu-west-2) — live now. EU (Azure westeurope) and US (GCP us-east1) are coming soon. Inference processing and stored OKF knowledge remain in the selected region. Optional web tools may contact external websites or search providers. DSPT documentation is provided for UK deployments on request.
Yes. Your deployment scales to zero after a configurable idle timeout (default 30 minutes). Scale-up from zero takes 60-120 seconds. The monthly fee covers the endpoint, tooling, and data residency guarantee regardless of usage. You are not billed per inference request.
Starter (1-8B): Qwen3 8B, Llama 3.2 3B, Phi-4 Mini. Team (9-20B): Phi-4 14B, Qwen 2.5 Coder 14B, DeepSeek-R1 14B. Advanced (21-70B): Llama 3.3 70B, DeepSeek-R1 70B, Devstral 22B. The deployment wizard recommends the best model for your use case. Model routing then selects the right model per query automatically.
Enterprise pricing is based on hardware reservation, model tier, region, and contractual commitment. It is quoted after a Discovery call and Architecture Review. Contact sales@getcasedesk.com to start the process.