Engineering teams kept hitting the same wall: AI tools were productive, but every prompt left the building. CaseDesk was built to fix that - dedicated inference that stays in your region, on your terms.
When a team uses a public AI API, every prompt - including proprietary code, customer data, and internal knowledge - is transmitted to a third-party server, usually in the United States. For most companies that is a compliance risk. For NHS organisations and regulated industries it is a non-starter.
Self-hosting solves the privacy problem but introduces a new one: you need GPU hardware, Kubernetes expertise, and a team willing to maintain it. Most engineering teams do not have that capacity, and they should not have to.
CaseDesk sits in the middle. We manage the infrastructure. You get a dedicated endpoint in the region you need, with the compliance posture your organisation requires, and the OpenAI-compatible API your tools already speak.
Data residency enforced at the infrastructure level is more trustworthy than a contractual commitment. Your prompts never route through CaseDesk servers.
We use open-source models and a standard API. If you want to move to your own cluster tomorrow, we will help you do it. The endpoint you use today works on self-hosted vLLM without code changes.
One endpoint, one API key, one region. We deliberately avoid building a platform that requires a dedicated admin. If it takes more than ten minutes to set up, we have failed.
GPU-hour billing with no hidden minimums. You pay for what you use. Scale to zero when idle. No surprise invoices at the end of the month.
To be clear about what you own: CaseDesk provisions and operates the GPU infrastructure. On standard tiers, the node pool is shared across customers. On Enterprise Isolation, the hardware is dedicated to you but still operated by us. You do not own the GPU and cannot take it with you - that is the managed service trade-off.
What you can take with you is everything that matters to your application. Every deployment uses open-source model weights - there is nothing proprietary in the model itself. The API is the standard OpenAI format. There is no custom SDK and no CaseDesk-specific code in your application.
If you decide to run your own infrastructure - whether on AWS, Azure, an on-premise server, or a dedicated GPU node - you provision your own hardware, install vLLM, load the same model, and point your application at the new base URL. The same API calls, the same system prompts, the same tool integrations work without modification. We will provide the exact vLLM configuration that matches your current deployment.
We think this is the only honest way to sell infrastructure to enterprises. You choose CaseDesk because the managed service is worth it right now, not because leaving would be painful.
We are a small, focused team. If you have a question that the docs do not answer, you will reach a human who built the product.