Onsites AI is 100% free forever — 3 seats and 100 MB included. Start free →
Learning Center

Bring your own AI model: OpenAI-compatible endpoints

By the Onsites AI team · Last updated · 5-minute read

OWN API KEY your account, your vendor + strongest models + zero GPU ops − per-token vendor bill INTERNAL GATEWAY your LLM service, your domain + data never leaves + one endpoint, many models − runs on your VMs FULLY OFFLINE weights inside the gap + nothing crosses out + no usage bill at all − you own the GPU Same desk either way: drafting, translation, summaries, handbook Q&A — configured in one settings screen.

On the managed cloud, Onsites AI arrives pre-wired: credits meter drafts and summaries, the model lives behind our endpoint, and you never think about it. Self-hosted reverses the deal — the desk ships without a brain and asks you to bring one: your API key to a provider, an endpoint on an internal LLM gateway, or weights on a GPU in your own rack. This is not a tax; it is the design that makes self-hosting honest about AI economics and data control. You pay for compute once instead of renting it per action, and not one token of customer text leaves your network unless your own key says so. Here is how to choose, connect and live with whichever shape fits.

The three ways to bring a model

Your own API key is the fastest: point the desk at your provider account (OpenAI, Anthropic, a regional hyperscaler — any OpenAI-compatible endpoint works), paste the key in settings, and the full Copilot experience lights up. You inherit the vendor's model quality and its bill, per-token; your data crosses to the vendor under your contract with them, not ours. An internal gateway is the enterprise shape: your organization already runs a central LLM service (or builds one on an inference stack), and the desk speaks to that endpoint like any other internal system. Data stays inside, the org's model policy applies uniformly, and switching models later is a gateway config change rather than a desk upgrade. Fully offline weights suit the air-gapped deployment: an open-weight model with a small inference server, entirely inside your perimeter, zero per-use cost forever, GPU care entirely yours.

Choosing between them: four questions

Where may data go? If the answer is "nowhere at all," the choice is already made — offline weights, full stop. If regulated-but-routable is acceptable under your DPA, the internal gateway keeps movement auditable and internal. If customer support text is ordinary commercial data, your own API key is reasonable and by far the cheapest to operate. Who operates hardware? No GPUs on staff argues for an API key; an ML team argues for a gateway you can share with other internal tools. What languages must the copilot handle? Translation quality tracks model size multilingually — test your real language pairs during the 60-day trial before committing. What does the desk actually need? The copilot's work is drafting, polishing, translating, summarizing and grounded Q&A over workspace context — demanding, but not frontier-reasoning; a mid-size current model comfortably clears the quality bar, which is why "biggest model" is rarely the right answer and "biggest model your ops can babysit" is.

Connecting it: the thirty-minute config

All three shapes converge on one screen. In the self-hosted admin, the AI settings take an endpoint URL, an API key (if any), a model identifier and sane limits; the deployment guide shows the exact fields and the env-var fallback for headless installs. Thirty minutes, in order: verify connectivity from the desk's container to the endpoint (the most common failure is a firewall rule, not a model); run the built-in test action — one draft on a dummy conversation; check the latency, because copilot UX is latency UX; then enable features for one pilot team before the whole workspace. Rollout advice that saves weekend calls: keep a second model identifier configured as fallback from day one, so a provider outage or a GPU reboot leaves the desk drafting on the small model instead of on human patience.

The economics, with real numbers

The honest comparison is per-action. On the cloud desk, a Copilot draft is 3 credits and a Harness-with-tools run is 24 — roughly $0.03 and $0.24 at published rates; a 10-person support team using AI liberally spends the upper tens of dollars per month. Self-hosted, the same actions cost you inference only: with an API key, mid-size models price drafts around fractions of a cent per action; with offline weights, the marginal action costs electricity and the monthly bill is the GPU's amortization — a few hundred euros on a single server that also serves every other internal AI need you point at it. The crossover you are hunting: if AI usage is light and sporadic, the API key's simplicity wins; when usage runs high and constant across many internal tools, owned compute wins and keeps winning. The AI-cost walkthrough for the cloud side gives you the spending baseline to compare against your key's first invoice.

Try before you license: the trial dress rehearsal

Because every self-hosted start begins as a 60-day trial of the full product with no payment up front, the model decision belongs inside the trial, not before it. A sequence that has worked for teams: week one, deploy via the Docker guide against a provider key you already own, import a slice of real history, and let two agents work live; week two, draft your evaluation set and score drafts and summaries on your actual language pairs; week three, if data policy pushes that way, trial the internal gateway or a small offline model on the same evaluation file and compare the diff; week four, write the decision memo — model choice, GPU or vendor commitment, fallback policy — while the license purchase is still reversible by simply not making it. The four-question checklist should be answered by the end of that memo too. Ten seats is the license floor, so the arithmetic to bring to the memo is a comparison you already know how to run: ten self-hosted seats versus the same workload metered on the cloud desk.

Living with it: governance and quality

Two habits keep a BYO model deployment trustworthy. Governance through config, not vibes: the desk consumes the model; your access policy (who may use AI actions, what data the endpoint may see) is enforced by the gateway and the desk's own audit trails, and it should be written down the way your backup policy is. Quality checks: run a small evaluation file of your real conversations through draft and summarize monthly — ten threads is enough — and keep a human review habit per hallucination guardrails: AI proposes, a person reads, the desk records who sent what. If quality drifts after a model change, you will see it in that diff, not in customer complaints first, and the fix might be a config line rather than a procurement cycle.

One reassurance to end on: BYO model is a requirement of self-hosting, not a penalty of it. Every AI feature the cloud desk has — drafting, tone, translation, summaries, Harness Q&A with visible sources, agent tools — works the same against your endpoint. You are choosing whose computer does the thinking, not whether the desk can think.

Frequently asked questions

Can I use my own AI model with self-hosted Onsites AI?
Yes — it is the design. Point the desk at your own provider API key, an internal LLM gateway, or fully offline open weights inside your network. All AI features (drafting, translation, summaries, Harness Q&A) run against the endpoint you configure.

Do we have to pay Onsites for AI on the self-hosted plan?
No. Self-hosted licenses carry no AI fees because you provide the model and its compute: per-token costs belong to your provider account, or owned-GPU electricity and amortization for offline weights.

Which model should we bring — should it be the biggest one?
Not necessarily. The copilot's work is drafting, polishing, translating, summarizing and grounded Q&A — a mid-size current model clears the quality bar. Choose by data policy and operating capacity: "the biggest model your ops can babysit" beats "the biggest model."

How hard is it to configure our own model?
One admin screen with endpoint URL, key, model identifier and limits — typically 30 minutes including a connectivity check and a test draft. The deployment guide documents exact fields; the most common failure is a firewall rule, not a model problem.

What happens when our model endpoint goes down?
Configure a fallback model identifier from day one so drafting degrades to the small model instead of failing outright. AI features degrade gracefully; the desk's core inbox, CRM and document tools never depend on the AI endpoint.

Create your free workspace →  See the pricing