# Running your own model or using an API: how to decide.

> In this field note, veridive sets out how to choose between a public API, a provider’s enterprise or regional offering, a managed cloud deployment and a self-hosted open-weight model. It explains when self-hosting pays off, what it takes to run, how quality and costs compare, and why the choice should be made per workflow and kept reversible.

Self-hosting an open-weight model makes sense when data must not leave your infrastructure, volumes are high and steady, or you need control over change. The price is operations. Decide per workflow with the evaluation set, not per company by principle.

## Key takeaways

- The options form a spectrum: public API, a provider’s enterprise or regional offering, managed cloud deployment and a self-hosted open-weight model.
- Self-host when data must not leave your infrastructure, volume is high and steady, or you must control when the model changes.
- Self-hosting means GPU capacity, serving, monitoring, patching, model upgrades and the people to run them, including on a bad night.
- Decide workflow by workflow on your own evaluation set, behind a model-agnostic layer, and take data residency questions to counsel.

In the security review, someone says no customer data may leave the company’s servers. In the business, someone wants the most capable model available. Within a meeting or two, the question has become a company-wide yes or no, and both sides are arguing about the wrong unit.

Where a model runs is a decision per workflow. Self-hosting an open-weight model (one whose weights you can download and run on your own hardware) makes sense when the data must not leave your infrastructure, when volumes are high and steady, or when you need control over when the model changes. The price is operations: hardware, serving, monitoring, patching, upgrades and the people who do all of it. Everything else can usually run on an API, and the evaluation set shows what that costs you in quality, if anything.

## What are the options between a public API and your own servers?

Four, and the middle two are the ones people forget.

| Option | Where the model runs | What you control | What you operate |
|---|---|---|---|
| Public API | The provider’s infrastructure, on standard terms | Your prompts and what you send | Your application |
| Enterprise or regional offering | The provider’s infrastructure, under enterprise terms, sometimes in a region you choose | Contract terms such as retention and processing region | Your application and the contract |
| Managed cloud deployment | A model hosted inside your own cloud account | Region, network, access and logs | Configuration, access and spend |
| Self-hosted open-weight model | Your data center, or servers you control | Everything, including when the model changes | Hardware, serving, monitoring, patching and upgrades |

A requirement that starts as “we must self-host” sometimes ends in the second or third row once the DPO and counsel have read the terms. What each row means for your data is their question, not the architecture diagram’s; [where your data goes when you use a language model](https://veridive.com/insights/llm-data-privacy/) lists what to ask.

## When does self-hosting make sense?

When at least one of these holds clearly, and the evaluation set shows an open-weight model is good enough for the task:

- **The data must not leave your infrastructure.** Regulation, a customer contract or internal policy rules out the other rows, after counsel has looked at them.
- **Volume is high and steady.** Hardware costs the same busy or idle. A steady, heavy load keeps it busy; a spiky or small one doesn’t.
- **You need control over change.** Providers update and retire models on their own schedule. A process where every change must be validated before use may need a model that changes only when you say so.
- **The site is isolated.** A plant with limited connectivity needs the model on site.

It rarely makes sense when volume is low or unpredictable, when your evaluation set shows that only the largest hosted models reach the bar, or when nobody in-house can run GPU infrastructure. Without one of the conditions above, the default is an API or a managed deployment.

## What does it really take to run?

More than a server with a GPU:

- **GPU capacity** sized for peak load, with room for failover and for testing the next model. Long documents need more memory, so context length is a sizing question.
- **Serving:** software that batches and queues requests, enforces limits and authentication, and exposes an internal API your applications call.
- **Monitoring** of latency, throughput, errors and GPU use, plus output quality measured against the evaluation set.
- **Security patching** of the operating system, drivers, serving software and containers, with model files taken only from sources you have verified.
- **Model upgrades.** New open-weight models appear often, and each needs an evaluation run, a capacity test and a rollout plan, or you stay on an older model by default.
- **People** who understand both the infrastructure and the models, on call when something breaks at month-end.

If that list has no owner, the answer is already no.

> Where a model runs is a decision per workflow, not a principle for the whole company.

## How do quality and language support compare?

Open-weight models vary widely, and more so outside English. A model that reads English contracts well may write stiff Turkish, mix formal and informal address, or miss domain terms. General benchmarks won’t tell you; your evaluation set will. Test each candidate on your own Turkish and English examples, scored by native speakers who know the work, and compare quality, cost per case and speed on the same cases.

Two things are worth checking. Smaller open-weight models are often good enough for narrow tasks such as extraction, classification and routing, where the output is structured and easy to verify. And the hardware you can afford limits context length and speed in practice: a model that has to truncate a long document may lose the part that holds the answer.

## How do costs behave in each option?

The shapes differ more than the prices:

- **Public APIs** charge per use. Cost starts near zero and grows with volume and with the length of what you send.
- **Enterprise and regional offerings** also charge per use, sometimes with commitments.
- **Managed deployments** are often billed for reserved capacity over time, so cost rises in steps, busy or not.
- **Self-hosting** is mostly fixed: hardware bought or leased, power, space and people. Each extra request costs little until capacity runs out, then cost jumps.

The comparison that matters is the fixed monthly cost of self-hosting, people included, against the API cost at your real volume, peaks included. The crossing point moves whenever prices or volumes change, so recalculate it rather than deciding once.

## How do you keep the choice reversible?

By never letting a workflow call a model directly:

- **A model-agnostic layer.** Applications call one internal interface; behind it, each workflow is routed to the model chosen for it, wherever that model runs.
- **Portable assets.** Prompts, output schemas and evaluation sets stay provider-neutral and are owned by you.
- **The evaluation set as the switch test.** Before a workflow moves, its set is re-run on the new option, and the results decide.
- **Exit terms.** Contracts say how data is exported and deleted when you leave.

Consider an illustrative insurer with three workflows. Claims document extraction reads medical reports and identity documents at high, steady volume; after review with the DPO and counsel, those documents stay in the company’s own data center, and an open-weight model meets the acceptance threshold on the evaluation set, so it runs self-hosted. Drafting replies to broker emails is lower in volume, carries less sensitive data and benefits from a larger model’s writing, so it uses a cloud API under enterprise terms, in a region counsel approved. An internal policy assistant runs on a managed deployment in the company’s cloud account. All three call the same internal interface, so any of them can move when its evaluation set says another option is better.

Where processing may happen, and whether a provider’s terms satisfy your KVKK or GDPR obligations, are questions for your DPO and counsel. The [questions to ask them](https://veridive.com/insights/kvkk-gdpr-ai-questions/) have their own note.

## Decide workflow by workflow

List your candidate workflows and ask three questions of each: must the data stay inside, is the volume high and steady, must you control when the model changes? Where all three answers are no, use an API or a managed deployment and spend the effort elsewhere.

Our [data and AI foundations](https://veridive.com/services/data-ai-foundations/) work compares candidate models on your examples and deploys where your data needs to live, as described in [how we work](https://veridive.com/approach/). To talk through a specific workflow, [send a short brief](https://veridive.com/contact/).

This note is general information, not legal advice.

## Frequently asked questions

### Is it better to self-host an LLM or use an API?

It depends on the workflow, not the company. Self-hosting an open-weight model makes sense when data must not leave your infrastructure, volume is high and steady, or you need to control when the model changes. Otherwise an API or a managed cloud deployment is usually simpler and cheaper to run. Test both on your own evaluation set before deciding.

### What do you need to run an LLM on your own servers?

You need GPU capacity sized for peak load, serving software that queues and batches requests, monitoring of latency, errors and output quality, security patching of the whole stack, a plan for model upgrades, and people who can run all of it around the clock. You also need an evaluation set to confirm the open-weight model meets your quality bar.

### Are open-weight models good enough in Turkish?

It varies widely by model, task and domain, so test rather than assume. Build an evaluation set from your own Turkish documents and questions, have native speakers who know the work score the answers, and compare open-weight candidates with hosted models on quality, cost per case and speed. Narrow tasks such as extraction or classification are the easiest place to start.
