Open Weights, Closed Weights: What Developers Actually Choose
The debate over releasing model weights openly is usually conducted in terms of principle. Practitioners choose on the basis of constraints, and their decisions are more predictable than the rhetoric suggests.
The debate over whether model weights should be released openly is usually conducted in terms of principle. Practitioners choose on the basis of constraints — where the data can sit, what a token costs, whether anybody on the team has run a serving stack — and their decisions are more predictable than the rhetoric suggests.
Five reasons a team runs its own weights
- Data cannot leave the organisation or the jurisdiction.
- The workload is high volume and cost-sensitive, so per-token pricing is unaffordable.
- The team needs to fine-tune for a narrow task.
- Vendor continuity risk matters more than peak capability.
- Latency requirements rule out a network round trip.
Four reasons a team rents one
- The task needs the highest available capability, which open models trail.
- The team does not want to operate inference infrastructure.
- Usage is spiky and paying per request is cheaper than provisioned capacity.
- Tooling, retrieval and evaluation are included in the platform.
The clause that restricts commercial use above a revenue threshold
The most practical difference between open-weight releases is not capability but licence terms. Some are permissive; others restrict commercial use above a revenue threshold, prohibit certain applications, or require attribution. Teams that check benchmark scores and skip the licence have occasionally discovered the problem after shipping.
The sensible default is to treat an open-weight model like any other third-party dependency: read the terms, record the version, and plan for the possibility that a future release changes them.
Owning the serving stack is a second job
Running a model yourself means owning the problem of serving it: procuring hardware or renting it, managing a serving stack, handling load spikes, patching, monitoring and capacity planning. For an organisation whose core competence is something else, that is a substantial distraction.
The teams that choose open weights successfully are usually those that already operate infrastructure. For everyone else, the relevant comparison is not open versus closed but self-hosted versus rented, and renting usually wins until volume is very high.
A frontier model and a small local one, routed side by side
The most common production arrangement combines both: a hosted frontier model for requests that need maximum capability, and a smaller open model running locally for high-volume, low-complexity tasks or for data that cannot leave the premises.
Routing between them requires a shared evaluation harness, which is the piece most organisations underinvest in. Without it, the routing decision is made on intuition, and cost savings are hard to demonstrate.
The evaluation burden also falls unevenly. A large organisation can afford a private test suite, a routing layer and an internal benchmark team. A smaller one picks whichever model its engineers liked most and revisits the decision when something breaks. That gap in evaluation capability is becoming as significant a source of disadvantage as any gap in model access.
Image credit and licence details for every photograph on this site are listed on the credits page. This article is editorial content; it carries no sponsored material.