2026-06-12
Private LLM deployment: open vs. closed, and the math on VRAM, concurrency, and compliance
Why does private LLM selection keep stalling teams? This article breaks down three core yardsticks from an engineering standpoint: VRAM is not just about parameters, because the KV Cache determines the real concurrency ceiling; concurrency should be worked backward from latency targets and QPS to the number of GPUs; and compliance must turn "meets MLPS" (China's Multi-Level Protection Scheme) into verifiable hard constraints. Using this three-yardstick framework, it compares open-source and closed-source options on the merits, and includes a utilization break-even calculation for cloud versus owned hardware plus a scale-to-hardware reference table.
Why selection stalls for half a month: translating "open vs. closed" into three engineering yardsticks
The sticking point in private deployment decisions is rarely "too few model options." Quite the opposite: more and more open-weight models are runnable, and closed-source vendors can all offer private-deployment licensing. Yet the more options on the table, the harder it becomes to make a call. I once watched the technical lead at a manufacturing company evaluate more than a dozen options over the better part of a month and still go in circles—not because any model was clearly inadequate, but because there was no single ruler that could measure all the options on the same scale. The real reason decisions stall is that the dimensions aren't quantified, not that there isn't enough information.
Most selection guides offer a four-dimensional framework: data security level, concurrency and performance requirements, budget and cost, and team capabilities. The four dimensions are not wrong, and none can be omitted; the problem is that they are too coarse to land on a purchase order. "Data must not leave the domain" is a stance, not a constraint—it doesn't tell you how many GPUs to buy, whether the model can be quantized, or whether logs need to be retained for audit. "Performance must be good enough" is equally empty: what QPS counts as good enough, how many hundred milliseconds the first token must be kept under, how much p99 jitter can be tolerated—all of it is hidden under the words "good enough." When the dimensions stop at the level of adjectives, review meetings can only argue over leanings again and again, never reaching an actionable conclusion.
Translate this coarse framework one level down and it actually converges into three calculable engineering yardsticks. The first is VRAM: model parameters set your floor, but what really consumes VRAM is often the KV Cache during inference, which grows linearly with concurrency and context length—the parameter count is only the smaller part of the bill. The second is concurrency: first set the latency target and peak QPS, then work backward to how many streams a single GPU can handle and how many GPUs you need in total, rather than buying hardware first and then seeing how fast it runs. The third is compliance: break a statement like "meets MLPS" (China's Multi-Level Protection Scheme) into hard constraints you can check off one by one—whether data leaves the domain, whether model weights can reside in your own data center, how long the call chain must be retained for audit. What the three yardsticks have in common is that they all land on numbers and on a purchase order, so reviews no longer compare whose stance is firmer but whose numbers are more solid.
Once broken down this way, the old question of "open source or closed source" shifts position too. It should not be a premise of selection but a result that falls out naturally after the three yardsticks are calculated. For example: if the compliance yardstick requires the weights to stay local, with not even the model files hosted on the vendor side, then the closed-source API route is essentially ruled out, and what remains is choosing a size class among open weights. Conversely, if data leaving the domain is permitted in your industry, the budget is tight, and the team can't support GPU operations, then a closed-source private deployment or dedicated instance may actually be less hassle. Picking a side first and then finding reasons easily turns an engineering problem into a battle of preferences; calculating the yardsticks first and then looking at the results returns open vs. closed to what it really is—a choice of delivery form, not a matter of faith.
So this article won't hand you an "open source vs. closed source" comparison table and ask you to pick a side. Instead, it first puts the three yardsticks (VRAM, concurrency, and compliance) in your hands. The following sections take them one at a time: how to calculate actual VRAM usage from parameters and the KV Cache, how to work backward from latency targets to the number of GPUs, and how to break compliance down into a checklist you can verify; they then combine all three into cost and payback, and finally produce a scale-to-hardware reference table you can take straight to procurement. Once all three yardsticks are worked out, you will most likely already have your own answer to the open vs. closed question, with no need to agonize over it for half a month.
The VRAM yardstick: parameters are only the floor; the KV Cache is the real bill for concurrency
Many teams doing private deployment selection for the first time match the model's parameter count against GPU specs, only to hit OOM as soon as concurrency ramps up after launch. The root cause is conflating "can fit the model" with "can handle concurrency." VRAM demand has two layers: loading the parameters is the floor, and the KV Cache is the part that actually grows elastically with the business.
Layer 1: the static minimum for loading weights
At FP16 precision, each parameter takes 2 bytes. Adding runtime framework overhead, rule-of-thumb estimates are as follows:
| Parameters | FP16 weight estimate | With runtime overhead |
|---|---|---|
| 7B | ~14 GB | ~20 GB |
| 14B | ~28 GB | ~40 GB |
| 70B | ~140 GB | ~170 GB, multi-GPU required |
What these numbers represent is the minimum threshold for a single request under no load. An 80 GB A100 holds a 7B model with room to spare and is enough for 14B as well, but this is only the starting point, not the end point.
Layer 2: the KV Cache bill amplified by concurrency
During inference, every active request needs to maintain its own KV Cache (key-value cache, which stores the intermediate states of the attention mechanism). The complete formula for actual VRAM usage is:
Total VRAM = weight VRAM + number of concurrent requests × KV Cache size per request
The KV Cache size per request depends on the number of model layers, the number of attention heads, and the sequence length. For a mid-sized model, the per-request KV Cache takes up a considerable share of VRAM as the sequence length grows. This means that after the weights are loaded, the remaining VRAM can support only a limited number of concurrent streams—and that is the ceiling; actual scheduling must also reserve headroom for the system.
Now look at the common pitfall in reverse: if the purchase was based on "7B only needs 20 GB" and you bought a 24 GB consumer card, then after the weights are loaded only 4 GB remain for the KV Cache. Two or three concurrent requests fill up the VRAM and trigger OOM outright. The parameter spec wasn't wrong, but by ignoring the concurrency layer, the purchase was.
Quantization: the lever that trades precision for VRAM
When the target hardware's VRAM can't hold FP16 weights, quantization is the most common engineering technique. INT8 quantization compresses weights from 16 bits to 8 bits, reducing VRAM usage by about 75%, while most evaluations show accuracy loss within 2%—an acceptable price for most internal enterprise Q&A and document processing scenarios.
In practice: after INT8 quantization, the VRAM usage of a large language model (LLM) drops noticeably, going from "must span multiple machines" to "runs on a single machine." This is not a free lunch with lossless precision, but a clear engineering trade-off: exchanging a quantifiable loss of precision for a dramatic reduction in hardware cost.
INT4 quantization can push VRAM even lower, but the accuracy drop is usually more pronounced. Whether it is usable depends on how sensitive the specific task is to accuracy, and it must be tested on your own datasets rather than judged by generic benchmarks alone.
Capacity baseline: measured reference for 7B on an A100 80G
A 7B-parameter model fully loaded with FP16 weights on a single A100 80GB can keep inference latency stable below 200ms—a capacity baseline worth remembering. Its significance is that you can use it as an anchor to extrapolate upward to how many GPUs 14B and 70B need, and downward to the latency cost of running a quantized 7B on smaller cards (such as the A10 24G).
Run these numbers before selection
Before getting to the "how many GPUs to buy" decision, it's advisable to pin down the following three numbers first:
- Peak concurrency: the highest number of simultaneous in-flight requests the business side can specify, which determines the KV Cache budget
- Average sequence length: the total token length of input plus output, which directly affects the per-request KV Cache size
- Precision tolerance: whether the business scenario can accept INT8 or requires FP16, which determines whether quantization is an option
Once these three numbers are set, plug them into the formula above to get the minimum VRAM requirement, then add a 20%–30% safety margin—that is the true procurement floor. Skipping this step and buying GPUs based on parameter count alone is one of the main reasons selection drags on for half a month or ends in rework.
The concurrency yardstick: working backward from latency targets and QPS to the number of GPUs
Many teams get stuck on selection not because they don't understand hardware, but because they skip a prerequisite step: write the latency SLA down as a number first. A description like "responses need to be fast" imposes no constraint on a purchase order; "time to first token no more than 1 second" is the engineering entry point from which you can work back to a GPU count.
Step 1: separate the scenarios—the SLA gap determines the configuration gap
The two types of scenarios require resources of entirely different magnitudes, and estimating them together leads either to severe overprovisioning or to a system that collapses on launch:
- Real-time conversation (consumer-facing or embedded in business processes): time to first token must be kept within 1 second, or users will clearly notice lag. This constraint directly limits the batch size—larger batches mean higher throughput but also longer queuing times, and there is a hard tension between the two, so you can't just focus on the throughput figure.
- Internal tools / low-concurrency background tasks: knowledge base Q&A, code review, and report generation with a few dozen people online at once, where a latency of 3–5 seconds is perfectly acceptable. The bottleneck in these scenarios is VRAM capacity rather than latency, so you can trade larger batches for higher throughput; on the same hardware, the number of people actually served is several times that of real-time conversation scenarios.
Applying the same configuration standard to both types means either buying overly expensive low-latency hardware for low-concurrency internal tools, or loading real-time conversation onto a throughput-oriented machine and leaving users waiting for the first word.
Step 2: the inference framework determines how much throughput you can squeeze out
Hardware sets the ceiling; the inference framework determines what fraction of that ceiling you can actually reach. The three mainstream frameworks are clearly positioned, and mixing them up or choosing the wrong one can leave the same hardware delivering half the results or worse:
- Ollama: launches with a single command, suitable for local development and proofs of concept. It has no production-grade concurrency management, so don't use it for performance benchmarks under real load.
- vLLM: currently the most active production inference framework in the community. Its core is the PagedAttention mechanism—paging the KV Cache on demand rather than preallocating contiguous blocks—which greatly reduces VRAM fragmentation and significantly increases the number of concurrent requests the same VRAM can support. The first choice for high-concurrency production.
- TensorRT-LLM: NVIDIA's official offering, with deep kernel optimizations for its own GPUs. Its latency and throughput ceilings are higher than vLLM's, but deployment complexity and maintenance costs are also higher. Worth the investment for scenarios with extremely demanding latency SLAs (such as real-time voice or financial trading assistance) and an operations team with CUDA experience.
Step 3: estimate VRAM with QPS × per-request KV, then validate batch size against latency
Working back to the GPU count follows two lines of calculation; take the larger result:
The VRAM constraint line: multiply the target concurrent QPS by the per-request KV Cache footprint (determined by context length and the number of model layers; see Section 2 of this series), add the VRAM base of the model weights themselves, and you get the peak VRAM requirement. This line tells you "the minimum VRAM you need."
The latency constraint line: the time-to-first-token SLA in turn limits the maximum batch size allowed in the prefill stage. The larger the batch, the longer prefill takes and the later the first token arrives. If your SLA is 1 second, you need to measure the largest batch size that meets the latency target at the target QPS, then confirm whether single-GPU or multi-GPU compute is sufficient at that batch size.
Only when both lines pass do you have enough GPUs. Look only at VRAM and ignore latency, and the first token will time out after launch; look only at latency benchmarks and ignore peak VRAM, and an OOM at peak will bring the service down.
Engineering reference for model size and GPU count
| Model size | Typical scenario | Recommended hardware | GPU count |
|---|---|---|---|
| 7B | Internal testing, small-team tools | RTX 4090 / A10 | Single GPU |
| 14B | Production, low-to-medium concurrency | A100 40G / H20 | 1–2 GPUs |
| 70B | Production, high concurrency | A100 80G / H800 | 4–8 GPUs |
The figures above come from commonly cited reference ranges in industry deployment practice; specific values should still be based on test results for your actual context length, peak concurrency, and batch size. When 70B spans 8 GPUs, inter-GPU communication bandwidth (NVLink vs. PCIe) becomes a new bottleneck, and the interconnect topology must be confirmed during evaluation—it's not just about assembling enough total VRAM.
The conclusion: selection on the concurrency yardstick is not a table lookup. It starts from your SLA numbers and passes through three filters—framework choice, KV estimation, and latency validation—to arrive at a minimum GPU count you can write into a purchase order. Skip any one of them and the number will be distorted.
The compliance yardstick: breaking "meets MLPS" into verifiable hard constraints
The previous two sections dealt with VRAM and concurrency—the question of "can it run, and can it run fast enough." But in heavily regulated industries such as finance, healthcare, and government, there is another gate that takes effect before any of that math: if a solution fails on compliance, no matter how good its performance, it is out. The trouble with this gate is that business stakeholders often phrase it as a vague "must meet MLPS," and engineering simply can't run acceptance testing against a slogan. To get it through review and audit, you first have to translate that phrase into a few hard constraints that can be checked off item by item.
First, why it deserves separate treatment. VRAM and concurrency are continuous quantities; if you fall a little short, you can compensate by adding GPUs, reducing concurrency, or compressing context. Compliance is not continuous—it is a Boolean. A solution either satisfies "data doesn't leave the internal network" or it doesn't; there is no "mostly satisfies" state in between. This means compliance constraints should sit at the very front of the selection process as a filter, rather than waiting until the architecture is finalized and the purchase order drawn up, only to be vetoed by security or legal—at that point, rework costs are measured in weeks, and a good part of that half month of selection agonizing gets stuck right here.
Three items you can check one by one
Unpack "meets MLPS," and three items are decisive for deployment architecture, each of which maps to a concrete verification action:
- Data localization: where training corpora, inference inputs and outputs, vector databases, and logs are physically stored. The verifiable question is "list every storage node that holds business data, along with its physical location and owning entity." If even one copy of the data sits in a data center or third-party account you don't control, this item fails.
- Staying within the internal network: whether any hop in the inference chain crosses the public internet. What's most easily overlooked here is not the main model call but the external capabilities wired in "along the way"—web search, third-party embedding services, calls to external APIs for tool augmentation. The verifiable question is "draw the complete call chain and mark the network boundary of each hop"; every arrow that goes out to the internet must be clearly explained or cut off.
- MLPS certification: the system completes grading, filing, and assessment under the Multi-Level Protection requirements and obtains a conclusion at the corresponding level. Core systems in finance and government usually require Level 3. This item checks process evidence rather than technical parameters, but it in turn constrains what kind of underlying infrastructure you can use.
For finance, healthcare, and government, private deployment selection must satisfy MLPS certification, data localization, and staying within the internal network—in 2024 industry practice, this is already a default premise rather than a bonus. Write it into the requirements baseline as a hard constraint, and later architecture discussions will have a shared boundary.
How compliance locks in the architecture in reverse
What really affects selection is the chain reaction of these three constraints. It becomes clearer if you look at it from a different angle: don't ask "what deployment form do I want to use," ask "given that data can't leave the internal network, which forms are still alive."
Working backward from that angle, public-cloud on-demand inference services are the first to go—their essence is sending requests to the vendor's data center for processing, so "not leaving the internal network" and "calling the public cloud on demand" are mutually exclusive by definition. The same goes for closed-source LLMs in SaaS form: whatever they're called, as long as inference happens outside your internal network, they fail the first check. After this filter, essentially only two categories remain on the table: installing a closed-source model in your own data center under a private-deployment license, or simply deploying an open-source model yourself.
At this point, the much-debated question of "open or closed source" has already had half its option space cut away by compliance constraints. What's left is not a battle of sides but continued math, using the VRAM and concurrency yardsticks from the earlier sections, between "closed-source private-deployment licensing" and "self-deployed open source." Compliance hasn't made the choice for you, but it has crossed off a large swath of options you would otherwise have agonized over—that is its most practical value.
Two more details that tend to trip up deployment are worth watching in advance. The first is the temptation of hybrid architectures: handling everyday requests with a small local model and passing hard problems "as a fallback" to a large public-cloud model. This design looks good on both performance and cost, but as soon as that fallback path carries real business data out of the network, it drags down the compliance of the entire solution in one stroke. The second is the update channel: how are the model weights and dependency images of a private deployment updated? If they're pulled directly from the public internet, that is another hidden outbound path that must also be explained during audits. Think these two points through, and your hardware purchase order won't have to be torn up after a single security review.
To sum up in one sentence: the compliance yardstick doesn't take part in the discussion of "is the performance good"; it only answers "does this solution even get a seat at the table." First run the candidates through the three items—data localization, staying within the internal network, and MLPS certification—then do the detailed math on VRAM and concurrency. Get the order backward, and the more detailed your earlier calculations, the more painful the rework later.
Open vs. closed source: calculable judgments from the three yardsticks, not taking sides
The part of selection discussions most likely to go off track is debating open vs. closed source as a matter of stance. In engineering terms, the question has no camps—only which path has the lower TCO and more controllable risk given your VRAM budget, concurrency targets, and compliance boundaries.
First, lay out the cost structure of both paths
Open-source models (Llama 3, DeepSeek, Qwen, GLM, and others) carry no license fee—that's a fact. But "free" only means zero royalties, not zero deployment cost. The real bill has three parts:
- Hardware capital expenditure: model parameters set the VRAM floor, and the KV Cache and concurrency targets determine the actual purchase volume—this money is entirely on your side, with no one to share it.
- Operations staffing: inference framework selection (vLLM / TGI / SGLang), quantization precision trade-offs, VRAM fragmentation management, version upgrades and regression testing—each needs someone watching over the long term. Based on common industry observation, a private deployment project built from scratch to stable operation, including tuning and on-call duty, usually requires several engineers with GPU cluster experience on an ongoing basis, with especially heavy involvement early on.
- Opportunity cost: a team's energy has limits. Every person-day spent on infrastructure is a person-day not spent on business features.
The bill structure of closed-source/commercial solutions is the exact opposite: capital expenditure is low (or converted into subscriptions), and operational complexity shifts to the vendor, but room for customization is bound by contracts and interfaces, and the monthly bill is an ongoing, rigid expense—the higher the business volume, the higher the cost, with no concept of a "buyout."
The switching line under the three yardsticks
Rather than a conclusion, here is the decision logic. Plug in the three yardsticks from the previous sections directly:
| Yardstick | Signals pointing to open-source private deployment | Signals pointing to closed-source/managed |
|---|---|---|
| Compliance | Data can't leave the internal network, industry regulation requires the model to be auditable, proof of on-premises deployment is required | Compliance requirements are loose, or can be met by calling external APIs after data masking |
| VRAM/concurrency | Concurrency is high and stable, utilization of owned GPUs can stay above 60% over the long term, and you can optimize quantization and scheduling | Concurrency fluctuates widely, the peak-to-trough ratio exceeds 5×, and you can't afford the capital waste of idle VRAM |
| Team operations capability | ≥1 engineer familiar with inference frameworks and GPU cluster operations who is willing to own them long term | AI infrastructure isn't the core business, the team is mainly business developers, and you can't create a dedicated infrastructure role |
The switching line isn't either-or; it's weighted. If all three yardsticks point the same way, the judgment is clear-cut. If two point to open source and one to closed source, that one "pointing to closed source" is usually the bottleneck and needs to be solved separately (for example, by hiring an operations engineer, or by absorbing fluctuating concurrency with elastic cloud GPUs).
The most common miscalculation behind "open source is free"
A typical faulty decision path goes like this: seeing that Qwen or DeepSeek can be downloaded for free, the team enters 0 for model licensing fees at project approval, then underestimates the hardware procurement and staffing budgets, and costs far exceed expectations after launch.
A more accurate TCO calculation should be:
- VRAM procurement = GPU count calculated with the Section 2 yardstick × cost per GPU + spare redundancy
- Operations staffing = engineer annual salary × share of time allocated to the project, starting at no less than 1 person-year
- Opportunity cost = the expected output of the same staffing invested in product features (hard to quantify, but can't be ignored)
- Comparison baseline = annual fees for a closed-source API or managed solution at the same call volume
When VRAM utilization on the owned-hardware route stays below 50% for an extended period, the TCO break-even point usually tips toward a managed solution. This point isn't a fixed number—it depends on your concurrency distribution and hardware depreciation cycle—but it can be calculated with the utilization model in Section 6.
Where open source's real advantages lie
Strip away the halo of "free," and open source offers two core values:
The first is data sovereignty. Weights stay local, inference happens on the internal network, and logs never leave the data center—for finance, healthcare, and government scenarios, this isn't a bonus; it's the entry threshold. Under this constraint, closed-source APIs can't even compete, so there is no "which is better" comparison.
The second is room for deep customization. Fine-tuning on training data, modifying inference logic, choosing how to connect to a private knowledge base, adjusting output formats—the freedom open-source models offer on these dimensions is something no amount of prompt engineering on a commercial API can replace. If your business scenario requires highly controllable model behavior, the value of this freedom grows over time.
The real advantage of closed source is equally clear: outsourcing infrastructure risk. Inference framework upgrades, model version iterations, hardware failure recovery—these aren't unimportant; you just don't have to do them yourself. For business teams where AI is a tool rather than a core competency, this outsourcing logic holds up completely.
An actionable decision process
Before getting into specific selection, it's advisable to answer three questions in the following order:
- Can data and models leave the internal network? If yes, closed-source APIs remain candidates; if no, lock in private deployment directly, then discuss open source versus a commercial private-deployment edition.
- Does the team currently have someone who can own inference framework operations? If yes, the open-source route is viable; if not, and you can't hire in the short term, prioritize commercial solutions with managed support.
- Can concurrency utilization justify the payback on buying your own hardware? Calculate with the break-even model in Section 6: if you clear the line, buy your own; if not, pay on demand in the cloud.
Once all three questions have clear answers, the choice between open and closed source has essentially been decided by the constraints, and there's no need for judgments at the level of "technology philosophy."
Cost and payback: cloud or owned hardware? Calculate the utilization break-even point
By this point, the three yardsticks have already pinned down "how much compute you need." Only one question remains: rent these GPUs or buy them. Many teams pick sides on instinct here—feeling that "owning is safer" or "the cloud is more flexible"—but this is a pure arithmetic problem, and the variable is utilization. Work out the utilization and the conclusion will surface on its own.
Start with the most common scenario, a single GPU running 7B, since it's the starting point for most private deployment projects. By rough current market estimates, renting a cloud GPU instance for a full year costs on the order of RMB 30,000–50,000; buying the GPU and machine outright requires an upfront investment of roughly RMB 100,000–150,000. Note that these two numbers aren't on the same dimension: the former is an operating expense you renew every year, while the latter is an asset you pay for once and amortize year by year. Supporting costs such as bandwidth, storage, and data center space must be counted separately on both sides, and they're often underestimated on the owned-hardware side: once a server goes into the data center, the hidden bill for electricity, operations staff, and spare parts keeps running.
What really determines the conclusion is the payback period, and the payback period is derived from utilization. Here's a more intuitive way to calculate it: a local server with an A100, including three years of maintenance costs, comes to a total investment of roughly $80,000. The same compute on a public cloud with on-demand billing costs about $5 per hour. Divide $80,000 by $5 per hour and you get about 16,000 hours of equivalent usage. That's the break-even point—as long as your real workload can consume that many machine-hours within two years, buying starts to beat renting; the longer and fuller it runs, the more you save.
The key lies in the words "real workload." A figure of 16,000 hours sounds like a lot, but if a GPU runs an always-on service 24/7, that's over 8,000 hours a year, so two years covers it exactly. In other words, for a stable inference service running at full load, buying your own hardware pays for itself in two years, and from the third year on it's almost pure savings. In reality, though, few workloads stay at full load. If your average GPU utilization is only 30%, those 16,000 hours will take five or six years to use up—and GPUs are due for replacement within three to five years, so the payback window may never arrive.
So the break-even judgment can be summarized into two workload profiles:
- High utilization, long-running: typical examples include company-wide internal Q&A, fixed batch-processing pipelines, and online services with continuous throughput in production. These workloads saturate the GPUs and run for a long time, so the one-time investment in owned hardware is diluted over time and is clearly cheaper in the long run. Data also never leaves the domain, so the compliance yardstick is satisfied along the way.
- Low-frequency, fluctuating, bursty: for example, quarterly report generation, occasional experimental projects, or human-machine interaction scenarios that are busy by day and idle at night. Buying a pile of GPUs for peak load that sit idle most of the year is the most expensive kind of waste; renting on demand lets costs track actual usage, and you can scale down easily when demand falls.
In real projects, both types often coexist. A pragmatic approach is a hybrid split: carry the stable, predictable core workloads that must stay on-premises on owned machines, and push the elastic, experimental, and peak portions to on-demand cloud. This holds down baseline costs without paying for fluctuations.
Two more items are easily missed when making the call. One is opportunity cost: the money spent on buying hardware is tied up and could have been invested elsewhere, and once the GPUs are bought, you're locked into that hardware generation; when the next generation doubles price-performance, you can only watch. The other is the trap of "inflated" utilization—many teams treat "the GPU is on" as "the GPU is in use," but idle VRAM produces no value, and what belongs in the payback formula is effective inference time, not power-on time. It's advisable to run on-demand in the cloud for a month or two after launch, pin down the real average daily request volume, peak-to-trough ratio, and average utilization, and then plug these measured numbers into the break-even formula above.
To sum up in one sentence: cloud vs. owned hardware is not a matter of preference but of utilization. First measure how full your workload really is and how long it runs, and let the numbers decide for you—rather than bending a procurement decision to fit a gut-feel expectation.
Scale-to-hardware reference table: turning the yardsticks into a concrete purchase order
The three yardsticks covered in the previous sections—VRAM, concurrency, and compliance—ultimately have to land on a purchase order. Below, models are divided into three tiers by parameter scale, with engineering judgments on hardware configuration and the business scenarios each tier suits.
Three tiers compared
| Parameter scale | Typical VRAM usage (FP16) | Recommended hardware | Use cases |
|---|---|---|---|
| 7B | ~14 GB | RTX 4090 (24 GB) × 1, or A10 (24 GB) × 1, or A100 40G × 1 | Internal tools, knowledge base Q&A, validation by small and midsize teams |
| 14B–40B | Increases with parameter scale | A100 40G × 2–4, or H20 × 1–2 | High-precision inference such as legal document analysis, research summaries, and radiology report generation |
| 70B–72B | ~140 GB | A100 80G × 8 (standard cluster), or H800 × 4–8 | Chinese long-text understanding, high-concurrency production, full-parameter fine-tuning |
7B: a reasonable starting point for validation
A 7B model's weights take about 14 GB of VRAM (FP16). A single 24 GB card has about 10 GB left for the KV Cache after loading the weights, which is basically enough for a dozen or so concurrent streams. The RTX 4090 suits budget-sensitive internal test environments, while the A10's ECC memory and data-center cooling specs make it better suited to always-on services. If the team already has an A100 40G, just reuse it—no need for additional procurement.
The core value of this tier is validation: Does the business process work end to end? Are there sticking points in prompt engineering? Is the data pipeline complete? Working through these questions on low-cost hardware is far more cost-effective than going straight to a large cluster and then having to start over.
14B–40B: precision requirements drive multi-GPU parallelism
When a business scenario has explicit accuracy requirements—such as comparing legal clauses, reviewing research literature, or generating medical imaging reports—7B's reasoning depth is often insufficient. Models in the 40B class have a structural advantage on these tasks, but the price is VRAM demand beyond a single GPU's capacity, which requires tensor parallelism or pipeline parallelism.
Take radiology report generation as an example: supervised fine-tuning of an LLM on specialist-annotated data can significantly improve terminology accuracy compared with the baseline. In many clinical scenarios, the size of this improvement is the dividing line for whether a system can go live. Apply the same fine-tuning strategy to 7B, and the room for improvement is much narrower because of the model's limited capacity.
On the hardware side, larger models require a corresponding increase in the number of GPUs for multi-GPU parallelism. If the budget allows, the H20 has higher memory bandwidth and delivers better throughput for long-sequence inference.
70B–72B: the top-end configuration for Chinese long text
A 70B model's weights come close to 140 GB, so FP16 loading must span multiple GPUs. A multi-GPU cluster is currently the most common production configuration—after accounting for weights and runtime overhead, there is still considerable room for the KV Cache, enough to support a certain level of concurrency with long contexts while keeping latency at the level of seconds.
Chinese-optimized 72B models (such as Qwen-72B) perform noticeably better on long-document understanding and multi-turn conversational coherence than non-Chinese-optimized versions of the same size, which significantly affects the real-world experience of deployments at Chinese enterprises. If your business involves a large amount of Chinese long-text processing, this tier is an option worth serious evaluation.
Note that the procurement and operations costs of an 8-GPU cluster are considerable. Before committing to this configuration, use the concurrency yardstick from the earlier sections to work out your actual QPS needs. If peak concurrency isn't high, choosing 70B is a decision driven more by precision requirements than by throughput requirements.
The cost-reduction path: distillation, not downgrading
After getting the business working on 70B, the real problem many teams face is that inference costs are too high, but switching directly to 7B drops accuracy unacceptably. Knowledge distillation is the engineering way out of this dilemma—using the 70B model already fine-tuned on business data as the teacher to guide the training of a 7B model. In industry practice, this path retains about 90% of performance while cutting inference compute to roughly one-tenth of the original.
Distillation is no silver bullet: it needs high-quality teacher model outputs as training signals and places extra demands on annotated data and the training process. But for teams that have already gotten their business logic working on a large model and are entering the cost-reduction phase, it is a safer path than starting selection over.
The order of selection decisions
- Set the precision floor first: the business's minimum accuracy requirement determines the lower bound on parameter scale. This step can't be based on gut feel; benchmark with real business data.
- Then calculate the concurrency budget: use the VRAM formula from Section 2 to work back from target QPS and context length to the number of GPUs required.
- Check compliance constraints: hard constraints such as data localization, MLPS requirements, and audit logs affect the deployment architecture and must be clarified before procurement.
- Assess the feasibility of distillation: if precision requirements point to a large model but cost pressure is significant, factor in the engineering cost of the distillation path before making the final decision.
There's no universal answer for hardware configuration, but the three yardsticks provide a framework into which you can repeatedly plug concrete numbers. Fill in your business parameters, and the purchase order follows naturally.
FAQ: the four most frequently asked questions in private deployment selection
I bought GPUs based on "7B = 20GB." Why did VRAM blow up right after launch?
Because that estimate counted only the floor—the weights—and not the big consumers of VRAM during inference. Whether a card can hold up depends on adding four items together: model weights, KV Cache, the footprint of the runtime framework itself, and waste caused by fragmentation.
Weights are fixed. At FP16 precision, 7B parameters really do come to around 14GB, and adding some resident overhead brings it to 20GB—that part is correct. The problem lies in the other three items. The KV Cache grows linearly with the number of concurrent requests and the context length—with a dozen or so long-context sessions running at once, this part can easily eat up well over ten GB. The inference framework itself reserves a VRAM pool and caches compiled artifacts, which also takes a few GB. Add the fragmentation from dynamic allocation, and there's often a noticeable gap between nominal free memory and what's actually usable.
So the safer approach is to work backward: first decide the concurrency and maximum context length you need to support, calculate the peak KV Cache requirement from those, add the weights and framework overhead, and finally leave a 15% to 20% margin. Working out the peak for your launch scenario is far more reliable than applying the rule of thumb "parameter count times two." If you find VRAM sitting right at the threshold, you've essentially left no margin, and the first traffic peak will break through it.
Are heavily regulated industries limited to open-source private deployment, with no option to use commercial platforms?
No. This ties the two concepts of "compliance" and "open source" together, but they fundamentally address different things. What compliance actually requires are a few hard constraints: data doesn't leave the domain, there's an audit trail, accountability can be assigned and traced, and model behavior is controllable. Whether these can be met depends on the deployment form and contract terms, not on whether the model weights are public.
Many commercial LLMs also offer private or dedicated deployment, placing the entire inference service inside your own data center or a dedicated network zone so that data never leaves the boundary. In this form, their ability to keep data within the domain is not fundamentally different from deploying an open-source model. What really needs to be checked are a few specific questions: whether the service runs entirely within your network, whether any telemetry is sent back, whether logs and audits can integrate with your existing security system, and whether the vendor can cooperate in assigning accountability when problems arise.
Conversely, open-source models aren't automatically compliant either. If you deploy an open-source model yourself, you still have to build auditing, access control, and data masking yourself. So the right question isn't "open source or commercial" but "can this delivery solution meet each of my hard constraints." Listing out your compliance checklist and comparing against it is clearer-headed than taking sides based on a model's origins.
Is a time to first token under 1 second a hard requirement? Does it apply to internal use too?
It isn't a universal hard requirement; it depends entirely on the use case. In external real-time conversations and customer service Q&A, users are staring at the screen waiting for a response, and a slightly slow first token ruins the experience—only then is latency a hard constraint. But many internal tasks don't care about it at all.
The classic counterexample is batch processing: overnight document summarization, bulk report generation, offline data labeling. These tasks care about how many items can be processed per unit of time—that is, throughput—not how quickly a single request returns. For such workloads, you can make batches larger and trade latency for throughput, getting several times more work out of the same GPUs. Configuring them to meet "first token under 1 second" would actually waste compute.
So before setting targets, distinguish the workload types: interactive workloads look at time to first token and tokens per second, while batch workloads look at overall throughput and unit cost. For mixed workloads, it's best to separate the two types of requests at the scheduling layer, reserving a low-latency lane for interactive requests and slotting batch jobs into idle periods. Holding every scenario to a single latency standard either shortchanges the experience or wastes money.
How do you choose between the cloud and buying your own hardware? Is there a simple formula?
There's a good-enough approximation: calculate utilization. The core idea is to spread the total cost of ownership of owned hardware over its lifecycle, convert it into an equivalent hourly cost, and compare that with the unit price of pay-as-you-go cloud billing.
Concretely, first estimate your real workload curve: on average, how many hours a day the GPUs are actually running. If your business runs high loads around the clock, with utilization consistently high, the unit cost of owned hardware will usually be noticeably lower than renting in the cloud, and the upfront investment can be amortized within one to two years. If the workload has distinct peaks and troughs (busy by day and idle at night), or the total volume is still small and demand is still changing, then the elasticity of the cloud is more cost-effective, since you won't keep paying for idle GPUs.
Beyond this utilization break-even point, there are several factors that don't go into the formula but still require a decision: owning the hardware means handling hardware operations, failure replacement, data center space, and power yourself; the cloud is less hassle but may cost more in total over the long run, and you also need to confirm it's acceptable from a data compliance standpoint. A safe compromise is hybrid: use owned hardware for the stable base load and cloud elasticity for peak overflow. Don't rush to place an order; draw out your workload curve for the next year, and the break-even point will reveal itself.