Teverant AI · Insights

2026-08-31

Choosing VRAM for on-premises LLMs: enterprise configurations compared

A detailed guide to calculating VRAM for on-premises LLM deployment. Taking into account model size, quantization method, context length, and concurrency requirements, it compares single-GPU, multi-GPU, Mac, and CPU options to help enterprises make sound selection and procurement decisions.

1. Choosing VRAM for the enterprise: don't look only at parameter count

When enterprises evaluate on-premises large language models (LLMs), the most common mistake is converting "parameter scale" directly into "how much VRAM is needed." The parameter count only indicates the base volume of the weight data; it doesn't represent the full footprint once the service is running. An actionable estimate has to consider four variables at once: model size, weight precision, request concurrency, and context length. Leave out any one of them, and the procurement conclusion may stop at "the model barely starts" rather than "the business runs stably."

VariableMain impactQuestions to confirm during selection
Model sizeSets the starting point for the weight footprint; more parameters usually mean more base space is needed to load the modelIs what you're deploying a full model, a distilled model, or a mixture-of-experts architecture? Are the parameters that need to be loaded the same as the nominal total?
Quantization methodChanges the storage cost per parameter, and may introduce extra scaling factors, metadata, or operator limitationsHalf precision, low-bit integer, or mixed precision? Do the inference framework and GPU support the corresponding kernels?
ConcurrencyIncreases the caches, batch data, and intermediate tensors resident at the same time, determining peak usage and the throughput ceilingDoes concurrency mean online users, simultaneous generation requests, or total queued tasks? How long do peaks last?
Context lengthThe longer the input and generated content, the larger the attention cache usually is; long-text requests keep occupying runtime spaceDoes the business use average or maximum length? Are there long inputs such as contracts, codebases, or knowledge base documents?

These four aren't independent of one another. Low-bit quantization can compress weights, but it won't eliminate runtime overhead by the same proportion; if you shrink the model but raise concurrency and the context limit at the same time, the VRAM you saved can quickly be consumed by the cache. Conversely, a model that runs fine in single-user interactive testing isn't necessarily able to handle several long-text requests generating at once. So enterprises should define the service load first and then work back to the hardware, rather than buying GPUs first and then looking for a model that will fit.

In procurement discussions, you also need to distinguish three different standards so that technical validation isn't confused with launch goals:

  • Loadable: the weights fit in VRAM, and the inference process can start and complete a small number of requests. This only verifies the minimum operating conditions and usually says nothing about whether latency, throughput, and stability meet requirements.
  • Stably usable: beyond the weights, there's room for the attention cache, the inference framework's workspace, temporary tensors, and necessary system usage, so the service can run continuously within the target context range without frequently overflowing during brief peaks.
  • Production deployment: on top of stable operation, it should cover the business's concurrency peaks and reserve room for model version changes, inference engine upgrades, prompt growth, and monitoring components. A production configuration shouldn't keep VRAM pinned at its limit over the long term.

In engineering order, VRAM is usually the first threshold on-premises inference has to clear, but it isn't the only component that determines performance. The CPU affects tokenization, request scheduling, and some operator processing; insufficient system memory affects model loading, weight mapping, and even process stability; storage performance determines cold start and model switching speed; and in multi-GPU environments, inter-GPU communication capability and topology directly affect the efficiency of sharded inference. Two machines with the same total VRAM can differ significantly in actual throughput if one relies on a slow interconnect and the other can transfer tensors efficiently.

So this section arrives at a simple rule: use model size and quantization precision to estimate the "static base," use concurrency and context to estimate the "dynamic increment," then add framework overhead and an engineering margin—that is the VRAM standard enterprises should use. The configuration comparisons that follow also revolve around this standard, rather than using whether a certain GPU can start a certain model as the final basis for purchasing.

2. How to calculate VRAM: model weights are only the first item on the bill

In on-premises LLM deployment, the most common misjudgment is treating the model file size as the GPU requirement. The model file only reflects the static weights; actual inference also has to leave room for intermediate tensors during input and output, the context cache, and runtime management structures. So "the file fits" only means the model may load successfully—it doesn't mean it can reliably serve requests.

As a first step, you can estimate weight VRAM as follows:

Weight VRAM ≈ parameter count × bytes per parameter

Storage or compute precisionRough conversionUseful for judging
FP16About 2 bytes per parameterUnquantized deployments or those that prioritize precision
INT8About 1 byte per parameterOptions that need a smaller footprint while keeping good quality
INT4About 0.5 bytes per parameterDeployments with limited VRAM that can accept some quantization loss

For example, for a model with 7B parameters, the theoretical FP16 weight size is about 14GB, INT8 about 7GB, and INT4 about 3.5GB. The word "about" matters here: quantized formats usually also include scaling factors, grouping information, and alignment overhead, so the actual file size won't strictly equal the parameter count times the ideal bytes. Format conversion or extra copies may also occur during model loading, so you can't compare this result directly with a GPU's nominal VRAM.

Record three types of VRAM separately

A more practical approach is to break out three metrics in your deployment sheet, rather than recording a single model size.

  • Static weight VRAM: the base portion occupied long-term after the model is loaded, usually determined by parameter scale, quantization method, and loading format.
  • Per-request peak VRAM: the extra VRAM a single request consumes during prefill and generation, affected by input length, output length, and model architecture.
  • Peak VRAM at target concurrency: the total peak when the business's concurrency target is reached, mainly affected by the KV Cache, the distribution of request lengths, and the scheduling strategy.

Beyond the weights, activations are temporary space used during computation and are usually more pronounced in the prefill stage; the KV Cache grows with context length and the number of requests processed simultaneously. For workloads such as long-document Q&A, code completion, and batch summarization, the KV Cache is often more likely to become the bottleneck than the short-lived activations of a single request. The inference framework itself also takes up some VRAM, including operator workspaces, communication buffers, memory management structures, and necessary temporary tensors.

If all you have right now is the model's parameter count and quantization format, you can add about 30% on top of the weight figure as a screening line for low concurrency and typical context lengths. For example, if weights are estimated at 14GB, look at a preliminary budget of about 18GB rather than treating a 14GB GPU as an adequate configuration. This is only an early screening tool, not a production capacity commitment. For long contexts, multi-user parallelism, or highly variable output lengths, the extra space needed may well exceed this proportion.

For formal planning, it's advisable to set up a simple capacity log: model version, quantization format, weight footprint, input length, maximum output length, per-request peak, target concurrency, and total peak. Load testing should cover at least short requests, typical requests, and requests near the limit, and record peak VRAM rather than averages. Only by bringing the real distribution of business requests into testing can you tell whether VRAM runs out first or throughput and queuing latency hit the bottleneck first.

Using an inference framework with PagedAttention lets you manage the KV Cache in pages, reducing fragmentation and improving VRAM utilization across concurrent requests. It addresses the efficiency of cache allocation and reuse; it doesn't reduce the resident footprint of the model weights themselves. If the weights already exceed a single GPU's capacity, switching frameworks can't substitute for quantization, sharding across GPUs, or new hardware. Final capacity should be based on load-test results for the target model, context limit, and concurrency.

3. Model size × quantization: an enterprise VRAM configuration reference table

The parameter count only determines the approximate size of the weights; it can't be equated directly with the VRAM a machine needs. The figures below are estimated on an engineering basis of "weights loaded onto the GPU, a common inference framework, and basic runtime headroom reserved"; they are suitable for initial procurement screening but should not replace real load testing. Actual usage is also affected by quantization format, inference framework, context length, and the number of concurrent requests.

Model tierINT4 weightsFP16 weightsRecommended configurationFit
7BAbout 4–5GBAbout 16GB8GB minimum; 12–16GB more stableOffice Q&A, summarization, text classification, lightweight RAG
13B–14BAbout 9GBAbout 28GB12GB is worth a try; 16GB better for long-term operationMore complex knowledge base Q&A, structured content generation
32BAbout 20GBAbout 64GB24GB is a runnable tier; 32GB gives more headroomScenarios that prioritize reasoning quality, context, and stable single-user response
70BAbout 40GBAbout 140GBINT4: 48GB or more recommended; FP16 usually multi-GPUComplex analysis, higher-quality generation, businesses with clear model capability requirements

7B is the easiest tier to deploy. The INT4 weights themselves take about 4–5GB, but 8GB of VRAM doesn't mean every scenario will run stably: with longer contexts or several requests handled at once, cache and framework overhead quickly squeeze the available space. So 8GB is better suited to single-user, short-input internal tools; when light concurrency is needed, 12–16GB is more reasonable. The FP16 version usually needs about 16GB of weight space, and it's best to leave extra headroom for actual operation.

13B–14B is a common middle ground for enterprise workstations. INT4 needs about 9GB, and a 12GB GPU can handle a basic deployment, but with little headroom; 16GB is better for keeping longer conversation history, incorporating retrieval results, or handling a small amount of concurrency. If you insist on FP16, VRAM requirements rise to about 28GB, and a single ordinary consumer GPU is usually no longer suitable.

32B's INT4 weights are about 20GB. It can run on 24GB of VRAM, but that's a "can start, needs careful control" tier: overly long contexts, larger batches, or serving several requests at once may all hit the limit. 32GB makes it easier to leave room for cache and runtime. The FP16 version is about 64GB and usually requires a high-VRAM single GPU or sharding across multiple GPUs—it shouldn't be purchased as an ordinary workstation setup.

Even with INT4, 70B weights come close to 40GB, and from an engineering standpoint 48GB or more should be seen as a more reliable starting point. FP16 is about 140GB and often requires two 80GB-class data center GPUs, or multi-GPU parallelism and sharded deployment. Here, "runs on multiple GPUs" doesn't equal "usable by multiple users": inter-GPU communication, bandwidth, and framework compatibility all affect actual throughput.

In terms of selection order, first determine the model tier the business permits, then choose among INT4, INT8, and FP16. INT4 suits deployments with limited VRAM focused on generation and Q&A; INT8 is usually a more conservative balance between resources and precision; FP16 is mainly for scenarios with ample VRAM, a need for more stable numerical behavior, or existing data center resources. The VRAM figures in the table are only a starting boundary; the final answer still has to be confirmed by load testing at the target context length and target concurrency.

4. Concurrency and context: why "it runs" doesn't mean "it's ready for production"

A short single-turn conversation returning normally only proves that the model weights have been loaded into VRAM; it doesn't show that the machine is suitable for providing service over the long term. Weights usually stay relatively stable after the model is loaded; what actually changes with requests is the KV Cache kept during inference. The longer each request's input and the more it outputs, the larger the cache footprint; the more requests processed at once, the more the cache stacks up per request. The number of model layers, the attention head configuration, and the data precision used by the KV Cache also change this overhead.

So the VRAM budget should be split into at least three parts: model weights, the KV Cache, and extra space for the inference framework and runtime. The last two determine the gap between "loaded successfully" and "runs stably." A test that inputs only a few hundred characters and serves one person at a time often won't reveal the pressure from long-document retrieval, follow-up questions, and multiple people making requests at once. This is especially true for RAG: retrieval results rapidly lengthen the input context, and if the model is also required to produce fairly complete answers, the cache keeps growing during generation.

Business typeMain riskLoad-testing focusConfiguration approach
Personal or internal assistantFew requests, but occasional long documents and long answersSingle-user long context, continuous multi-turn conversationEnsure per-request headroom first, then set a context limit
Department-level knowledge base Q&AMany people retrieving at once, with significant variation in input lengthConcurrent retrieval, peak VRAM and latency during generationDesign for working-hours concurrency and leave room for queuing
Customer service or model APIConcentrated request peaks; timeouts create backlogsPeak concurrency, sustained throughput, response time after queuingDefine the service level first, then work back to VRAM and instance count

Using a rule of thumb like "each additional concurrent request adds a fixed number of GB" isn't recommended. Different models have different numbers of layers and attention structures, quantization formats may compress only the weights and not the cache, and how the inference framework implements cache allocation and batching also affects the result. A more reliable approach is to lock down the model file, quantization scheme, inference framework, maximum input length, and maximum output length, and then test concurrency step by step—for example, starting at 1, 2, 4, and 8 concurrent requests. At each step, record peak VRAM, time to first token, full response time, generation throughput, and the proportion of requests that fail or queue.

Testing should also cover two kinds of boundaries: short inputs with high concurrency, to see whether the cache is quickly exhausted when many people access the system at once; and long inputs with low concurrency, to verify whether a single request triggers OOM at the context limit. If the business uses RAG, also fix the number of retrieved chunks or a total token cap; otherwise the input size varies between tests and the conclusions can't be compared. The final configuration shouldn't look only at averages but at peak VRAM during busy periods and the slowest requests.

When VRAM is close to the limit, the usual order of response is to first tighten the maximum context and maximum generation length, then limit the batch size or the number of requests processed at once; for traffic bursts, you can introduce queuing so the system trades controlled waiting for not crashing. For frameworks that support paged management of the KV Cache, mechanisms like PagedAttention are worth evaluating. The vLLM project designed this mechanism specifically to reduce cache fragmentation and improve VRAM utilization under concurrency, but it can't eliminate the resource demands that come with model weights, long contexts, and high concurrency themselves.

Only when the business still can't reach its target throughput after limiting context, controlling concurrency, and optimizing cache management should you consider switching models, adding GPUs, or splitting the service. Conclusions about VRAM selection should come from load-test curves under fixed conditions, not from a single demo that "can generate an answer."

5. Choose configurations by enterprise use case, not by blindly chasing bigger models

VRAM planning should work backward from a task's cost of errors, call frequency, and data shape, rather than starting by settling on the largest model. Internal Q&A, contract review, code generation, and core API services have different requirements for model capability, response stability, and concurrency. The same model can correspond to completely different hardware budgets when used by one person on a trial basis versus accessed by many people at once.

Use caseRecommended model and quantizationStarting VRAMConfiguration judgment
Office writing, summarization, basic knowledge Q&AAround 7B, INT48GBSuitable for low-concurrency validation; 12–16GB better for formal use
Confidential documents, initial contract screening, department-level RAG13B–14B, INT416GBPrioritize stable latency and output consistency; don't sacrifice usability for a bigger model
Code generation, mathematical reasoning, structured extraction14B minimum; consider 32B for high-quality needs, INT416GB; upgrade to 24–32GBMust verify with the enterprise's real task set whether a model upgrade delivers actual gains
Complex analysis, general reasoning, core APIsAround 70B, INT448GB or moreSuited to high-value services with higher quality requirements; for low-frequency calls, calculate the total cost of cloud versus on-premises

1. Office tasks: 8GB is the validation line, not necessarily the launch line

Email rewriting, meeting minutes, document condensing, and Q&A on company policies usually don't require deploying a large-parameter model from the start. A 7B-class INT4 model can serve as an entry-level option, and 8GB of VRAM is enough for single-user or low-concurrency validation. But an actual deployment also has to accommodate longer contexts, retrieval-augmentation components, the serving framework, and future version switches, so 12–16GB usually leaves headroom more easily. The "headroom" here isn't about chasing higher specs; it's about avoiding a situation where the model barely fits and a slightly longer input causes a VRAM overflow.

2. Confidential documents and departmental knowledge bases: ensure control first, then pursue parameter scale

Initial contract screening, policy retrieval, and department-level knowledge Q&A usually involve fixed formats, specialized terminology, and sensitive internal content. A 13B–14B INT4 configuration with 16GB of VRAM is a fairly safe starting point. Rather than compressing a larger model until it has almost no room to run, leaving enough VRAM for context, vector retrieval, and concurrent requests is often better for controlling latency. During engineering acceptance, focus on whether citations are accurate, whether refusal boundaries are stable, and whether answers to similar questions are consistent, rather than looking only at general benchmark scores.

3. Code and complex reasoning: VRAM upgrades must correspond to observable gains

Code generation, math problems, and complex structured extraction are more sensitive to the completeness of the reasoning chain and the ability to follow formats. You can start from 14B with 16GB of VRAM; if internal testing shows the smaller model frequently fails on key tasks, then evaluate 32B INT4 with 24–32GB of VRAM. The test set should come from real code repositories, historical tickets, financial spreadsheets, or business documents, and should record first-pass rate, manual editing time, format error rate, and response time. If a bigger model doesn't reduce rework, you shouldn't buy more VRAM on the strength of parameter count alone.

4. Core inference services: evaluate a 70B plan together with call economics

Only complex analysis, cross-document judgment, and internal APIs with high quality requirements justify considering a 70B INT4-class deployment, which usually needs 48GB of VRAM or more. With low-frequency calls, local multi-GPU equipment may sit idle for long periods; with peak-period calls, you also have to consider VRAM usage, inter-GPU communication, failover, and operations staffing. So local equipment depreciation, electricity, data center space, spare parts, and maintenance costs should be compared side by side with pay-as-you-go cloud GPU fees in the same table. The stronger the model, the less you can use "whether it starts" as the acceptance criterion.

The final choice should be settled by business load testing: fix the model version, quantization method, context length, and concurrency, and record time to first token, full response time, peak VRAM, and error rate for each before deciding whether more VRAM is needed. What you get this way is a task-oriented configuration, not a GPU list divorced from actual use.

6. Weighing single-GPU, multi-GPU, Mac, and CPU options

The number of GPUs isn't a direct measure of deployment capability. The value of a single GPU lies in a short pipeline, less configuration, and fewer points of failure; the value of multiple GPUs lies in splitting up a model that won't fit on one card, or increasing throughput through parallelism. Apple Silicon and CPU-only setups are better suited to specific cost, power, or offline constraints. When selecting, first settle the model, quantization format, context length, and concurrency target, then decide on the hardware form.

OptionSuitable model rangeMain advantagesIssues to watch out for
8–16GB single GPU7B–14B, mainly INT4Short deployment pipeline, easy latency controlAs context and concurrency grow, the KV Cache quickly eats the headroom
24–32GB single GPU32B INT4; 14B FP16Suitable for single-machine serving of mid-sized modelsLong-context, multi-request scenarios still need load testing
Single GPU around 48GBEntry configuration for 70B INT4No dependence on cross-GPU communication, relatively simple maintenance"Can load" doesn't mean it can meet target concurrency
Multi-GPUWhen a single GPU can't hold the model, or throughput needs to scaleCan scale out VRAM and computeFramework, interconnect bandwidth, and communication overhead determine the actual benefit
Apple Silicon unified memoryQuantized 32B; 70B worth trying with larger memoryMemory shared between CPU and GPU, suitable for local R&D and low-concurrency useServing tools, operator support, and concurrency capability must be verified separately
CPU + system memorySmall models such as 7B INT4No discrete GPU needed; can serve as an offline or disaster-recovery nodeGeneration speed is usually unsuitable for real-time multi-user access

Single GPU: minimize launch complexity first

If the model fits on a single GPU with enough headroom, a single GPU is usually the enterprise's first choice. With 8–16GB of VRAM, you can cover many 7B to 14B INT4 deployments; 24–32GB is better for fitting 32B INT4 onto one card or running the FP16 version of 14B. Only when VRAM reaches around 48GB do you enter the feasible range for quantized 70B.

The "headroom" here can't be calculated from weight size alone. The inference process also allocates runtime buffers, CUDA workspace, and the KV Cache. If the goal is a fixed single user with short contexts, barely fitting may be acceptable; if you need to serve many people's requests, include VRAM utilization, time to first token, and sustained generation speed in acceptance, rather than only checking whether the model loads successfully.

Multi-GPU: VRAM can work together, but it doesn't simply add up

Multiple GPUs address two kinds of problems: model weights and runtime memory exceed a single card's capacity, or a single card's throughput is insufficient. Before adopting tensor parallelism, you must confirm that the inference framework and model implementation support that form of parallelism; some software can only have different processes each use one GPU and can't efficiently split a single model.

Cross-GPU transfers add extra cost. When PCIe lanes are narrow, frequent synchronization between layers or tensors can cancel out the added compute; servers with high-speed interconnects are usually better suited to this kind of deployment. Nor can the VRAM of two cards be unconditionally merged in every framework—actual usable space depends on the sharding strategy, VRAM allocation method, and communication buffers. Before purchasing, run a real model through the target framework rather than just adding up VRAM numbers.

Mac and CPU: for low concurrency, offline work, and R&D

Apple Silicon's unified memory can be used by both the model and the graphics processor, so a machine with 64GB of unified memory can run a quantized 32B model and attempt a fairly tight 70B configuration; the 128GB tier leaves more comfortable room for a quantized 70B. However, unified memory isn't the same as professional GPU VRAM. After the model loads, it still competes for memory with the operating system, the context cache, and application processes, and the inference framework's operator coverage, batching capability, and serving interfaces need to be tested separately.

CPU plus system memory is a viable fallback. Running 7B INT4 with 32GB of memory suits offline document processing, development and debugging, low-frequency Q&A, or disaster-recovery nodes; but generation speed is usually only a few tokens per second, a clear gap from the dozens or more that GPUs commonly achieve. So a CPU option should be clearly positioned for low-frequency tasks or disaster recovery and shouldn't directly handle real-time, high-concurrency interfaces.

A practical rule: for a single GPU, look first at stability; for multiple GPUs, at the framework and interconnect; for a Mac, at the ecosystem and memory headroom; for CPU, at whether the business can tolerate waiting. The final choice should be decided by the measured throughput and latency of the target model, not by the device's theoretical total VRAM.

7. Procurement and launch: confirm final VRAM with load-test results

GPU procurement shouldn't start from "can this card load the model" but from "under what conditions will business requests push VRAM to its limit." A model starting successfully only shows that the weights and basic runtime environment fit; it doesn't prove it can withstand real users' concurrency, long contexts, and complex toolchains. Before purchasing, it's best to freeze an anonymized test set and have candidate hardware, model versions, inference frameworks, and quantization configurations all use the same inputs, to avoid faulty comparisons caused by different test conditions.

The test set can't consist of just a few short Q&A pairs. It should cover at least the following kinds of requests:

  • Short prompts and long prompts near the business limit, to observe how context length affects the KV Cache;
  • Typical answer lengths and longer outputs, to avoid estimating generation-stage resource use from average responses alone;
  • Q&A with retrieved chunks, including the number of documents, chunk length, and metadata in the test;
  • Tasks that require calling external tools, executing in steps, or returning structured results, to verify whether intermediate states cause extra peaks;
  • Sensitive business tasks involving permission checks, contract review, financial data processing, and the like, to confirm that the actual prompt templates don't change resource requirements.

Each scenario should be tested repeatedly, rather than recording only a single best result. It's advisable to group requests by target concurrency, observe performance at single-request, low-concurrency, and target-concurrency levels, and keep the raw logs. Acceptance records should include at least: VRAM usage after the model finishes loading, peak VRAM during a single request, peak VRAM at target concurrency, time to first token, subsequent generation speed, overall throughput, and failures such as timeouts, abnormal exits, and incomplete results.

What to observeKey judgmentFailure signals
VRAM after loadingWhether there's headroom for weights, runtime components, and base cacheAlready near capacity right after startup
Peak VRAMWhether long inputs, long outputs, and retrieval injection change resource boundariesOccasional requests trigger OOM outright
Concurrency performanceWhether throughput still meets business targets as concurrency increasesQueuing time rises rapidly or latency jitter is significant
Failure rateWhether the service still returns reliably under resource pressurePersistent timeouts, truncation, or retries

During load testing, pay special attention to the system's "degradation path." When VRAM runs high, does the inference framework move some data to system memory, does that trigger CPU offloading, or does it simply make time to first token and generation speed suddenly slow down? The latter two may not raise errors right away, but they can break your production SLA. Also simulate peak requests arriving together with ordinary ones, and observe whether a single long-context task slows down the whole batch.

The final configuration can't be purchased to fit snugly around the highest usage seen during testing. The peak is a risk boundary, not a recommended long-term operating point; you also need to reserve a safety margin for drivers, process fluctuations, cache growth, hot model updates, and sudden concurrency. How much margin to leave should be decided based on test results, the latency the business can tolerate, and the expansion cycle, rather than applying a fixed ratio divorced from the scenario.

Procurement evaluation also can't compare GPU unit prices alone. Total cost should include whole-system power consumption and power supply specs, cooling capacity, available chassis space, and how well drivers and inference frameworks are supported; if multiple GPUs are needed, also check inter-GPU communication, motherboard lanes, and deployment complexity. Finally, estimate operations effort and the cost of future model upgrades: larger models, longer contexts, or new quantization schemes may change VRAM and software stack requirements. Only by putting this round of load-test results together with the full operating cost can you judge whether a configuration merely "starts" or is truly ready for stable production.

8. FAQ: four common questions about choosing VRAM for on-premises deployment

Can you really deploy an enterprise LLM on 8GB of VRAM?

Yes, but more precisely: 8GB of VRAM is suitable for deploying compressed small models or handling low-concurrency, short-context internal tasks; it doesn't mean it can reliably host every enterprise model. VRAM first has to hold the model weights, and at runtime it also has to reserve space for the context cache, inference framework, temporary tensors, and concurrent requests. Weights that just barely fit usually mean that as soon as the context gets longer or a second request comes in, it starts to overflow.

In practice, you can judge in three steps: first confirm the model's quantization format and file size, then leave headroom for runtime overhead, and finally test consecutive requests with real business prompts. Devices with 8GB are usually better suited to short-input tasks such as knowledge base Q&A, text classification, and field extraction; for code generation, long-document summarization, and multi-user concurrency, prioritize adding VRAM or reducing model size. Don't just check that "the model loads"—check whether it can keep serving at the target context length.

Are two 16GB GPUs equivalent to one 32GB GPU?

No. Two cards can jointly host a model through model sharding or tensor parallelism, but the usable capacity doesn't deliver the experience of a single card with the combined capacity. Cross-GPU transfers add communication overhead, and motherboard slots, PCIe bandwidth, the interconnect between GPUs, and the inference framework's parallelism implementation all affect the outcome; some frameworks, because of their allocation strategy, also cause one card to hit its limit first.

If the model fits entirely on one 16GB card, two cards mainly add concurrency capacity rather than automatically doubling single-request speed. If the model must be split across two cards, confirm that the target framework supports the corresponding quantization format and multi-GPU inference, and check whether VRAM is balanced. Before purchasing, it's best to load test with the target model, target context, and target concurrency; "total VRAM meets the requirement" only shows that a feasible path exists—it doesn't directly represent production performance.

Does INT4 quantization noticeably hurt business results?

It does have an impact, but how much depends on the task, not on the label "INT4" alone. Quantization represents weights at lower precision, which usually reduces VRAM usage significantly and lowers the barrier to on-premises inference. The trade-off is that on complex reasoning, rigorous code, long-horizon planning, and edge cases, error rates may rise more readily than with higher-precision versions.

For ordinary retrieval Q&A, summarization, classification, and structured extraction, INT4 is often a starting point worth validating first. For high-risk processes such as financial calculations, specialized regulations, and code review, keep a high-precision baseline alongside it, and use a set of anonymized business samples to compare factual accuracy, format compliance, refusal quality, and the share of manual rework. If quality falls short, you can first switch to a more robust quantization scheme or increase model size, rather than immediately buying a bigger GPU. Q8 is usually closer to the original model's quality, but it also costs more in VRAM and throughput.

Without a discrete GPU, can you deploy using system memory or a Mac's unified memory?

Yes. CPU inference frameworks can place the model in system memory, and Apple Silicon can use unified memory to hold the model and runtime data. However, system memory capacity is only the condition for "fitting"; memory bandwidth and processor capability determine how well it "runs." Compared with GPUs, CPU options usually have weaker time to first token and sustained generation speed, making them suitable for low-frequency, offline, single-user tasks but not as the default configuration for real-time multi-user services.

On ordinary computers, focus on memory headroom, swap space, and stability during long runs; don't let the model fill all available memory, or the system and framework will have no buffer. For a Mac, estimate usable capacity by subtracting the operating system, applications, and context cache from total unified memory. Mac models with more unified memory can run bigger quantized models, but response speed still needs to be tested with real prompts. If the business requires stable concurrency, prioritize discrete GPUs or dedicated multi-GPU hosts rather than judging by memory capacity alone.