Model quantization replaces 32-bit floating-point weights and activations with lower-precision formats like INT8 or INT4, cutting memory footprint and compute cost at the price of some accuracy. Done well, INT8 gets you close to a [4x reduction in model size

and memory bandwidth](https://docs.nvidia.com/deeplearning/tensorrt/10.x.x/architecture/capabilities.html) versus FP32, while 4-bit and FP8 push further. The rule of thumb: pick weight-only quantization when you’re memory-constrained, and weight+activation quantization when you need raw throughput.


TL;DR:

  • INT8 quantization offers about a fourfold reduction in model size and memory bandwidth compared to FP32, with throughput improvements of two to four times on supported hardware.
  • Post-training quantization (PTQ) is generally faster and more suitable for large models when retraining isn’t feasible, but it depends heavily on representative calibration data to avoid accuracy loss.
  • Weight-only 4-bit quantization is easiest to implement and offers significant savings, ideal for constrained hardware, while weight+activation schemes prioritize throughput but require kernel support.
  • Techniques like SmoothQuant, GPTQ, AWQ, and FP8 each address specific challenges like outliers, accuracy preservation, or hardware support, and should be selected based on the target deployment environment.
  • Compatibility between quantization algorithms, runtime support, and formats like GGUF or SafeTensors is crucial; mismatches can cause performance issues or prevent models from executing correctly.

Table of Contents

Why quantize: benefits, precisions, and hardware limits

We quantize models for one practical reason: fewer bits per parameter means less memory, less bandwidth, and often faster inference. A 70-billion-parameter model in FP32 needs roughly 280 gigabytes just to hold its weights. Drop to INT8 and that number falls dramatically, since INT8 quantization provides up to a 4x reduction in model size and memory bandwidth compared to FP32, with throughput gains of 2 to 4 times faster compute on supported hardware. INT4 and weight-only schemes compress further still, often at a steeper accuracy cost depending on the model and task.

The precision you choose depends on what you’re optimizing for:

  • FP16/BF16 serves as the common baseline for training and a safe fallback when quantized accuracy isn’t acceptable.
  • INT8 is the standard compression target for both CPU and GPU deployment, with mature tooling support.
  • FP8 targets throughput on newer GPUs that have native FP8 tensor cores.
  • INT4 and weight-only formats maximize memory savings, which matters most when you’re trying to fit a large model onto a single consumer GPU or edge device.

Hardware constrains what’s actually achievable. CPUs rely on vector instructions tuned for INT8 arithmetic, while GPUs need specialized kernels that match the quantization scheme you picked. A model quantized for one runtime’s qengine or qconfig won’t necessarily run efficiently, or at all, on another. This is the detail that trips up a lot of teams: they quantize successfully in a notebook, then discover the target runtime has no kernel for that specific bit-width and packing format.

Core approaches: PTQ vs QAT, and how to decide

Two workflows dominate practical quantization work, and picking between them is mostly a question of how much accuracy you can afford to lose and how much compute you’re willing to spend getting it back.

Post-training quantization (PTQ) takes a trained model and converts its weights (and sometimes activations) to lower precision using a small calibration dataset to estimate value ranges. It’s fast, usually a matter of hours, and it’s the default approach for large language models where retraining is expensive or the original training data isn’t available. The quality of your calibration set matters enormously here: a calibration set that doesn’t represent your real inference traffic will produce quantization ranges that are wrong for the inputs you actually care about.

Quantization-aware training (QAT) simulates the effects of quantization during training or fine-tuning, letting the model adjust its weights to compensate for the precision loss ahead of time. PyTorch’s quantization documentation describes QAT alongside dynamic and static PTQ as one of three core quantization workflows, and notes that PyTorch 2’s export quantization path improves full-graph capture for these flows. QAT typically yields better accuracy retention than PTQ, but it costs a full or partial retraining cycle.

When choosing between them, weigh three factors:

  1. Accuracy sensitivity. If your task tolerates a few points of degradation, PTQ is usually sufficient and far cheaper.
  2. Access to training data and compute. QAT only makes sense if you can retrain, which rules it out for many third-party or open-weight models.
  3. Deployment velocity. If you need to ship this week, PTQ wins almost every time; QAT is an investment for models you’ll maintain long-term.

Weight-only vs weight+activation quantization in production

The split between weight-only and weight+activation quantization is the single biggest architectural decision in a quantization project, and it maps directly to what’s bottlenecking your deployment.

Two paths for model quantization

Weight-only quantization (the approach behind GPTQ and AWQ) compresses only the model’s weights, typically to 4 bits, while leaving activations in higher precision during computation. This gives you the largest memory savings with the simplest offline workflow, since you don’t need to calibrate activation ranges across a representative traffic sample. For large models, weight-only approaches often retain accuracy better than you’d expect, because the activations, where a lot of quantization error originates, stay untouched.

Weight+activation quantization (W8A8, or FP8 schemes) quantizes both sides of the matrix multiply, which is what actually unlocks faster compute, not just smaller storage. The catch is that it demands real runtime and kernel support, plus more careful calibration, since activation distributions are harder to predict than weight distributions.

  • Choose weight-only 4-bit when you’re deploying to a single GPU or edge device and memory is your hard constraint.
  • Choose W8A8 or FP8 when you’re serving at scale and inference latency or server throughput is what you’re being measured on.
  • Reserve FP16 fallback for layers where either approach causes unacceptable quality loss.

As the Intel documentation on weight-only quantization puts it, weight-only methods fit best when GPU memory is the bottleneck, while W8A8 or FP8 variants make more sense when latency and throughput are the priority and the hardware kernels exist to support them.

Pro Tip: Start with weight-only 4-bit quantization for any model you’re trying to fit onto cheaper hardware. It’s the lowest-effort path to real savings, and you can layer in activation quantization later if throughput becomes the bottleneck.

State-of-the-art methods: SmoothQuant, GPTQ, AWQ, and FP8

Four names come up constantly in quantization literature, and each solves a different piece of the accuracy-versus-efficiency puzzle.

SmoothQuant addresses the core problem with W8A8 quantization: activation outliers that blow up quantization error. It works by mathematically shifting the quantization difficulty from activations to weights through a smoothing transformation, enabling full INT8 post-training quantization without retraining. The original SmoothQuant paper reports up to a 1.56x inference speedup and roughly 2x memory reduction for large language models with negligible accuracy loss in its experiments.

GPTQ takes a weight-only approach, using second-order information (an approximation of the loss curvature) to reconstruct weights layer by layer after quantization, minimizing the error introduced at each step. It has become a go-to method for compressing large language models to 4 bits without the cost of retraining, and it preserves accuracy well across a wide range of model families.

AWQ (Activation-aware Weight Quantization) builds on the insight that not all weights matter equally: a small fraction of weights, identified by looking at activation magnitudes, disproportionately affect output quality. AWQ protects these salient weights during quantization rather than treating every weight identically. According to the AWQ paper, this approach often outperforms GPTQ on instruction-tuned models and reports speedups exceeding 3x on some hardware implementations.

SmoothQuant+ extends the smoothing idea specifically to group-wise 4-bit weight-only PTQ.

FP8 is the newest entrant gaining real tooling support. NVIDIA’s TensorRT documents FP8 support alongside INT8 and INT4 for server-class GPUs, positioning it as a middle ground: better accuracy retention than INT8 in many cases, with throughput closer to what you’d get from aggressive integer quantization. A broad experimental evaluation across 7B to 405B parameter models found that quantized larger models frequently outperform smaller FP16 models on many benchmarks, though outcomes depend heavily on the specific method, model size, and bit-width chosen, and the same analysis found quantized models can still underperform on tasks like instruction-following and hallucination detection even when other metrics look fine.

State-of-the-art methods: SmoothQuant, GPTQ, AWQ, and FP8 — overview diagram

Tooling and runtime support: PyTorch, Optimum, TensorRT, vLLM, GGUF

Picking an algorithm is only half the job. The runtime you deploy to determine which formats and kernels are actually available to you, and mismatches here are where projects stall.

  • PyTorch supports dynamic quantization, static PTQ, and QAT natively, and PyTorch 2’s export quantization workflow improves full-graph capture, which makes it easier to apply consistent quantization across an entire model graph rather than module by module.
  • Hugging Face Optimum provides helper APIs that wrap INT8 and other quantization flows for popular model families, reducing the boilerplate needed to go from a trained checkpoint to a quantized one.
  • NVIDIA TensorRT documents FP8, INT8, and INT4 support through its Model Optimizer, aimed at high-throughput GPU serving where every millisecond of latency matters.
  • vLLM focuses on efficient serving of quantized LLMs at scale, commonly paired with GPTQ or AWQ checkpoints.
  • GGUF and SafeTensors serve different purposes: GGUF is the common format for local, CPU-friendly inference with built-in quantization variants, while SafeTensors stores full-precision weights in a format built for safe, fast loading in GPU serving and training pipelines.

Before you commit to a toolchain, run through this compatibility checklist:

Check Why it matters
qconfig/qengine match A quantized model built for one backend’s qengine may not run on another
Kernel availability Some bit-widths (INT4, FP8) need specialized kernels your target hardware may lack
Format conversion Converting between GGUF, SafeTensors, and framework-native formats can introduce errors
Framework version Quantization APIs change fast; pin versions tested against your target model

Practical recipes: quantizing to INT8 and 4-bit weight-only

Here’s a working checklist for taking a model from full precision to a quantized, deployable artifact.

Before you quantize anything:

  1. Record baseline metrics: latency, memory footprint, and task accuracy on a held-out set.
  2. Assemble a calibration dataset that mirrors real inference traffic, not a generic benchmark sample.
  3. Decide your target precision based on the weight-only versus weight+activation tradeoff above.

INT8 PTQ recipe:

  1. Load the trained model in its native framework (PyTorch or a Hugging Face checkpoint).
  2. Run calibration: pass a few hundred representative samples through the model to collect activation statistics, following the static PTQ flow described in PyTorch’s quantization documentation.
  3. Apply SmoothQuant-style activation smoothing first if you’re seeing large outliers in specific layers, since this is what makes W8A8 PTQ viable without retraining.
  4. Convert to INT8 using your framework’s conversion API or Hugging Face Optimum’s quantization helpers.
  5. Validate against your baseline metrics before shipping.

4-bit weight-only recipe (GPTQ or AWQ):

  1. Choose a toolchain (AutoGPTQ-style libraries or AWQ-specific implementations) matched to your model family.
  2. Run the quantization pass, which reconstructs weights layer by layer using calibration data.
  3. Export to a serving format compatible with your runtime (vLLM, for example, has first-class support for GPTQ and AWQ checkpoints).
  4. Evaluate on downstream tasks, not just perplexity, since weight-only methods can show deceptively small perplexity shifts while task accuracy moves more.

If quality regresses past your threshold, don’t throw out the whole effort. Try mixed precision first: keep a handful of sensitive layers (often the first and last few) in FP16 while quantizing the rest.

Pro Tip: Always keep your FP16 baseline checkpoint on hand as a rollback path. Quantization regressions are often isolated to a few layers, and a hybrid approach beats reverting the entire model.

Evaluation and benchmarks: what to measure and how to read results

Quantization claims in papers rarely transfer cleanly to your model and your traffic, so build your own evaluation suite rather than trusting a single published number.

Track these dimensions together, not in isolation:

  • Latency and throughput, measured under realistic batch sizes and sequence lengths, not synthetic single-request tests.
  • Memory footprint, both model size on disk and peak runtime memory during inference.
  • Perplexity, as a quick sanity check, though it’s a weak proxy for real task quality.
  • Downstream task accuracy across a diverse suite spanning question answering, math reasoning, and instruction following.
  • Hallucination and instruction-following behavior, which the large-scale evaluation across 7B to 405B models found can degrade even when other metrics hold steady.

When you read a benchmark claim from a paper or vendor, check the model family, whether it was instruction-tuned, and what calibration data was used. These details shift results more than the quantization method name alone would suggest, and a number that looks impressive for a base model can mean something very different for a chat-tuned one.

Best practices and troubleshooting common failures

Most quantization failures trace back to a handful of recurring causes, and most have straightforward fixes.

  • Unrepresentative calibration data produces quantization ranges tuned to the wrong distribution. Pull calibration samples from actual production traffic, and resist the urge to reuse a convenient public benchmark instead.
  • Activation outliers in specific layers can blow up quantization error for W8A8 schemes. Smoothing techniques like those in SmoothQuant redistribute this difficulty into the weights, and selective clipping can help when smoothing alone isn’t enough.
  • Kernel and runtime mismatches cause models that quantize cleanly to fail or run slowly in production. Confirm your qengine and target runtime actually support the bit-width and packing format you chose before you invest time in calibration.
  • Packing errors, especially with 4-bit formats, can silently corrupt weights if the export step doesn’t match what your inference runtime expects.

Pro Tip: When accuracy regresses after quantization, diagnose layer by layer before reaching for a full revert. Reverting two or three sensitive layers to FP16 often recovers most of the lost quality while keeping the bulk of your memory savings.

If none of these fixes close the gap, widen your calibration diversity first, then consider whether QAT is worth the retraining cost for this particular model.

How an engineering partner approaches a quantization project

When taking on a quantization project, it’s best to follow a clear sequence: audit the current deployment and baseline metrics, prototype a quantized version against representative traffic, bench it against production-like load, then deploy with a rollback path in place. That order matters because skipping the audit step is where most quantization efforts go sideways, teams quantize first and only discover their real bottleneck (memory, latency, or cost) after the fact.

Our AI Engineering Assistance, AI Code Reviews & Optimization, and Local AI Model Development services map directly onto this flow: engineering assistance for the prototype and benchmarking work, code review for catching kernel and runtime mismatches before they reach production, and local model development for teams that need quantized models running on-premises for privacy or compliance reasons.

Where quantization is headed, and the one habit worth building now

FP8 support and better 4-bit PTQ methods are widening what’s deployable on a given GPU budget, and that trend looks likely to continue as more runtimes add native FP8 kernels. The mistake we see most often is optimizing for a single metric, usually perplexity, because it’s easy to measure, while the production task actually depends on instruction-following or domain accuracy that perplexity doesn’t capture well. Our advice: run a small, representative pilot on your actual traffic before committing to a full migration. A week spent validating beats a quarter spent unwinding a bad default.

— Chad

Get your quantization project production-ready

Quantization research moves fast, but shipping a quantized model that holds up under real traffic is a different problem than reproducing a paper’s benchmark. We combine AI engineering expertise with experienced human review to get models from prototype to production without the guesswork of mismatched kernels or silent accuracy regressions.

Bowtie

If you’re weighing whether to quantize in-house or need a second set of eyes on an existing deployment, our AI Engineering Assistance, AI Code Reviews & Optimization, and Local AI Model Development services cover the full path from audit to deployment. For teams further along, a Senior Developer Review starts at $449 and gives you a concrete read on where your current implementation stands before you invest further. Check our pricing and service bundles to find the right starting point for your team.

FAQ

When should you quantize a model?

Quantize when memory, latency, or serving cost is limiting deployment, typically when a model is too large for your target hardware or too slow to meet latency requirements. It’s less useful early in development when you still need maximum flexibility to retrain or adjust the architecture.

Can you explain quantization in a simple way?

Quantization is the process of representing a model’s numbers with fewer bits, similar to rounding a long decimal to save space. This makes the model smaller and often faster to run, usually with a small, manageable loss in accuracy.

How are LLM models quantized?

Large language models are typically quantized using post-training methods like GPTQ or AWQ, which compress weights to 4 bits using calibration data rather than retraining the full model. Some workflows also apply SmoothQuant to enable full weight-and-activation INT8 quantization by smoothing activation outliers first.

What is the primary goal of model quantization?

The primary goal is reducing memory footprint and compute cost by lowering numerical precision, which makes models cheaper and faster to deploy. The best quantization approach balances these efficiency gains against an acceptable level of accuracy loss for the task at hand.

Sources