Quantizing the Key-Value (KV) cache from standard FP16 down to FP8 (E4M3 or E5M2) or INT4 in SGLang slashes context window VRAM consumption by 50% to 75% with under 0.5% perplexity degradation. This permits local inference servers to double concurrent batch sizes and run 32k to 128k context windows on standard 24GB RTX 3090/4090 GPUs.
What Is KV Cache Quantization in SGLang?
KV Cache Quantization in SGLang compresses the stored attention keys and values generated during multi-turn LLM inference into lower-precision 8-bit or 4-bit floating point formats, drastically reducing GPU memory consumption during long-context generation.
When hosting large language models (LLMs) locally, VRAM is consumed by two distinct elements: static model weights and the dynamic Key-Value (KV) cache. While model weights remain fixed once loaded, the KV cache grows linearly with every token processed and scales multiplicatively with concurrent batch requests. In a 32,000-token context window, unquantized FP16 KV cache can easily exceed 16GB of VRAM per request, causing out-of-memory (OOM) crashes on workstation GPUs.
| Precision Format | VRAM per 8k Context (Llama-3-70B) | Perplexity Impact | Throughput (Tokens/sec) |
|---|---|---|---|
| Native FP16 | ~5.2 GB per sequence | Baseline (0.00%) | 100% (Baseline) |
| FP8 (E4M3 Format) | ~2.6 GB per sequence (50% reduction) | < 0.15% (Indistinguishable) | 118% (Reduced memory bandwidth) |
| FP8 (E5M2 Format) | ~2.6 GB per sequence (50% reduction) | ~0.32% (Slight drop) | 115% |
| INT4 (AWQ Scale) | ~1.3 GB per sequence (75% reduction) | ~1.10% (Measurable degradation) | 95% (Dequantization overhead) |
How to Enable FP8 KV Cache in SGLang CLI
To enable FP8 KV Cache in SGLang, launch the server using the --kv-cache-dtype fp8_e4m3 flag during server instantiation.
SGLang’s implementation integrates seamlessly with its native RadixAttention engine. Unlike standard inference servers that dump cached tokens between API calls, SGLang retains prefix attention trees in quantized memory, allowing subsequent multi-turn agentic requests to skip re-computation.
Here is the production command to launch a quantized inference server on an Ada Lovelace or Hopper GPU:
python3 -m sglang.launch_server \
--model-path meta-llama/Meta-Llama-3.1-70B-Instruct \
--kv-cache-dtype fp8_e4m3 \
--mem-fraction-static 0.85 \
--context-length 32768 \
--port 30000
If you run on older Ampere hardware (such as RTX 3090 or A100 GPUs lacking native FP8 tensor cores), SGLang automatically performs runtime casting, delivering the exact same 50% VRAM memory savings with negligible compute overhead.
For production local AI deployments, FP8 (E4M3) KV caching should be enabled by default. It provides a free 50% memory expansion with zero perceptible degradation in coding or reasoning tasks. Avoid INT4 KV caching unless memory constraints make it physically impossible to load your target sequence length, as INT4 dequantization introduces compute latency and noticeable degradation on math benchmarks.
People Also Ask
Does FP8 KV Cache work on consumer NVIDIA RTX 3090 GPUs?
Yes. While Ampere RTX 3090s lack native hardware FP8 tensor cores, SGLang stores the cache in 8-bit memory to cut VRAM use and casts to FP16 in CUDA cores during attention calculation with minimal speed loss.
What is the difference between FP8 E4M3 and E5M2?
E4M3 uses 4 bits for exponent and 3 for mantissa, prioritizing numerical precision over dynamic range, which is optimal for KV caches. E5M2 prioritizes dynamic range and is typically used for neural network gradient backpropagation.
Does SGLang KV quantization work with RadixAttention?
Yes. SGLang’s RadixAttention prefix tree natively supports quantized cache nodes, allowing multi-turn chats and shared system prompts to remain cached in FP8 format indefinitely.

