755GB to 395GB: How Bitdeer AI Halves GLM-5.3's GPU Footprint
Author: Dang Hoang Duy
TL;DR
GLM-5.3's 753 billion parameters ship in FP8 as 755 GB of weights, too big for the 640 GB an 8xH100 node holds, especially once the KV cache needs room too. We quantized it in-house to W4AFP8 (4-bit experts, FP8 activations): the checkpoint is about 395 GB, it serves on one node instead of two, and it matches the unquantized model on every benchmark where we have a figure to compare against.
We built two pipelines AWQ and expert-parallel GPTQ and recommend the latter: it cuts per-GPU calibration memory from ~76 GB to under 10 GB, produces its checkpoint in ~37% less wall-clock time, and reaches the same answers using 5.5–18.2% fewer output tokens.
In this post:
• What we quantized
• The hard part was memory
• Does it still answer correctly?
• Same scores, fewer tokens
• Deploying it
GLM-5.3 is big. Its 753 billion parameters ship as 755 GB of FP8 weights spread across 141 files, and an 8xH100 node has 640 GB of GPU memory. The weights don't fit, and that's before you leave any room for the KV cache, which holds the state of every in-flight request.
So we quantized it ourselves, down to 4-bit experts with FP8 activations. The result is a checkpoint of about 395 GB that serves on a single node. On every benchmark where we have a reference to compare against, it matches the unquantized model.
All numbers in this post were measured on nodes of eight NVIDIA H100-80GB GPUs unless we say otherwise. To serve GLM-5.3 at its full 1M-token context, we recommend 8xH200 instead.

Figure 1. Weight footprint against what one 8-GPU node holds. Anything left of a vertical line still has to share that node with the KV cache.
What we quantized
W4AFP8 stands for 4-bit integer weights with 8-bit floating-point activations, but we don't apply it everywhere. Each of GLM-5.3's 75 MoE layers has 256 routed experts, and together they account for nearly all of the model's parameters. That's where 4 bits pays off, so that's where we use INT4. Everything else is either too small to matter or too sensitive to touch.
We didn't pick this split ourselves. It's exactly what SGLang's w4afp8 path expects, right down to details like the DSA (sparse-attention) indexer, which the engine refuses to load in BF16. Get the layout wrong and you're looking at another conversion pass over 400 GB before the model will serve at all. So both of our pipelines write SGLang's layout directly, MTP draft layer included.
The hard part was memory
We built two quantization paths. AWQ protects the salient weight channels, the ones that see the largest activations, by scaling them before quantization. GPTQ quantizes one column at a time and uses second-order information to compensate for the error in the weights it hasn't quantized yet. Both used the same calibration data: 256 samples of 2,048 tokens from UltraChat, with all 256 experts in every layer calibrated. Both also calibrate layer by layer, one decoder layer at a time. They have to, because the BF16 source weights come to 1.51 TB and a node has 640 GB.
GPTQ is where memory really bites. For every linear module it quantizes, it keeps a Hessian built from that module's calibration inputs. One GLM-5.3 MoE layer has 768 of them (256 experts, three projections each), and in the standard data-parallel implementation every GPU accumulates all of them before they're combined. That adds up to roughly 76 GiB of working memory on an 80 GB card, before a single weight or activation is loaded. As it stands, GPTQ simply can't run on this model.
The fix is expert parallelism instead of data parallelism. IST-DASLab's MoE-Quant used the same approach on DeepSeek-V3, for the same reason. Each GPU owns a subset of the experts and holds only their Hessians. The calibration activations are all-gathered first, so every GPU runs its experts over every token and each Hessian still covers the full calibration set. Peak memory came to 26.3 GB across 4 GPUs and 13.6 GB across 8. The two runs produced identical greedy output, so how you split the work doesn't change the answer. And on the same hardware, this path produced its checkpoint in about 37% less wall-clock time than AWQ.

Figure 2. Hessian workspace per GPU for one MoE layer, on H100-80GB cards.
Does it still answer correctly?
To find out, we compared three W4AFP8 checkpoints of the same model: our expert-parallel GPTQ build, our AWQ build, and PhalaCloud's independent W4AFP8 release. Same harness, same prompts; only the weights differ. We reproduced two benchmarks following Artificial Analysis' published methodology: GPQA Diamond (198 questions, 5 repeats each) and AA-LCR v1.1, a long-context reasoning benchmark (100 questions, 3 repeats, around 97k tokens of input each).

Figure 3. Whiskers are one standard error. The vertical line is Artificial Analysis' published figure for unquantized GLM-5.3.
Public-methodology reproductions on our own endpoint, not official Artificial Analysis runs. PhalaCloud's own card reports 91.9 and 73.0 under a different protocol.
The three can't be told apart. On GPQA Diamond, the best and worst checkpoints are 0.41 percentage points apart, against error bars of about +/-0.9. On AA-LCR all three answered the same 300 items, so we could compare each pair directly with a paired test, and none of the pairs differs (p = 0.55 to 0.87).
A six-task capability suite, scored with lm-evaluation-harness on the full test sets, tells the same story.
No two checkpoints are 2 points apart on any task. The largest gap, 1.85 points on IFEval, is within measurement noise at this sample size.
We also ran Terminal-Bench, which tests whether a model can complete real tasks at a command line. This time we ran unquantized GLM-5.3 alongside as a reference, so it's the one result in this post that compares a quantized checkpoint against the original model rather than against another quantization.
Run separately from the benchmarks above. Scores are the mean over 3 seeds, and +/- here is the standard deviation across seeds, not a standard error as in the other tables. Terminal-Bench is large, so we only ran 3 seeds, and with this much variance a 3-seed standard deviation is a rough guide at best. Our AWQ checkpoint matched the unquantized model's mean; PhalaCloud's 3.0-point gap is smaller than its own spread.
Same scores, fewer tokens
So the scores don't separate the three checkpoints. What does separate them is how much they write to get there. Our EP-GPTQ build reaches the same answers with shorter reasoning. On GPQA Diamond it spends 5.5% fewer output tokens per answer than our AWQ build and 14.6% fewer than PhalaCloud's. On AA-LCR the gap is 18.2% against our AWQ build and 8.0% against PhalaCloud's.
On a served endpoint, output tokens are the bill, so that's a direct saving. How big it is depends on the workload, though: on a short-answer suite it didn't show up at all.

Figure 4. Run totals at the same accuracy. The percentages above are per answer.
Deploying it
Both checkpoints are on Hugging Face under Bitdeer AI, and both run the same way, so pick either.
You'll need SGLang 0.5.17 or newer, one 8-GPU Hopper node, and about 400 GB of free disk.
Send requests at temperature 1.0 and top_p 0.95, the settings the model card recommends. Both checkpoints also carry GLM-5.3's MTP draft layer, assembled at save time from the same source release, so EAGLE speculative decoding works with no separate draft model. Add --speculative-algorithm EAGLE to turn it on. Each model card has the full flag reference, including the context-length and memory-fraction settings you'll want to tune per node.
We validated this configuration on SGLang 0.5.17 at tensor parallelism 8 with an FP8 KV cache. The quantized checkpoints inherit GLM-5.3's licence.
The short version
- 753 billion parameters, 755 GB in FP8, down to about 395 GB at W4AFP8: one 8-GPU node instead of two.
- Quality matches Artificial Analysis' published GLM-5.3 figures on both benchmarks, and is level with an independent W4AFP8 release.
- On Terminal-Bench, the one benchmark we ran against unquantized GLM-5.3, our AWQ checkpoint matched its mean score of 85.4% over three seeds.
- Expert-parallel GPTQ cuts the calibration memory a 256-expert layer needs from roughly 76 GiB per GPU to under 10 GiB, without changing what the model answers.
Scaling down massive MoE models shouldn't mean sacrificing accuracy or hardware efficiency. With Expert-Parallel GPTQ, we’ve demonstrated that GLM-5.3 can be crammed into a single 8xH100 node while maintaining uncompromised evaluation scores and generating noticeably fewer reasoning tokens.
Both checkpoints are now live on Hugging Face. We recommend expert-parallel GPTQ as it costs less to run. Grab the commands above, launch your single-node instance, and let us know how it performs in your production environment!