English | 中文
A dependency-light, transformers-free toolkit for post-training, weight-only quantization of HuggingFace checkpoints.
quantkit reads .safetensors directly — no model definition, no forward pass, no calibration data. Quantize a checkpoint on CPU or GPU, split the work across machines at the file level, and get a standard HF checkpoint back.
Scope: weight-only, data-free (absmax / round-to-nearest) quantization. Not a calibration- or training-based tool — for GPTQ / AWQ / static-activation schemes use
llm-compressor.
| scheme | element | scale | granularity | tensor names |
|---|---|---|---|---|
FP8_BLOCK |
E4M3 | fp32 | 128×128 block | weight, weight_scale |
FP8_CHANNEL |
E4M3 | fp32 | per output channel | weight, weight_scale |
MXFP8 |
E4M3 | E8M0 | 1D group of 32 | weight, weight_scale |
MXFP4 |
E2M1 | E8M0 | 1D group of 32 | weight (packed), weight_scale |
All schemes derive the scale from the weight tensor itself, so no calibration dataset is required. MXFP8/MXFP4 are OCP microscaling formats; the E8M0 scale is a power of two.
pip install -e . # torch + safetensors + numpy + pyyaml
pip install -e ".[dev]" # + pytestRequires Python ≥ 3.10 and PyTorch ≥ 2.1.
quantkit quantize \
--model /path/to/hf_bf16 \
--output /path/to/hf_mxfp4 \
--scheme MXFP4 \
--include '.*\.(gate|up|down)_proj\.weight$' \
--device cudaOr from a YAML/JSON config (--config examples/quantize_mxfp4.yaml); CLI flags override config values.
Python API:
from quantkit import quantize_model
quantize_model(
model_dir="hf_bf16",
out_dir="hf_mxfp4",
scheme="MXFP4",
include=r".*\.(gate|up|down)_proj\.weight$",
device="cuda",
)--include REGEX/--exclude REGEXmatch on tensor names, e.g.--include '.*\.down_proj\.weight$'.- Only
.weighttensors with ≥ 2 dims are candidates; norms, biases and other 1-D tensors are left untouched. - Quantizing fewer layers (e.g. only MoE experts) also cuts I/O linearly — useful on very large checkpoints.
An HF checkpoint: *.safetensors (quantized weights + scales),
model.safetensors.index.json, and config.json with a quantization_config.
Non-weight sidecars (tokenizer, …) are copied over unchanged.
Two metadata schemas:
--emit hf(default) — the flatquantization_configwithquant_method(fp8/mxfp4), read by HF Transformers and vLLM.--emit compressed-tensors— theconfig_groupsschema read by vLLM, llm-compressor and HF.
The packed-weight tensor name follows the schema — see Compatibility.
Parallelism is file-level: the input is split into num_shards disjoint slices of safetensors files. Worker i quantizes slice i and writes the same-named output shard, so the union of all workers reconstructs the full checkpoint. Workers only need a barrier and a weight-map gather (gloo) — no tensor communication.
The sharding knobs (--num-shards / --shard-index) fall back to env vars, and further to the standard launcher variables:
| value | CLI | env (preferred) | env (fallback) |
|---|---|---|---|
| number of workers | --num-shards N |
NUM_SHARDS |
WORLD_SIZE |
| this worker's index | --shard-index i |
SHARD_INDEX |
RANK |
In addition, the rendezvous variables MASTER_ADDR / MASTER_PORT enable the process-group merge (see below).
torchrun --standalone --nproc_per_node=8 \
-m quantkit quantize --model M --output O --scheme MXFP4--standalone sets up a local rendezvous (MASTER_ADDR / MASTER_PORT / RANK / WORLD_SIZE) for you.
Run the same command on every machine, changing only --node_rank:
# machine 0 (rendezvous host)
torchrun --nnodes=4 --nproc_per_node=8 --node_rank=0 \
--master_addr=10.0.0.1 --master_port=29500 \
-m quantkit quantize --model /shared/hf_bf16 --output /shared/hf_mxfp4 --scheme MXFP4
# machine 1
torchrun --nnodes=4 --nproc_per_node=8 --node_rank=1 \
--master_addr=10.0.0.1 --master_port=29500 \
-m quantkit quantize --model /shared/hf_bf16 --output /shared/hf_mxfp4 --scheme MXFP4
# machine 2, 3: same, with --node_rank=2, 3--master_addris machine 0;WORLD_SIZE = nnodes × nproc_per_node.- One worker per machine is fine too:
--nnodes=4 --nproc_per_node=1.
On machine i of N:
RANK=$i WORLD_SIZE=$N MASTER_ADDR=10.0.0.1 MASTER_PORT=29500 \
quantkit quantize --model /shared/hf_bf16 --output /shared/hf_mxfp4 --scheme MXFP4Set RANK / WORLD_SIZE (not just NUM_SHARDS / SHARD_INDEX): they drive both the sharding and the process-group rendezvous.
The index (model.safetensors.index.json) is written by worker 0 and must cover every shard. Whether it is complete depends on whether the workers form a process group (torch.distributed):
- With a process group (e.g.
torchrun, or whenMASTER_ADDRis set): quantkit initializes a gloo group, all-gathers each worker's weight map, and worker 0 writes a complete index. This is the recommended path. - Without one (independent processes): workers cannot exchange weight maps, so worker 0's index covers only its own shards and a warning is emitted.
A group only forms if the workers share a rendezvous; the launchers above set MASTER_ADDR / MASTER_PORT / RANK / WORLD_SIZE for you.
- Shared filesystem: every worker must see the same
--modeland--outputpaths (a network/shared filesystem), since they write different shards of one checkpoint. - Throughput ceiling is aggregate storage bandwidth — usually the real limit; adding workers past that point does not help.
- Sharding is by file, which is balanced when the checkpoint's shards are roughly equal-sized (the usual case). Very uneven shards can cause stragglers.
- gloo is CPU-only, so no GPUs are required for the multi-worker path.
| loader | schema | MXFP4 packed tensor |
|---|---|---|
| HF Transformers / vLLM (native) | quant_method: fp8 / mxfp4 |
weight |
| compressed-tensors (vLLM, llm-compressor, HF) | quant_method: compressed-tensors |
weight_packed |
Validated against compressed-tensors: the emitted config parses and both
MXFP4PackedCompressor.decompress and MXFP8QuantizationCompressor.decompress
recover the original weights (MXFP8 ≈ 3.6%, MXFP4 ≈ 12% max abs error / absmax;
MXFP4 is coarse by design).
quantkit overlaps a slice of llm-compressor:
- llm-compressor is the broader tool, but pulls in
transformers/accelerate/datasets, and its model-free path (model_free_ptq) is single-node. quantkitis transformers-free, dependency-light, file-level multi-node, and limited to data-free weight-only quantization — the smallest thing that produces a loadable MX / FP8 checkpoint.
pip install -e ".[dev]"
pytestMIT