Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

quantkit

English | 中文

A dependency-light, transformers-free toolkit for post-training, weight-only quantization of HuggingFace checkpoints.

quantkit reads .safetensors directly — no model definition, no forward pass, no calibration data. Quantize a checkpoint on CPU or GPU, split the work across machines at the file level, and get a standard HF checkpoint back.

Scope: weight-only, data-free (absmax / round-to-nearest) quantization. Not a calibration- or training-based tool — for GPTQ / AWQ / static-activation schemes use llm-compressor.

Formats

scheme element scale granularity tensor names
FP8_BLOCK E4M3 fp32 128×128 block weight, weight_scale
FP8_CHANNEL E4M3 fp32 per output channel weight, weight_scale
MXFP8 E4M3 E8M0 1D group of 32 weight, weight_scale
MXFP4 E2M1 E8M0 1D group of 32 weight (packed), weight_scale

All schemes derive the scale from the weight tensor itself, so no calibration dataset is required. MXFP8/MXFP4 are OCP microscaling formats; the E8M0 scale is a power of two.

Install

pip install -e .            # torch + safetensors + numpy + pyyaml
pip install -e ".[dev]"     # + pytest

Requires Python ≥ 3.10 and PyTorch ≥ 2.1.

Usage

quantkit quantize \
    --model  /path/to/hf_bf16 \
    --output /path/to/hf_mxfp4 \
    --scheme MXFP4 \
    --include '.*\.(gate|up|down)_proj\.weight$' \
    --device cuda

Or from a YAML/JSON config (--config examples/quantize_mxfp4.yaml); CLI flags override config values.

Python API:

from quantkit import quantize_model

quantize_model(
    model_dir="hf_bf16",
    out_dir="hf_mxfp4",
    scheme="MXFP4",
    include=r".*\.(gate|up|down)_proj\.weight$",
    device="cuda",
)

Selecting which tensors to quantize

  • --include REGEX / --exclude REGEX match on tensor names, e.g. --include '.*\.down_proj\.weight$'.
  • Only .weight tensors with ≥ 2 dims are candidates; norms, biases and other 1-D tensors are left untouched.
  • Quantizing fewer layers (e.g. only MoE experts) also cuts I/O linearly — useful on very large checkpoints.

Output

An HF checkpoint: *.safetensors (quantized weights + scales), model.safetensors.index.json, and config.json with a quantization_config. Non-weight sidecars (tokenizer, …) are copied over unchanged.

Two metadata schemas:

  • --emit hf (default) — the flat quantization_config with quant_method (fp8 / mxfp4), read by HF Transformers and vLLM.
  • --emit compressed-tensors — the config_groups schema read by vLLM, llm-compressor and HF.

The packed-weight tensor name follows the schema — see Compatibility.

Multi-worker / multi-node

Parallelism is file-level: the input is split into num_shards disjoint slices of safetensors files. Worker i quantizes slice i and writes the same-named output shard, so the union of all workers reconstructs the full checkpoint. Workers only need a barrier and a weight-map gather (gloo) — no tensor communication.

The sharding knobs (--num-shards / --shard-index) fall back to env vars, and further to the standard launcher variables:

value CLI env (preferred) env (fallback)
number of workers --num-shards N NUM_SHARDS WORLD_SIZE
this worker's index --shard-index i SHARD_INDEX RANK

In addition, the rendezvous variables MASTER_ADDR / MASTER_PORT enable the process-group merge (see below).

Single node, multiple workers

torchrun --standalone --nproc_per_node=8 \
    -m quantkit quantize --model M --output O --scheme MXFP4

--standalone sets up a local rendezvous (MASTER_ADDR / MASTER_PORT / RANK / WORLD_SIZE) for you.

Multiple machines (torchrun)

Run the same command on every machine, changing only --node_rank:

# machine 0 (rendezvous host)
torchrun --nnodes=4 --nproc_per_node=8 --node_rank=0 \
    --master_addr=10.0.0.1 --master_port=29500 \
    -m quantkit quantize --model /shared/hf_bf16 --output /shared/hf_mxfp4 --scheme MXFP4

# machine 1
torchrun --nnodes=4 --nproc_per_node=8 --node_rank=1 \
    --master_addr=10.0.0.1 --master_port=29500 \
    -m quantkit quantize --model /shared/hf_bf16 --output /shared/hf_mxfp4 --scheme MXFP4

# machine 2, 3: same, with --node_rank=2, 3
  • --master_addr is machine 0; WORLD_SIZE = nnodes × nproc_per_node.
  • One worker per machine is fine too: --nnodes=4 --nproc_per_node=1.

Manual, via environment (no torchrun)

On machine i of N:

RANK=$i WORLD_SIZE=$N MASTER_ADDR=10.0.0.1 MASTER_PORT=29500 \
    quantkit quantize --model /shared/hf_bf16 --output /shared/hf_mxfp4 --scheme MXFP4

Set RANK / WORLD_SIZE (not just NUM_SHARDS / SHARD_INDEX): they drive both the sharding and the process-group rendezvous.

How the index is merged

The index (model.safetensors.index.json) is written by worker 0 and must cover every shard. Whether it is complete depends on whether the workers form a process group (torch.distributed):

  • With a process group (e.g. torchrun, or when MASTER_ADDR is set): quantkit initializes a gloo group, all-gathers each worker's weight map, and worker 0 writes a complete index. This is the recommended path.
  • Without one (independent processes): workers cannot exchange weight maps, so worker 0's index covers only its own shards and a warning is emitted.

A group only forms if the workers share a rendezvous; the launchers above set MASTER_ADDR / MASTER_PORT / RANK / WORLD_SIZE for you.

Requirements & caveats

  • Shared filesystem: every worker must see the same --model and --output paths (a network/shared filesystem), since they write different shards of one checkpoint.
  • Throughput ceiling is aggregate storage bandwidth — usually the real limit; adding workers past that point does not help.
  • Sharding is by file, which is balanced when the checkpoint's shards are roughly equal-sized (the usual case). Very uneven shards can cause stragglers.
  • gloo is CPU-only, so no GPUs are required for the multi-worker path.

Compatibility

loader schema MXFP4 packed tensor
HF Transformers / vLLM (native) quant_method: fp8 / mxfp4 weight
compressed-tensors (vLLM, llm-compressor, HF) quant_method: compressed-tensors weight_packed

Validated against compressed-tensors: the emitted config parses and both MXFP4PackedCompressor.decompress and MXFP8QuantizationCompressor.decompress recover the original weights (MXFP8 ≈ 3.6%, MXFP4 ≈ 12% max abs error / absmax; MXFP4 is coarse by design).

Relation to other tools

quantkit overlaps a slice of llm-compressor:

  • llm-compressor is the broader tool, but pulls in transformers / accelerate / datasets, and its model-free path (model_free_ptq) is single-node.
  • quantkit is transformers-free, dependency-light, file-level multi-node, and limited to data-free weight-only quantization — the smallest thing that produces a loadable MX / FP8 checkpoint.

Testing

pip install -e ".[dev]"
pytest

License

MIT

About

A dependency-light, transformers-free toolkit for data-free weight-only quantization (FP8 blockwise, MXFP8, MXFP4), with file-level multi-node support and HF / compressed-tensors output.

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages