wejoncy / QLLM

A general 2-8 bits quantization toolbox with GPTQ/AWQ/HQQ, and export to onnx/onnx-runtime easily.
Apache License 2.0
149 stars 15 forks source link

support `MARLIN` pack_mode #106

Closed wejoncy closed 8 months ago