diff --git a/docs/source/en/quantization/auto_round.md b/docs/source/en/quantization/auto_round.md index fb19214b19f8..314b4f5e84c2 100644 --- a/docs/source/en/quantization/auto_round.md +++ b/docs/source/en/quantization/auto_round.md @@ -15,7 +15,7 @@ rendered properly in your Markdown viewer. It leverages sign gradient descent to fine-tune both rounding values and min-max clipping thresholds in just 200 steps. Designed for broad compatibility, it seamlessly supports a wide range of LLMs and is actively expanding to cover more VLMs as well. It also supports quantization and inference across multiple hardware platforms, including CPU, XPU, and CUDA. -AutoRound also offers a variety of useful features, including mixed-bit tuning and inference, lm-head quantization, support for exporting to formats like GPTQ/AWQ/GGUF, and flexible tuning recipes. +AutoRound also offers a variety of useful features, including automatic mixed-bit tuning and inference, support for MXFP4 and NVFP4 data types, model-free quantization, export to formats such as GPTQ, AWQ, GGUF, and LLM-Compressor, and flexible tuning recipes. For a comprehensive overview and the latest updates, check out the AutoRound [README](https://github.com/intel/auto-round). AutoRound was originally developed as part of the [Intel Neural Compressor](https://github.com/intel/neural-compressor), serving as a general-purpose model compression library for deep learning. @@ -30,13 +30,13 @@ pip install auto-round ## Supported Quantization Configurations -AutoRound supports several quantization configurations: +AutoRound supports the following quantization configurations: -- **Int8 Weight Only** -- **Int4 Weight Only** -- **Int3 Weight Only** -- **Int2 Weight Only** -- **Mixed bits Weight only** +1. **INT2–INT8 Weight-Only** +2. **Mixed-Bit Weight-Only** +3. **GGUF Q\*_K** +4. **MXFP** (limited support) +5. **NVFP4** (limited support) ## Hardware Compatibility @@ -53,16 +53,49 @@ Currently, only offline mode is supported to generate quantized models. ```bash auto-round \ - --model facebook/opt-125m \ - --bits 4 \ + --model Qwen/Qwen3-0.6B \ + --scheme "W4A16" \ --group_size 128 \ --output_dir ./tmp_autoround ``` -AutoRound also offer another two recipes, `auto-round-best` and `auto-round-light`, designed for optimal accuracy and improved speed, respectively. -For 2 bits, we recommend using `auto-round-best` or `auto-round`. +AutoRound also offer another four recipes, `auto-round-best`, `auto-round-light`,`auto-round-opt-rtn`,`auto-round-rtn`, designed for optimal accuracy and improved speed, respectively. +For 2 bits, we recommend using `auto-round-best` with `--enable_alg_ext`. + + + + + + +### AutoScheme Usage + +AutoScheme is a feature that automatically selects the best quantization scheme from the available options for each layer to be quantized in minutes, subject to a target average bit width. + + +```bash +auto-round \ + --model Qwen/Qwen3-0.6B \ + --options "W4A16,W2A16G64" \ + --target_bits 3.5 \ + --output_dir ./tmp_autoround +``` + + +### Algorithm Combinations + +AutoRound supports combining multiple algorithms, such as AutoRound, AutoRound + AWQ, and AutoRound + AWQ + Hadamard (very limited support). We are currently expanding support for more algorithms. This enables further optimization of quantization results, especially for scenarios involving quantized activations. + +```bash +auto-round \ + --model Qwen/Qwen3-0.6B \ + --algs "autoround,awq" \ + --output_dir ./tmp_autoround +``` + + + ### AutoRound API Usage @@ -70,28 +103,21 @@ For 2 bits, we recommend using `auto-round-best` or `auto-round`. This setting offers a better trade-off between accuracy and tuning cost, and is recommended in all scenarios. ```python -from transformers import AutoModelForCausalLM, AutoTokenizer from auto_round import AutoRound -model_name = "facebook/opt-125m" -model = AutoModelForCausalLM.from_pretrained(model_name, dtype="auto") -tokenizer = AutoTokenizer.from_pretrained(model_name) -bits, group_size, sym = 4, 128, True +model_name = "Qwen/Qwen3-0.6B" # mixed bits config # layer_config = {"model.decoder.layers.6.self_attn.out_proj": {"bits": 2, "group_size": 32}} -autoround = AutoRound( - model, - tokenizer, - bits=bits, - group_size=group_size, - sym=sym, +ar = AutoRound( + model_name, + scheme="W4A16", # enable_torch_compile=True, # layer_config=layer_config, ) output_dir = "./tmp_autoround" -# format= 'auto_round'(default), 'auto_gptq', 'auto_awq' -autoround.quantize_and_save(output_dir, format='auto_round') +# format= 'auto_round'(default), 'llm_compressor', "gguf:q4_k_m", 'auto_gptq', 'auto_awq' +ar.quantize_and_save(output_dir, format='auto_round') ``` @@ -103,26 +129,18 @@ autoround.quantize_and_save(output_dir, format='auto_round') This setting provides the best accuracy in most scenarios but is 4–5× slower than the standard AutoRound recipe. It is especially recommended for 2-bit quantization and is a good choice if sufficient resources are available. ```python -from transformers import AutoModelForCausalLM, AutoTokenizer from auto_round import AutoRound -model_name = "facebook/opt-125m" -model = AutoModelForCausalLM.from_pretrained(model_name, dtype="auto") -tokenizer = AutoTokenizer.from_pretrained(model_name) -bits, group_size, sym = 4, 128, True -autoround = AutoRound( - model, - tokenizer, - bits=bits, - group_size=group_size, - sym=sym, +model_name = "Qwen/Qwen3-0.6B" +ar = AutoRound( + model_name, + scheme="W4A16", nsamples=512, iters=1000, - low_gpu_mem_usage=True ) output_dir = "./tmp_autoround" -autoround.quantize_and_save(output_dir, format='auto_round') +ar.quantize_and_save(output_dir, format='auto_round') ``` @@ -137,22 +155,15 @@ This setting offers the best speed (2 - 3X faster than AutoRound), but it may ca from transformers import AutoModelForCausalLM, AutoTokenizer from auto_round import AutoRound -model_name = "facebook/opt-125m" -model = AutoModelForCausalLM.from_pretrained(model_name, dtype="auto") -tokenizer = AutoTokenizer.from_pretrained(model_name) -bits, group_size, sym = 4, 128, True -autoround = AutoRound( - model, - tokenizer, - bits=bits, - group_size=group_size, - sym=sym, +model_name = "Qwen/Qwen3-0.6B" +ar = AutoRound( + model_name, iters=50, lr=5e-3, ) output_dir = "./tmp_autoround" -autoround.quantize_and_save(output_dir, format='auto_round') +ar.quantize_and_save(output_dir, format='auto_round') ``` @@ -176,7 +187,7 @@ AutoRound automatically selects the best available backend based on the installe ### CPU -Supports 2, 4, and 8 bits. We recommend using the AutoRound Kernel (ARK) backend for inference. PyTorch 2.8.0 or later is required with ARK. +Supports 2, 4 and 8 bits. We recommend using the AutoRound Kernel (ARK) backend for inference. PyTorch 2.8.0 or later is required with ARK. ```python from transformers import AutoModelForCausalLM, AutoTokenizer @@ -195,7 +206,7 @@ print(tokenizer.decode(model.generate(**inputs, max_new_tokens=50, do_sample=Fal ### XPU -Supports 4 and 8 bits. We recommend using the AutoRound Kernel (ARK) backend for inference. PyTorch 2.8.0 or later is required with ARK. +Supports 2, 4 and 8 bits. We recommend using the AutoRound Kernel (ARK) backend for inference. PyTorch 2.8.0 or later is required with ARK. ```python from transformers import AutoModelForCausalLM, AutoTokenizer @@ -214,7 +225,7 @@ print(tokenizer.decode(model.generate(**inputs, max_new_tokens=50, do_sample=Fal ### CUDA -Supports 2, 3, 4, and 8 bits. We recommend using GPTQModel for 4 and 8 bits inference. +Supports 2-8 bits. We recommend using GPTQModel for 4 and 8 bits inference. ```python from transformers import AutoModelForCausalLM, AutoTokenizer diff --git a/src/transformers/utils/quantization_config.py b/src/transformers/utils/quantization_config.py index fd3617b4b6d8..68ea144c3c7c 100644 --- a/src/transformers/utils/quantization_config.py +++ b/src/transformers/utils/quantization_config.py @@ -211,7 +211,7 @@ class AutoRoundConfig(QuantizationConfigMixin): Args: bits (`int`, *optional*, defaults to 4): - The number of bits to quantize to, supported numbers are (2, 3, 4, 8). + The number of bits to quantize to, supported numbers are (2, 3, 4, 5, 6, 7, 8). group_size (`int`, *optional*, defaults to 128): Group-size value sym (`bool`, *optional*, defaults to `True`): Symmetric quantization or not backend (`str`, *optional*, defaults to `"auto"`): The inference backend. By default, AutoRound selects a compatible backend based on the device, quantization settings, and installed libraries. See [Specify inference backend](https://github.com/intel/auto-round/blob/main/docs/step_by_step.md#specify-inference-backend) for all backend options. @@ -238,7 +238,7 @@ def __init__( def post_init(self): r"""Safety checker that arguments are correct.""" - if self.bits not in [2, 3, 4, 8]: + if self.bits not in [2, 3, 4, 5, 6, 7, 8]: raise ValueError(f"Only support quantization to [2,3,4,8] bits but found {self.bits}") if self.group_size != -1 and self.group_size <= 0: raise ValueError("group_size must be greater than 0 or equal to -1")