Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
113 changes: 62 additions & 51 deletions docs/source/en/quantization/auto_round.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,7 @@ rendered properly in your Markdown viewer.
It leverages sign gradient descent to fine-tune both rounding values and min-max clipping thresholds in just 200 steps. Designed for broad compatibility, it seamlessly supports a wide range of LLMs and is actively expanding to cover more VLMs as well.
It also supports quantization and inference across multiple hardware platforms, including CPU, XPU, and CUDA.

AutoRound also offers a variety of useful features, including mixed-bit tuning and inference, lm-head quantization, support for exporting to formats like GPTQ/AWQ/GGUF, and flexible tuning recipes.
AutoRound also offers a variety of useful features, including automatic mixed-bit tuning and inference, support for MXFP4 and NVFP4 data types, model-free quantization, export to formats such as GPTQ, AWQ, GGUF, and LLM-Compressor, and flexible tuning recipes.
For a comprehensive overview and the latest updates, check out the AutoRound [README](https://github.com/intel/auto-round).

AutoRound was originally developed as part of the [Intel Neural Compressor](https://github.com/intel/neural-compressor), serving as a general-purpose model compression library for deep learning.
Expand All @@ -30,13 +30,13 @@ pip install auto-round

## Supported Quantization Configurations

AutoRound supports several quantization configurations:
AutoRound supports the following quantization configurations:

- **Int8 Weight Only**
- **Int4 Weight Only**
- **Int3 Weight Only**
- **Int2 Weight Only**
- **Mixed bits Weight only**
1. **INT2–INT8 Weight-Only**
2. **Mixed-Bit Weight-Only**
3. **GGUF Q\*_K**
4. **MXFP** (limited support)
5. **NVFP4** (limited support)

## Hardware Compatibility

Expand All @@ -53,45 +53,71 @@ Currently, only offline mode is supported to generate quantized models.

```bash
auto-round \
--model facebook/opt-125m \
--bits 4 \
--model Qwen/Qwen3-0.6B \
--scheme "W4A16" \
--group_size 128 \
--output_dir ./tmp_autoround
```

AutoRound also offer another two recipes, `auto-round-best` and `auto-round-light`, designed for optimal accuracy and improved speed, respectively.
For 2 bits, we recommend using `auto-round-best` or `auto-round`.
AutoRound also offer another four recipes, `auto-round-best`, `auto-round-light`,`auto-round-opt-rtn`,`auto-round-rtn`, designed for optimal accuracy and improved speed, respectively.
For 2 bits, we recommend using `auto-round-best` with `--enable_alg_ext`.

</hfoption>


<hfoption id="auto scheme cmd">

### AutoScheme Usage

AutoScheme is a feature that automatically selects the best quantization scheme from the available options for each layer to be quantized in minutes, subject to a target average bit width.


```bash
auto-round \
--model Qwen/Qwen3-0.6B \
--options "W4A16,W2A16G64" \
--target_bits 3.5 \
--output_dir ./tmp_autoround
```
</hfoption>

<hfoption id="alg combinations">

### Algorithm Combinations

AutoRound supports combining multiple algorithms, such as AutoRound, AutoRound + AWQ, and AutoRound + AWQ + Hadamard (very limited support). We are currently expanding support for more algorithms. This enables further optimization of quantization results, especially for scenarios involving quantized activations.

```bash
auto-round \
--model Qwen/Qwen3-0.6B \
--algs "autoround,awq" \
--output_dir ./tmp_autoround
```
</hfoption>


<hfoption id="quantization auto-round api">

### AutoRound API Usage

This setting offers a better trade-off between accuracy and tuning cost, and is recommended in all scenarios.

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
from auto_round import AutoRound

model_name = "facebook/opt-125m"
model = AutoModelForCausalLM.from_pretrained(model_name, dtype="auto")
tokenizer = AutoTokenizer.from_pretrained(model_name)
bits, group_size, sym = 4, 128, True
model_name = "Qwen/Qwen3-0.6B"
# mixed bits config
# layer_config = {"model.decoder.layers.6.self_attn.out_proj": {"bits": 2, "group_size": 32}}
autoround = AutoRound(
model,
tokenizer,
bits=bits,
group_size=group_size,
sym=sym,
ar = AutoRound(
model_name,
scheme="W4A16",
# enable_torch_compile=True,
# layer_config=layer_config,
)

output_dir = "./tmp_autoround"
# format= 'auto_round'(default), 'auto_gptq', 'auto_awq'
autoround.quantize_and_save(output_dir, format='auto_round')
# format= 'auto_round'(default), 'llm_compressor', "gguf:q4_k_m", 'auto_gptq', 'auto_awq'
ar.quantize_and_save(output_dir, format='auto_round')
```

</hfoption>
Expand All @@ -103,26 +129,18 @@ autoround.quantize_and_save(output_dir, format='auto_round')
This setting provides the best accuracy in most scenarios but is 4–5× slower than the standard AutoRound recipe. It is especially recommended for 2-bit quantization and is a good choice if sufficient resources are available.

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
from auto_round import AutoRound

model_name = "facebook/opt-125m"
model = AutoModelForCausalLM.from_pretrained(model_name, dtype="auto")
tokenizer = AutoTokenizer.from_pretrained(model_name)
bits, group_size, sym = 4, 128, True
autoround = AutoRound(
model,
tokenizer,
bits=bits,
group_size=group_size,
sym=sym,
model_name = "Qwen/Qwen3-0.6B"
ar = AutoRound(
model_name,
scheme="W4A16",
nsamples=512,
iters=1000,
low_gpu_mem_usage=True
)

output_dir = "./tmp_autoround"
autoround.quantize_and_save(output_dir, format='auto_round')
ar.quantize_and_save(output_dir, format='auto_round')
```

</hfoption>
Expand All @@ -137,22 +155,15 @@ This setting offers the best speed (2 - 3X faster than AutoRound), but it may ca
from transformers import AutoModelForCausalLM, AutoTokenizer
from auto_round import AutoRound

model_name = "facebook/opt-125m"
model = AutoModelForCausalLM.from_pretrained(model_name, dtype="auto")
tokenizer = AutoTokenizer.from_pretrained(model_name)
bits, group_size, sym = 4, 128, True
autoround = AutoRound(
model,
tokenizer,
bits=bits,
group_size=group_size,
sym=sym,
model_name = "Qwen/Qwen3-0.6B"
ar = AutoRound(
model_name,
iters=50,
lr=5e-3,
)

output_dir = "./tmp_autoround"
autoround.quantize_and_save(output_dir, format='auto_round')
ar.quantize_and_save(output_dir, format='auto_round')
```

</hfoption>
Expand All @@ -176,7 +187,7 @@ AutoRound automatically selects the best available backend based on the installe

### CPU

Supports 2, 4, and 8 bits. We recommend using the AutoRound Kernel (ARK) backend for inference. PyTorch 2.8.0 or later is required with ARK.
Supports 2, 4 and 8 bits. We recommend using the AutoRound Kernel (ARK) backend for inference. PyTorch 2.8.0 or later is required with ARK.

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
Expand All @@ -195,7 +206,7 @@ print(tokenizer.decode(model.generate(**inputs, max_new_tokens=50, do_sample=Fal

### XPU

Supports 4 and 8 bits. We recommend using the AutoRound Kernel (ARK) backend for inference. PyTorch 2.8.0 or later is required with ARK.
Supports 2, 4 and 8 bits. We recommend using the AutoRound Kernel (ARK) backend for inference. PyTorch 2.8.0 or later is required with ARK.

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
Expand All @@ -214,7 +225,7 @@ print(tokenizer.decode(model.generate(**inputs, max_new_tokens=50, do_sample=Fal

### CUDA

Supports 2, 3, 4, and 8 bits. We recommend using GPTQModel for 4 and 8 bits inference.
Supports 2-8 bits. We recommend using GPTQModel for 4 and 8 bits inference.

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
Expand Down
4 changes: 2 additions & 2 deletions src/transformers/utils/quantization_config.py
Original file line number Diff line number Diff line change
Expand Up @@ -211,7 +211,7 @@ class AutoRoundConfig(QuantizationConfigMixin):

Args:
bits (`int`, *optional*, defaults to 4):
The number of bits to quantize to, supported numbers are (2, 3, 4, 8).
The number of bits to quantize to, supported numbers are (2, 3, 4, 5, 6, 7, 8).
group_size (`int`, *optional*, defaults to 128): Group-size value
sym (`bool`, *optional*, defaults to `True`): Symmetric quantization or not
backend (`str`, *optional*, defaults to `"auto"`): The inference backend. By default, AutoRound selects a compatible backend based on the device, quantization settings, and installed libraries. See [Specify inference backend](https://github.com/intel/auto-round/blob/main/docs/step_by_step.md#specify-inference-backend) for all backend options.
Expand All @@ -238,7 +238,7 @@ def __init__(

def post_init(self):
r"""Safety checker that arguments are correct."""
if self.bits not in [2, 3, 4, 8]:
if self.bits not in [2, 3, 4, 5, 6, 7, 8]:
raise ValueError(f"Only support quantization to [2,3,4,8] bits but found {self.bits}")
if self.group_size != -1 and self.group_size <= 0:
raise ValueError("group_size must be greater than 0 or equal to -1")
Expand Down
Loading