TECHNICAL•2026-08-21•13 min read
Quantizing Small Models to 3 and 2 Bits with ORA-QAT
We push Qwen3-4B from 16 bits to 3 and 2 with our novel algorithm for efficient quantization-aware training. At 3 bits, the model occupies only 22% of the BF16 size and keeps almost 97% of full-precision quality, beating the GPTQ baseline by 5.3 points on average; at 2 bits it occupies only 18% and more than doubles GPTQ's score. The training runs on a single GPU and produces a quantized checkpoint in just an hour.
Read more→