ORA MODELS
Compressed Models,
Ready to Deploy.
All models are compressed using OraCompress and verified against baseline benchmarks. Available on Hugging Face.
Qwen3.8-27B-OQ3
From: Qwen 3.8 27B
LanguageMultilingual
10.5 GB
Memory footprint
81%
Smaller than original
3 bit
Mixed quantization
4.5%
Accuracy drop
Qwen3-4B-ORA-W3
From: Qwen 3 4B
LanguageMultilingual
2.2 GB
Memory footprint
73%
Smaller than original
4.37 bit
QAT quantization
3.5%
Accuracy drop
ORA-Llama-47B
From: Meta Llama 3.1 70B
LanguageInstruction-tuned
47B
Parameters
4.1×
Throughput vs. original
72%
Lower cost per token
1 GPU
vs. 4 GPUs
ORA-Qwen-3.5 9B
From: Qwen 3.5 9B
LanguageMultilingual
5.7 GB
Memory footprint
70%
Smaller than original
3.9-bit
Mixed quantization
5%
Accuracy drop