Q-Lite (Standard): Standard INT8 quantization optimized for fast inference. Recommended as the default choice.
Q-Pro (Advanced): High-precision quantization with fine-tuning to maximize accuracy. Compile with the --use_q_pro option. (Note: Requires longer compilation time.)
Q-Master (Ultimate/QAT): Quantization-Aware Training (QAT) pipeline that fine-tunes the model while simulating quantization to recover accuracy loss. (Note: Requires training dataset and longer execution time.)
* Note: Performance results may vary depending on the specific hardware configuration.
Accuracy Champions (Within 1%p)
Definition: Cases where Quantized (INT8) Accuracy stays within 1 percentage point of Original (FP32) Accuracy, judged by metric direction. For higher-is-better metrics (Top-1, mAP, mIoU, PSNR, SSIM, AP, etc.): Quantized ≥ Original - 1.0. For lower-is-better metrics (RMSE, NME, MNAE, ADD, etc.): Quantized ≤ Original + 1.0.
Qualifying accuracy values in the table below are shown in bold.
License Notice
Please review the following before downloading the model:
I have reviewed the license terms for the model I want to download.
If I distribute a commercial product using this model, I may be obligated to disclose my source code.
DEEPX is not responsible for any license disputes arising from the use of this model.