- automatically published
Quantization in Machine Learning Models

Introduction
- Quantization refers to the process of reducing the precision of the numbers that represent the weights and activations of a machine learning model without significantly reducing its accuracy.
- It is a critical technique for deploying models on resource-constrained devices like mobile phones, embedded systems, and IoT devices.
- Quantization (huggingface.co)
Benefits of Quantization
- Memory Efficiency: Reduces the model size, enabling it to fit in the limited memory of small devices.
- Computational Efficiency: Lower precision operations are faster and consume less power.
- Bandwidth Reduction: Smaller models require less data to be transferred when downloaded or updated.
Strategies for Quantization
Sparsification
- Description: Involves reducing the number of non-zero elements in the model’s weights, effectively compressing the model.
- Techniques:
- Weight Pruning: Removing weights that have little impact on the output.
- Structured Pruning: Removing entire channels or filters that are not contributing significantly to the model’s performance.
- References:
Removal of Least Significant Bits (LSB)
- Description: This strategy involves truncating the least significant bits from the weights’ binary representation.
- Approach:
- Fixed-Point Quantization: Converts floating-point numbers to fixed-point format, removing the least significant bits.
- Dynamic Quantization: Adjusts the quantization parameters dynamically based on the distribution of the parameters.
- Benefits:
- Reduces the precision of weights with minimal impact on accuracy.
- Simplifies the hardware implementation of mathematical operations.
- References:
Uniform and Non-Uniform Quantization
- Uniform Quantization:
- Applies the same quantization step size across all values.
- Easier to implement but might not be optimal for all distributions of model parameters.
- Non-Uniform Quantization:
- Adapts the quantization step size according to the distribution of the parameters.
- Can achieve better accuracy for the same level of compression.
- References:
Tools and Frameworks for Quantization
- TensorFlow Lite: Provides tools for post-training quantization and quantization-aware training.
- TensorFlow Lite Guide
- PyTorch Quantization: Supports dynamic quantization, static quantization, and quantization-aware training.
- PyTorch Quantization
- ONNX Runtime: Offers support for quantized models, enabling optimized inference on different hardware.
- ONNX Runtime Quantization
Challenges and Considerations
- Accuracy Trade-offs: Finding the right balance between model size reduction and accuracy preservation.
- Hardware Compatibility: Ensuring quantized models are compatible with the target hardware’s instruction set.
- Quantization Granularity: Deciding between per-layer, per-channel, or per-tensor quantization for optimal performance.
- Quantized Neural Networks (QNNs)
- Goal: Reduce model size without sacrificing accuracy.
- Concept: Lower precision representation of weights and activations (e.g., from 32-bit floats to 8-bit integers).
Key Techniques
Overview of GGUF quantization methods : LocalLLaMA (reddit.com)
-
Quantization:
-
Rounding of weights and activations to lower precision representation.
-
Example:
Quantized Weight = Round(Original Weight / Scale)
-
-
Binary Quantization:
-
Extremely aggressive quantization to binary values (1 or -1).
-
Example:
Binary Weight = Sign(Original Weight)
-
-
Ternary Quantization:
- Weights quantized to -1, 0, or 1.
- Offers better information retention than binary quantization.
-
Quantization-Aware Training (QAT):
-
Integrate quantization effects into the training process for smoother transitions and less accuracy loss.
Quantization Schemes
-
-
Fixed-Point: Quantization into fixed bit-width representations.
-
Logarithmic: Leverages logarithmic scale for wider dynamic range.
-
Quantized Inference
-
Running inference using the quantized model.
-
Often requires integer math operations, leading to computational efficiency gains.
-
Dequantization: Process of converting quantized output back to a familiar floating-point representation.
Benefits of QNNs
- Smaller model sizes: Ideal for memory-constrained devices.
- Faster inference: Lower precision often leads to faster computations.
- Reduced power consumption: Benefits embedded systems and mobile devices.
- Mobile and edge devices
- Real-time applications
- Resource-constrained environments
Hyperparameter Tuning (LinkedIn Thread)
- 𝐇𝐲𝐩𝐞𝐫𝐩𝐚𝐫𝐚𝐦𝐞𝐭𝐞𝐫 𝐎𝐩𝐭𝐢𝐦𝐢𝐳𝐚𝐭𝐢𝐨𝐧: 𝟏𝟎 𝐓𝐨𝐩 𝐏𝐲𝐭𝐡𝐨𝐧 𝐋𝐢𝐛𝐫𝐚𝐫𝐢𝐞𝐬 𝐟𝐨𝐫 𝐒𝐞𝐜𝐫𝐞𝐭 𝐈𝐧𝐠𝐫𝐞𝐝𝐢𝐞𝐧𝐭 𝐢𝐧 𝐌𝐚𝐜𝐡𝐢𝐧𝐞 𝐋𝐞𝐚𝐫𝐧𝐢𝐧𝐠 𝐒𝐮𝐜𝐜𝐞𝐬𝐬
- Hyperparameter optimization plays a crucial role in determining the performance of a machine learning model. They are one the 3 components of training.
- 𝟛 ℂ𝕠𝕞𝕡𝕠𝕟𝕖𝕟𝕥𝕤 𝕠𝕗 𝕄𝕠𝕕𝕖𝕝:
- 1️⃣ Training data: Training data is what the algorithm leverages (think: instructions to build a model) to identify patterns
- 2️⃣ Parameters: Algorithm ‘learns’ by adjusting parameters, such as weights, based on training data to make accurate predictions, which are saved as part of the final model.
- 3️⃣ Hyperparameters: Hyperparameters are variables that regulate the process of training and are constant during the training process.
- 𝔻𝕚𝕗𝕗𝕖𝕣𝕖𝕟𝕥 𝕋𝕪𝕡𝕖𝕤 𝕠𝕗 𝕊𝕖𝕒𝕣𝕔𝕙:
- 🔎Grid Search : Training models with every possible combination of the provided hyperparameter values a time-consuming process.
- 🔎Random Search: Training models with randomly samples hyperparameter values from the defined distributions, a more effective search.
- 🔎 Having Grid Search: Training models with all values, and then repeatedly “halving” the search space by only considering the parameter values that performed the best in the previous round.
- 🔎 Bayesian Search: Starting with an initial guess of values, using performance of the model to the values. It’s like how a detective might start with a list of suspects, then use new information to narrow down the list.
- I found these 𝟏𝟎 𝐩𝐲𝐭𝐡𝐨𝐧 𝐥𝐢𝐛𝐫𝐚𝐫𝐢𝐞𝐬 𝐟𝐨𝐫 𝐇𝐲𝐩𝐞𝐫𝐩𝐚𝐫𝐚𝐦𝐞𝐭𝐞𝐫 𝐎𝐩𝐭𝐢𝐦𝐢𝐳𝐚𝐭𝐢𝐨𝐧:
- 📚 Optuna
- You can tune estimators of almost any ML, DL package/framework, including Sklearn, PyTorch, TensorFlow, Keras, XGBoost, LightGBM, CatBoost, etc with a real-time Web Dashboard called optuna-dashboard.
- 📚Hyperopt
- Optimizing using Bayesian optimization, including conditional dimensions.
- 📚 Scikit-learn
- different searches such as GridSearchCV or HalvingGridSearchCV.
- 📚 Auto-Sklearn
- AutoML and a drop-in replacement for a scikit-learn estimator.
- 📚 Hyperactive
- Very easy to learn but extremly versatile providing intelligent optimization.
- 📚 Optunity
- Provides distinct approaches such plethora of score functions.
- 📚 HyperparameterHunter
- Automatic save/learn from Experiments for persistent optimization
- 📚 MLJAR
- AutoML creating Markdown reports from ML pipeline
- 📚 KerasTuner
- with Bayesian Optimization, Hyperband, and Random Search algorithms built-in
- 📚 Talos
- Hyperparameter Optimization for TensorFlow, Keras and PyTorch
- Extra:
- 📚 Sweeps
- 📚 Scikit-optimize
- 📚 PyCaret
- [The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction
- Microsoft Research](https://www.microsoft.com/en-us/research/publication/the-truth-is-in-there-improving-reasoning-in-language-models-with-layer-selective-rank-reduction/)
- pratyushasharma/laser: The Truth Is In There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction (github.com)
- huggingface/optimum-nvidia (github.com)
- [width=0.06]./figs/logo EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty (arxiv.org)
- run-ai/llmperf (github.com) Tensor vs serving frameworks
- Bitnet and the rise of the 1bit model