• automatically published

Quantization in Machine Learning Models

image.png

Introduction

  • Quantization refers to the process of reducing the precision of the numbers that represent the weights and activations of a machine learning model without significantly reducing its accuracy.
  • It is a critical technique for deploying models on resource-constrained devices like mobile phones, embedded systems, and IoT devices.
  • Quantization (huggingface.co)

Benefits of Quantization

  • Memory Efficiency: Reduces the model size, enabling it to fit in the limited memory of small devices.
  • Computational Efficiency: Lower precision operations are faster and consume less power.
  • Bandwidth Reduction: Smaller models require less data to be transferred when downloaded or updated.

Strategies for Quantization

Sparsification

  • Description: Involves reducing the number of non-zero elements in the model’s weights, effectively compressing the model.
  • Techniques:
    • Weight Pruning: Removing weights that have little impact on the output.
    • Structured Pruning: Removing entire channels or filters that are not contributing significantly to the model’s performance.
  • References:

Removal of Least Significant Bits (LSB)

  • Description: This strategy involves truncating the least significant bits from the weights’ binary representation.
  • Approach:
    • Fixed-Point Quantization: Converts floating-point numbers to fixed-point format, removing the least significant bits.
    • Dynamic Quantization: Adjusts the quantization parameters dynamically based on the distribution of the parameters.
  • Benefits:
    • Reduces the precision of weights with minimal impact on accuracy.
    • Simplifies the hardware implementation of mathematical operations.
  • References:

Uniform and Non-Uniform Quantization

  • Uniform Quantization:
    • Applies the same quantization step size across all values.
    • Easier to implement but might not be optimal for all distributions of model parameters.
  • Non-Uniform Quantization:
    • Adapts the quantization step size according to the distribution of the parameters.
    • Can achieve better accuracy for the same level of compression.
  • References:

Tools and Frameworks for Quantization

  • TensorFlow Lite: Provides tools for post-training quantization and quantization-aware training.
  • TensorFlow Lite Guide
    • PyTorch Quantization: Supports dynamic quantization, static quantization, and quantization-aware training.
  • PyTorch Quantization
    • ONNX Runtime: Offers support for quantized models, enabling optimized inference on different hardware.
  • ONNX Runtime Quantization

Challenges and Considerations

  • Accuracy Trade-offs: Finding the right balance between model size reduction and accuracy preservation.
  • Hardware Compatibility: Ensuring quantized models are compatible with the target hardware’s instruction set.
  • Quantization Granularity: Deciding between per-layer, per-channel, or per-tensor quantization for optimal performance.
  • Quantized Neural Networks (QNNs)
  • Goal: Reduce model size without sacrificing accuracy.
  • Concept: Lower precision representation of weights and activations (e.g., from 32-bit floats to 8-bit integers).

Key Techniques

Overview of GGUF quantization methods : LocalLLaMA (reddit.com)

  • Quantization:

    • Rounding of weights and activations to lower precision representation.

    • Example:

      Quantized Weight = Round(Original Weight / Scale)
      
  • Binary Quantization:

    • Extremely aggressive quantization to binary values (1 or -1).

    • Example:

      Binary Weight = Sign(Original Weight)
      
  • Ternary Quantization:

    • Weights quantized to -1, 0, or 1.
    • Offers better information retention than binary quantization.
  • Quantization-Aware Training (QAT):

    • Integrate quantization effects into the training process for smoother transitions and less accuracy loss.

      Quantization Schemes

  • Fixed-Point: Quantization into fixed bit-width representations.

  • Logarithmic: Leverages logarithmic scale for wider dynamic range.

  • Quantized Inference

  • Running inference using the quantized model.

  • Often requires integer math operations, leading to computational efficiency gains.

  • Dequantization: Process of converting quantized output back to a familiar floating-point representation.

Benefits of QNNs

  • Smaller model sizes: Ideal for memory-constrained devices.
  • Faster inference: Lower precision often leads to faster computations.
  • Reduced power consumption: Benefits embedded systems and mobile devices.
  • Mobile and edge devices
  • Real-time applications
  • Resource-constrained environments

Hyperparameter Tuning (LinkedIn Thread)

  • 𝐇𝐲𝐩𝐞𝐫𝐩𝐚𝐫𝐚𝐦𝐞𝐭𝐞𝐫 𝐎𝐩𝐭𝐢𝐦𝐢𝐳𝐚𝐭𝐢𝐨𝐧: 𝟏𝟎 𝐓𝐨𝐩 𝐏𝐲𝐭𝐡𝐨𝐧 𝐋𝐢𝐛𝐫𝐚𝐫𝐢𝐞𝐬 𝐟𝐨𝐫 𝐒𝐞𝐜𝐫𝐞𝐭 𝐈𝐧𝐠𝐫𝐞𝐝𝐢𝐞𝐧𝐭 𝐢𝐧 𝐌𝐚𝐜𝐡𝐢𝐧𝐞 𝐋𝐞𝐚𝐫𝐧𝐢𝐧𝐠 𝐒𝐮𝐜𝐜𝐞𝐬𝐬
  • Hyperparameter optimization plays a crucial role in determining the performance of a machine learning model. They are one the 3 components of training.
  • 𝟛 ℂ𝕠𝕞𝕡𝕠𝕟𝕖𝕟𝕥𝕤 𝕠𝕗 𝕄𝕠𝕕𝕖𝕝:
  • 1️⃣ Training data: Training data is what the algorithm leverages (think: instructions to build a model) to identify patterns
  • 2️⃣ Parameters: Algorithm ‘learns’ by adjusting parameters, such as weights, based on training data to make accurate predictions, which are saved as part of the final model.
  • 3️⃣ Hyperparameters: Hyperparameters are variables that regulate the process of training and are constant during the training process.
  • 𝔻𝕚𝕗𝕗𝕖𝕣𝕖𝕟𝕥 𝕋𝕪𝕡𝕖𝕤 𝕠𝕗 𝕊𝕖𝕒𝕣𝕔𝕙:
  • 🔎Grid Search : Training models with every possible combination of the provided hyperparameter values a time-consuming process.
  • 🔎Random Search: Training models with randomly samples hyperparameter values from the defined distributions, a more effective search.
  • 🔎 Having Grid Search: Training models with all values, and then repeatedly “halving” the search space by only considering the parameter values that performed the best in the previous round.
  • 🔎 Bayesian Search: Starting with an initial guess of values, using performance of the model to the values. It’s like how a detective might start with a list of suspects, then use new information to narrow down the list.
  • I found these 𝟏𝟎 𝐩𝐲𝐭𝐡𝐨𝐧 𝐥𝐢𝐛𝐫𝐚𝐫𝐢𝐞𝐬 𝐟𝐨𝐫 𝐇𝐲𝐩𝐞𝐫𝐩𝐚𝐫𝐚𝐦𝐞𝐭𝐞𝐫 𝐎𝐩𝐭𝐢𝐦𝐢𝐳𝐚𝐭𝐢𝐨𝐧:
  • 📚 Optuna
  • You can tune estimators of almost any ML, DL package/framework, including Sklearn, PyTorch, TensorFlow, Keras, XGBoost, LightGBM, CatBoost, etc with a real-time Web Dashboard called optuna-dashboard.
  • 📚Hyperopt
  • Optimizing using Bayesian optimization, including conditional dimensions.
  • 📚 Scikit-learn
  • different searches such as GridSearchCV or HalvingGridSearchCV.
  • 📚 Auto-Sklearn
  • AutoML and a drop-in replacement for a scikit-learn estimator.
  • 📚 Hyperactive
  • Very easy to learn but extremly versatile providing intelligent optimization.
  • 📚 Optunity
  • Provides distinct approaches such plethora of score functions.
  • 📚 HyperparameterHunter
  • Automatic save/learn from Experiments for persistent optimization
  • 📚 MLJAR
  • AutoML creating Markdown reports from ML pipeline
  • 📚 KerasTuner
  • with Bayesian Optimization, Hyperband, and Random Search algorithms built-in
  • 📚 Talos
  • Hyperparameter Optimization for TensorFlow, Keras and PyTorch
  • Extra:
  • 📚 Sweeps
  • 📚 Scikit-optimize
  • 📚 PyCaret

No alternative text description for this image