Quantitative measures indicating the relative contribution or influence of individual input features on a machine learning model’s predictions, enabling identification of the most critical variables driving model outputs. Methods include permutation importance, SHAP (SHapley Additive exPlanations) values, and tree-based Gini impurity scores, each providing global or local views of feature influence that support model debugging, data selection, and regulatory explainability requirements.

Semantic Classification

Content

I(j) = Influence of feature j on f's predictions

Properties:

  • Non-negative: I(j) ≥ 0

  • Normalised (optional): Σ I(j) = 1

  • Ranked: Features ordered by I(j)

    Types of Importance

    Global Importance: Across all predictions

    I_global(j) = E_X[Impact of feature j on f(X)]
    

    Local Importance: For specific instance x

    I_local(j, x) = Impact of feature j on f(x)
    

    Methods

    Intrinsic Feature Importance

    Linear Model Coefficients

    Linear Regression:

    y = β₀ + β₁x₁ + β₂x₂ + ... + βₚxₚ
    

    Importance:

    I(j) = |βⱼ| × σⱼ
    

    Where σⱼ is standard deviation of feature j (for comparable scales).

    Interpretation: Absolute standardised coefficient magnitude.

    Advantages:

  • Direct from model parameters

  • Computationally free

  • Clear interpretation

    Limitations:

  • Assumes linearity

  • Sensitive to multicollinearity

  • Not applicable to non-linear models

    Tree-Based Importance

    Decision Trees:

    I(j) = Σ (samples at node) × (impurity decrease) for all nodes splitting on feature j
       / (total samples × total impurity decrease)
    

    Impurity Measures:

  • Gini impurity: 1 - Σ pᵢ²

  • Entropy: -Σ pᵢ log(pᵢ)

  • Variance (regression): Var(y)

    Random Forest/Gradient Boosting:

    I(j) = Average importance of feature j across all trees
    

    Advantages:

  • Built into tree algorithms

  • Fast computation

  • Handles non-linearity and interactions

    Limitations:

  • Bias towards high-cardinality features: More split opportunities

  • Correlated features: Importance split between them

  • Unreliable for extrapolation: Training data dependent

    Model-Agnostic Feature Importance

    Permutation Importance

    Algorithm (Breiman, 2001):

    1. Compute baseline performance: S_orig = Score(f, X, y)
    2. For each feature j: a. Permute feature: X_perm = X with column j shuffled b. Recompute performance: S_perm = Score(f, X_perm, y) c. Importance: I(j) = S_orig - E[S_perm] (averaged over repeats)

    Properties:

  • Model-agnostic

  • Reflects true predictive importance

  • Accounts for feature interactions

    Advantages:

  • Unbiased (compared to tree importance)

  • Works with any model

  • Intuitive interpretation

    Limitations:

  • Requires access to validation data

  • Computationally expensive (many predictions)

  • Assumes feature independence (can create unrealistic data)

    Drop-Column Importance

    Algorithm:

    1. Train model on all features: f_all(X) → y
    2. For each feature j: a. Train model without feature j: f_{-j}(X_{-j}) → y b. Importance: I(j) = Score(f_all) - Score(f_{-j})

    Interpretation: Performance decrease when feature removed.

    Advantages:

  • Directly measures feature necessity

  • No unrealistic data generation

    Limitations:

  • Requires retraining p models

  • Computationally prohibitive for complex models

  • May underestimate importance if redundant features exist

    SHAP Feature Importance

    Definition (Lundberg & Lee, 2017):

I(j) = (1/n) Σ |φⱼ(xᵢ)|
           i=1

Where φⱼ(xᵢ) is SHAP value for feature j at instance i.

Interpretation: Average absolute contribution of feature j.

Advantages:

  • Consistent with local explanations

  • Theoretically grounded (Shapley values)

  • Handles feature interactions via interaction values

    Variants:

  • Mean absolute SHAP: (1/n) Σ |φⱼ(xᵢ)| (default)

  • Mean SHAP: (1/n) Σ φⱼ(xᵢ) (shows direction)

  • SHAP interaction importance: Sum of interaction effects

    Limitations:

  • Computationally expensive (especially Kernel SHAP)

  • Baseline dependence

  • Interpretation complexity for interactions

    Visualisations

    Bar Charts

    Structure:

  • Features on y-axis

  • Importance on x-axis

  • Sorted by magnitude

    Variants:

  • Standard: Single bar per feature

  • Grouped: Multiple models compared

  • Stacked: Positive/negative contributions

    Example (matplotlib):

    import matplotlib.pyplot as plt
     
    features = ['age', 'income', 'education', 'location']
    importances = [0.4, 0.3, 0.2, 0.1]
     
    plt.barh(features, importances)
    plt.xlabel('Feature Importance')
    plt.title('Permutation Importance')

    SHAP Summary Plots

    Structure:

  • Features on y-axis (sorted by importance)

  • SHAP values on x-axis

  • Each dot is an instance

  • Color indicates feature value (high/low)

    Interpretation:

  • Position: SHAP value (impact)

  • Color: Feature value

  • Density: Distribution of impacts

    Example:

    import shap
     
    shap.summary_plot(shap_values, X_test)

    Feature Importance with Confidence Intervals

    Permutation Importance Variance:

    Multiple permutations yield distribution:

    I(j) ~ N(μⱼ, σⱼ²)
    

    Visualisation:

  • Bar chart with error bars

  • Box plots showing distribution

  • Violin plots for full distribution

    Example (scikit-learn):

    from sklearn.inspection import permutation_importance
     
    result = permutation_importance(model, X_val, y_val, n_repeats=30)
     
    importances_mean = result.importances_mean
    importances_std = result.importances_std
     
    plt.barh(features, importances_mean, xerr=importances_std)

    Application Domains

    Feature Selection

    Use Case: Identify and retain only important features.

    Approach:

    1. Compute feature importance
    2. Threshold or select top-K features
    3. Retrain model on reduced feature set
    4. Evaluate performance

    Benefits:

  • Reduced overfitting

  • Faster training/inference

  • Improved interpretability

    Example:

    from sklearn.feature_selection import SelectFromModel
     
    # Using tree-based importance
    selector = SelectFromModel(RandomForestClassifier(), threshold='median')
    X_selected = selector.fit_transform(X_train, y_train)

    Model Debugging

    Use Cases:

  • Data leakage detection: Unexpected high importance

  • Sanity checks: Aligns with domain knowledge?

  • Bias detection: Protected attributes driving predictions?

    Example: Feature importance reveals customer_id has high importance → data leakage likely.

    Domain Insight

    Scientific Applications:

  • Hypothesis generation (which variables matter?)

  • Mechanism understanding (how do variables influence outcome?)

  • Prioritisation (which factors to intervene on?)

    Example (Healthcare): Feature importance shows “blood pressure” more important than “BMI” for heart disease prediction → clinical validation and insight.

    Regulatory Compliance

    Finance:

  • Fair lending: Ensure protected attributes not driving decisions

  • Model risk management: Understand key risk factors

    Example: Feature importance analysis shows race has near-zero importance → compliance with fair lending laws.

    Implementation Approaches

    Scikit-learn Tree Importance

from sklearn.ensemble import RandomForestClassifier
 
model = RandomForestClassifier().fit(X_train, y_train)
 
importances = model.feature_importances_
indices = np.argsort(importances)[::-1]
 
for i in range(X_train.shape[1]):
  print(f"{features[indices[i]]}: {importances[indices[i]]:.4f}")

Scikit-learn Permutation Importance

from sklearn.inspection import permutation_importance
 
result = permutation_importance(
  estimator=model,
  X=X_val,
  y=y_val,
  n_repeats=10,
  random_state=42,
  scoring='accuracy'
)
 
for i, (mean, std) in enumerate(zip(result.importances_mean, result.importances_std)):
  print(f"{features[i]}: {mean:.4f} ± {std:.4f}")

SHAP Feature Importance

import shap
 
explainer = shap.TreeExplainer(model)
shap_values = explainer.shap_values(X_test)
 
# Global feature importance
shap.summary_plot(shap_values, X_test, plot_type="bar")
 
# Or manually compute
feature_importance = np.abs(shap_values).mean(axis=0)

Custom Importance Function

def custom_feature_importance(model, X, y, metric, n_repeats=10):
  """
  Model-agnostic permutation importance with custom metric.
  """
  baseline_score = metric(y, model.predict(X))
  importances = {}
 
  for col in X.columns:
      scores = []
      for _ in range(n_repeats):
          X_perm = X.copy()
          X_perm[col] = np.random.permutation(X_perm[col])
          score = metric(y, model.predict(X_perm))
          scores.append(baseline_score - score)
 
      importances[col] = {
          'mean': np.mean(scores),
          'std': np.std(scores)
      }
 
  return importances

Evaluation & Validation

Consistency Checks

Across Methods: Compare rankings from different importance methods:

from scipy.stats import spearmanr
 
correlation = spearmanr(
  tree_importance_ranking,
  permutation_importance_ranking
)

Across Subsets: Importance should be stable across data subsets:

from sklearn.model_selection import KFold
 
importances_per_fold = []
for train_idx, val_idx in KFold(n_splits=5).split(X):
  # Compute importance on fold
  importances_per_fold.append(compute_importance(X[val_idx], y[val_idx]))
 
# Check variance
importance_std = np.std(importances_per_fold, axis=0)

Domain Validation

Expert Review:

  • Do important features align with domain knowledge?

  • Are there unexpected importances?

  • Are known important features captured?

    Hypothesis Testing:

  • Is I(j) significantly greater than zero?

  • Permutation test or bootstrap confidence intervals

    Example:

    # Permutation test for significance
    null_distribution = []
    for _ in range(1000):
    y_permuted = np.random.permutation(y)
    null_importance = compute_importance(X, y_permuted)
    null_distribution.append(null_importance)
     
    p_value = (null_distribution >= observed_importance).mean()

    Robustness Analysis

    Stability to Noise: Add random features and verify low importance:

    X_with_noise = X.copy()
    X_with_noise['random'] = np.random.randn(len(X))
     
    importances = compute_importance(X_with_noise, y)
    assert importances['random'] < threshold  # Near zero

    Sensitivity to Outliers: Recompute importance with outliers removed.

    Challenges & Limitations

    Methodological Challenges

    Correlated Features:

  • Tree importance: Split between correlated features

  • Permutation: Unrealistic combinations if features correlated

  • Solution: Conditional importance, clustered permutation

    Feature Interactions:

  • Standard importance: Doesn’t capture interactions

  • Solution: SHAP interaction values, H-statistic

    Causality:

  • Importance ≠ causal effect

  • Observational data limitations

  • Solution: Causal inference methods, interventional importance

    Computational Challenges

    Scalability:

  • Permutation: O(p × r × n) predictions

  • SHAP: Exponential (exact), polynomial (approximate)

  • Drop-column: Requires p model retrains

    Real-time Constraints:

  • Production systems: Pre-compute importance

  • Online learning: Incremental importance updates

    Interpretation Challenges

    Audience Dependence:

  • Technical vs. lay users

  • Absolute vs. relative importance

  • Positive vs. negative effects

    Multicollinearity:

  • Inflated coefficient variance (linear models)

  • Shared importance (tree models)

  • Solution: Regularisation, feature engineering

    Research Directions

    Emerging Areas

    Causal Feature Importance:

  • Interventional importance: I(j) = E[Y | do(X_j)] - E[Y]

  • Counterfactual reasoning

  • Structural causal models

    Conditional Importance:

  • Importance given realistic feature combinations

  • Conditional permutation schemes

  • Addressing feature dependence

    Temporal Feature Importance:

  • Time-varying importance (online learning)

  • Concept drift detection

  • Dynamic feature selection

    Multi-task Feature Importance:

  • Shared importance across tasks

  • Task-specific importance decomposition

    Industry Innovation

    Microsoft InterpretML:

  • Unified importance API

  • EBM feature importance (additive effects)

    Google Cloud Explainable AI:

  • Feature attribution aggregation

  • Integrated with Vertex AI

    H2O.ai Driverless AI:

  • Automated feature importance

  • Ensemble importance across models

    Best Practices

    Method Selection

    Decision Tree:

    1. Model type: Trees → intrinsic; Neural nets → permutation/SHAP
    2. Computational budget: Limited → tree intrinsic; Ample → SHAP
    3. Feature correlation: High → SHAP/conditional; Low → permutation
    4. Causality: Important → causal methods; Prediction → standard

    Implementation Guidelines

    Pre-analysis:

  • Check for correlated features (VIF, correlation matrix)

  • Validate data quality (missing values, outliers)

  • Establish domain priors (expected important features)

    Analysis:

  • Use multiple methods (robustness check)

  • Include confidence intervals (permutation variance)

  • Validate against domain knowledge

  • Test for statistical significance

    Post-analysis:

  • Document methodology and parameters

  • Visualise with clear labels

  • Highlight top-K features

  • Disclose limitations

    Visualisation Guidelines

    Clarity:

  • Sort by importance (descending)

  • Limit to top-K features (avoid clutter)

  • Include error bars (uncertainty)

  • Use color judiciously (positive/negative)

    Context:

  • Show feature scales (if relevant)

  • Include baseline (zero importance line)

  • Annotate unexpected results

  • Provide interpretation guide

    References

    Academic Literature

  • Breiman, L. (2001). “Random forests.” Machine Learning, 45(1), 5-32

  • Lundberg, S. M., & Lee, S. I. (2017). “A unified approach to interpreting model predictions.” NeurIPS

  • Strobl, C., et al. (2007). “Bias in random forest variable importance measures.” BMC Bioinformatics, 8(1), 25

  • Fisher, A., Rudin, C., & Dominici, F. (2019). “All models are wrong, but many are useful: Learning a variable’s importance by studying an entire class of prediction models simultaneously.” Journal of Machine Learning Research, 20(177), 1-81

    Standards

  • IEEE. (2023). IEEE P2976: Standard for eXplainable Artificial Intelligence

    Tools & Frameworks

  • Scikit-learn. (2023). Inspection module: permutation_importance

  • Lundberg, S. M. (2023). SHAP library

  • Molnar, C. (2022). Interpretable Machine Learning: Feature Importance

    See Also

  • Permutation Importance

  • SHAP

  • Feature Attribution

  • Global Explanation

  • Feature Selection

  • Partial Dependence Plot

Provenance