Quantitative measures indicating the relative contribution or influence of individual input features on a machine learning model’s predictions, enabling identification of the most critical variables driving model outputs. Methods include permutation importance, SHAP (SHapley Additive exPlanations) values, and tree-based Gini impurity scores, each providing global or local views of feature influence that support model debugging, data selection, and regulatory explainability requirements.
Semantic Classification
Content
-
Quantitative measures indicating the relative contribution or influence of individual input features on a machine learning model’s predictions, enabling identification of the most critical variables driving model outputs.
Related Terms
-
Broader: Global Explanation, Model Interpretability
-
Narrower: Permutation Importance, SHAP, Feature Attribution
-
Related: Feature Selection, Dimensionality Reduction
Formal Specification
Core Concept
Given model
f: X → Ywith featuresX = {x₁, x₂, ..., xₚ}, feature importanceI(j)quantifies:
I(j) = Influence of feature j on f's predictions
Properties:
-
Non-negative:
I(j) ≥ 0 -
Normalised (optional):
Σ I(j) = 1 -
Ranked: Features ordered by
I(j)Types of Importance
Global Importance: Across all predictions
I_global(j) = E_X[Impact of feature j on f(X)]Local Importance: For specific instance
xI_local(j, x) = Impact of feature j on f(x)Methods
Intrinsic Feature Importance
Linear Model Coefficients
Linear Regression:
y = β₀ + β₁x₁ + β₂x₂ + ... + βₚxₚImportance:
I(j) = |βⱼ| × σⱼWhere
σⱼis standard deviation of featurej(for comparable scales).Interpretation: Absolute standardised coefficient magnitude.
Advantages:
-
Direct from model parameters
-
Computationally free
-
Clear interpretation
Limitations:
-
Assumes linearity
-
Sensitive to multicollinearity
-
Not applicable to non-linear models
Tree-Based Importance
Decision Trees:
I(j) = Σ (samples at node) × (impurity decrease) for all nodes splitting on feature j / (total samples × total impurity decrease)Impurity Measures:
-
Gini impurity:
1 - Σ pᵢ² -
Entropy:
-Σ pᵢ log(pᵢ) -
Variance (regression):
Var(y)Random Forest/Gradient Boosting:
I(j) = Average importance of feature j across all treesAdvantages:
-
Built into tree algorithms
-
Fast computation
-
Handles non-linearity and interactions
Limitations:
-
Bias towards high-cardinality features: More split opportunities
-
Correlated features: Importance split between them
-
Unreliable for extrapolation: Training data dependent
Model-Agnostic Feature Importance
Permutation Importance
Algorithm (Breiman, 2001):
- Compute baseline performance:
S_orig = Score(f, X, y) - For each feature
j: a. Permute feature:X_perm = X with column j shuffledb. Recompute performance:S_perm = Score(f, X_perm, y)c. Importance:I(j) = S_orig - E[S_perm](averaged over repeats)
Properties:
- Compute baseline performance:
-
Model-agnostic
-
Reflects true predictive importance
-
Accounts for feature interactions
Advantages:
-
Unbiased (compared to tree importance)
-
Works with any model
-
Intuitive interpretation
Limitations:
-
Requires access to validation data
-
Computationally expensive (many predictions)
-
Assumes feature independence (can create unrealistic data)
Drop-Column Importance
Algorithm:
- Train model on all features:
f_all(X) → y - For each feature
j: a. Train model without featurej:f_{-j}(X_{-j}) → yb. Importance:I(j) = Score(f_all) - Score(f_{-j})
Interpretation: Performance decrease when feature removed.
Advantages:
- Train model on all features:
-
Directly measures feature necessity
-
No unrealistic data generation
Limitations:
-
Requires retraining
pmodels -
Computationally prohibitive for complex models
-
May underestimate importance if redundant features exist
SHAP Feature Importance
Definition (Lundberg & Lee, 2017):
I(j) = (1/n) Σ |φⱼ(xᵢ)|
i=1
Where φⱼ(xᵢ) is SHAP value for feature j at instance i.
Interpretation: Average absolute contribution of feature j.
Advantages:
-
Consistent with local explanations
-
Theoretically grounded (Shapley values)
-
Handles feature interactions via interaction values
Variants:
-
Mean absolute SHAP:
(1/n) Σ |φⱼ(xᵢ)|(default) -
Mean SHAP:
(1/n) Σ φⱼ(xᵢ)(shows direction) -
SHAP interaction importance: Sum of interaction effects
Limitations:
-
Computationally expensive (especially Kernel SHAP)
-
Baseline dependence
-
Interpretation complexity for interactions
Visualisations
Bar Charts
Structure:
-
Features on y-axis
-
Importance on x-axis
-
Sorted by magnitude
Variants:
-
Standard: Single bar per feature
-
Grouped: Multiple models compared
-
Stacked: Positive/negative contributions
Example (matplotlib):
import matplotlib.pyplot as plt features = ['age', 'income', 'education', 'location'] importances = [0.4, 0.3, 0.2, 0.1] plt.barh(features, importances) plt.xlabel('Feature Importance') plt.title('Permutation Importance')SHAP Summary Plots
Structure:
-
Features on y-axis (sorted by importance)
-
SHAP values on x-axis
-
Each dot is an instance
-
Color indicates feature value (high/low)
Interpretation:
-
Position: SHAP value (impact)
-
Color: Feature value
-
Density: Distribution of impacts
Example:
import shap shap.summary_plot(shap_values, X_test)Feature Importance with Confidence Intervals
Permutation Importance Variance:
Multiple permutations yield distribution:
I(j) ~ N(μⱼ, σⱼ²)Visualisation:
-
Bar chart with error bars
-
Box plots showing distribution
-
Violin plots for full distribution
Example (scikit-learn):
from sklearn.inspection import permutation_importance result = permutation_importance(model, X_val, y_val, n_repeats=30) importances_mean = result.importances_mean importances_std = result.importances_std plt.barh(features, importances_mean, xerr=importances_std)Application Domains
Feature Selection
Use Case: Identify and retain only important features.
Approach:
- Compute feature importance
- Threshold or select top-K features
- Retrain model on reduced feature set
- Evaluate performance
Benefits:
-
Reduced overfitting
-
Faster training/inference
-
Improved interpretability
Example:
from sklearn.feature_selection import SelectFromModel # Using tree-based importance selector = SelectFromModel(RandomForestClassifier(), threshold='median') X_selected = selector.fit_transform(X_train, y_train)Model Debugging
Use Cases:
-
Data leakage detection: Unexpected high importance
-
Sanity checks: Aligns with domain knowledge?
-
Bias detection: Protected attributes driving predictions?
Example: Feature importance reveals
customer_idhas high importance → data leakage likely.Domain Insight
Scientific Applications:
-
Hypothesis generation (which variables matter?)
-
Mechanism understanding (how do variables influence outcome?)
-
Prioritisation (which factors to intervene on?)
Example (Healthcare): Feature importance shows “blood pressure” more important than “BMI” for heart disease prediction → clinical validation and insight.
Regulatory Compliance
Finance:
-
Fair lending: Ensure protected attributes not driving decisions
-
Model risk management: Understand key risk factors
Example: Feature importance analysis shows
racehas near-zero importance → compliance with fair lending laws.Implementation Approaches
Scikit-learn Tree Importance
from sklearn.ensemble import RandomForestClassifier
model = RandomForestClassifier().fit(X_train, y_train)
importances = model.feature_importances_
indices = np.argsort(importances)[::-1]
for i in range(X_train.shape[1]):
print(f"{features[indices[i]]}: {importances[indices[i]]:.4f}")Scikit-learn Permutation Importance
from sklearn.inspection import permutation_importance
result = permutation_importance(
estimator=model,
X=X_val,
y=y_val,
n_repeats=10,
random_state=42,
scoring='accuracy'
)
for i, (mean, std) in enumerate(zip(result.importances_mean, result.importances_std)):
print(f"{features[i]}: {mean:.4f} ± {std:.4f}")SHAP Feature Importance
import shap
explainer = shap.TreeExplainer(model)
shap_values = explainer.shap_values(X_test)
# Global feature importance
shap.summary_plot(shap_values, X_test, plot_type="bar")
# Or manually compute
feature_importance = np.abs(shap_values).mean(axis=0)Custom Importance Function
def custom_feature_importance(model, X, y, metric, n_repeats=10):
"""
Model-agnostic permutation importance with custom metric.
"""
baseline_score = metric(y, model.predict(X))
importances = {}
for col in X.columns:
scores = []
for _ in range(n_repeats):
X_perm = X.copy()
X_perm[col] = np.random.permutation(X_perm[col])
score = metric(y, model.predict(X_perm))
scores.append(baseline_score - score)
importances[col] = {
'mean': np.mean(scores),
'std': np.std(scores)
}
return importancesEvaluation & Validation
Consistency Checks
Across Methods: Compare rankings from different importance methods:
from scipy.stats import spearmanr
correlation = spearmanr(
tree_importance_ranking,
permutation_importance_ranking
)Across Subsets: Importance should be stable across data subsets:
from sklearn.model_selection import KFold
importances_per_fold = []
for train_idx, val_idx in KFold(n_splits=5).split(X):
# Compute importance on fold
importances_per_fold.append(compute_importance(X[val_idx], y[val_idx]))
# Check variance
importance_std = np.std(importances_per_fold, axis=0)Domain Validation
Expert Review:
-
Do important features align with domain knowledge?
-
Are there unexpected importances?
-
Are known important features captured?
Hypothesis Testing:
-
Is
I(j)significantly greater than zero? -
Permutation test or bootstrap confidence intervals
Example:
# Permutation test for significance null_distribution = [] for _ in range(1000): y_permuted = np.random.permutation(y) null_importance = compute_importance(X, y_permuted) null_distribution.append(null_importance) p_value = (null_distribution >= observed_importance).mean()Robustness Analysis
Stability to Noise: Add random features and verify low importance:
X_with_noise = X.copy() X_with_noise['random'] = np.random.randn(len(X)) importances = compute_importance(X_with_noise, y) assert importances['random'] < threshold # Near zeroSensitivity to Outliers: Recompute importance with outliers removed.
Challenges & Limitations
Methodological Challenges
Correlated Features:
-
Tree importance: Split between correlated features
-
Permutation: Unrealistic combinations if features correlated
-
Solution: Conditional importance, clustered permutation
Feature Interactions:
-
Standard importance: Doesn’t capture interactions
-
Solution: SHAP interaction values, H-statistic
Causality:
-
Importance ≠ causal effect
-
Observational data limitations
-
Solution: Causal inference methods, interventional importance
Computational Challenges
Scalability:
-
Permutation:
O(p × r × n)predictions -
SHAP: Exponential (exact), polynomial (approximate)
-
Drop-column: Requires
pmodel retrainsReal-time Constraints:
-
Production systems: Pre-compute importance
-
Online learning: Incremental importance updates
Interpretation Challenges
Audience Dependence:
-
Technical vs. lay users
-
Absolute vs. relative importance
-
Positive vs. negative effects
Multicollinearity:
-
Inflated coefficient variance (linear models)
-
Shared importance (tree models)
-
Solution: Regularisation, feature engineering
Research Directions
Emerging Areas
Causal Feature Importance:
-
Interventional importance:
I(j) = E[Y | do(X_j)] - E[Y] -
Counterfactual reasoning
-
Structural causal models
Conditional Importance:
-
Importance given realistic feature combinations
-
Conditional permutation schemes
-
Addressing feature dependence
Temporal Feature Importance:
-
Time-varying importance (online learning)
-
Concept drift detection
-
Dynamic feature selection
Multi-task Feature Importance:
-
Shared importance across tasks
-
Task-specific importance decomposition
Industry Innovation
Microsoft InterpretML:
-
Unified importance API
-
EBM feature importance (additive effects)
Google Cloud Explainable AI:
-
Feature attribution aggregation
-
Integrated with Vertex AI
H2O.ai Driverless AI:
-
Automated feature importance
-
Ensemble importance across models
Best Practices
Method Selection
Decision Tree:
- Model type: Trees → intrinsic; Neural nets → permutation/SHAP
- Computational budget: Limited → tree intrinsic; Ample → SHAP
- Feature correlation: High → SHAP/conditional; Low → permutation
- Causality: Important → causal methods; Prediction → standard
Implementation Guidelines
Pre-analysis:
-
Check for correlated features (VIF, correlation matrix)
-
Validate data quality (missing values, outliers)
-
Establish domain priors (expected important features)
Analysis:
-
Use multiple methods (robustness check)
-
Include confidence intervals (permutation variance)
-
Validate against domain knowledge
-
Test for statistical significance
Post-analysis:
-
Document methodology and parameters
-
Visualise with clear labels
-
Highlight top-K features
-
Disclose limitations
Visualisation Guidelines
Clarity:
-
Sort by importance (descending)
-
Limit to top-K features (avoid clutter)
-
Include error bars (uncertainty)
-
Use color judiciously (positive/negative)
Context:
-
Show feature scales (if relevant)
-
Include baseline (zero importance line)
-
Annotate unexpected results
-
Provide interpretation guide
References
Academic Literature
-
Breiman, L. (2001). “Random forests.” Machine Learning, 45(1), 5-32
-
Lundberg, S. M., & Lee, S. I. (2017). “A unified approach to interpreting model predictions.” NeurIPS
-
Strobl, C., et al. (2007). “Bias in random forest variable importance measures.” BMC Bioinformatics, 8(1), 25
-
Fisher, A., Rudin, C., & Dominici, F. (2019). “All models are wrong, but many are useful: Learning a variable’s importance by studying an entire class of prediction models simultaneously.” Journal of Machine Learning Research, 20(177), 1-81
Standards
-
IEEE. (2023). IEEE P2976: Standard for eXplainable Artificial Intelligence
Tools & Frameworks
-
Scikit-learn. (2023). Inspection module: permutation_importance
-
Lundberg, S. M. (2023). SHAP library
-
Molnar, C. (2022). Interpretable Machine Learning: Feature Importance
See Also