A Notion of Feature Importance by Decorrelation and Detection of Trends by Random Forest Regression

by   Yannick Gerstorfer, et al.

In many studies, we want to determine the influence of certain features on a dependent variable. More specifically, we are interested in the strength of the influence – i.e., is the feature relevant? – and, if so, how the feature influences the dependent variable. Recently, data-driven approaches such as random forest regression have found their way into applications (Boulesteix et al., 2012). These models allow to directly derive measures of feature importance, which are a natural indicator of the strength of the influence. For the relevant features, the correlation or rank correlation between the feature and the dependent variable has typically been used to determine the nature of the influence. More recent methods, some of which can also measure interactions between features, are based on a modeling approach. In particular, when machine learning models are used, SHAP scores are a recent and prominent method to determine these trends (Lundberg et al., 2017). In this paper, we introduce a novel notion of feature importance based on the well-studied Gram-Schmidt decorrelation method. Furthermore, we propose two estimators for identifying trends in the data using random forest regression, the so-called absolute and relative transversal rate. We empirically compare the properties of our estimators with those of well-established estimators on a variety of synthetic and real-world datasets.


page 1

page 2

page 3

page 4


Geometry- and Accuracy-Preserving Random Forest Proximities

Random forests are considered one of the best out-of-the-box classificat...

Prediction Error Reduction Function as a Variable Importance Score

This paper introduces and develops a novel variable importance score fun...

Single Sample Feature Importance: An Interpretable Algorithm for Low-Level Feature Analysis

Have you ever wondered how your feature space is impacting the predictio...

Evaluating Feature Importance Estimates

Estimating the influence of a given feature to a model prediction is cha...

A Feature Importance Analysis for Soft-Sensing-Based Predictions in a Chemical Sulphonation Process

In this paper we present the results of a feature importance analysis of...

Intercomparison of Brown Dwarf Model Grids and Atmospheric Retrieval Using Machine Learning

Understanding differences between sub-stellar spectral data and models h...

Reconstructing dynamical networks via feature ranking

Empirical data on real complex systems are becoming increasingly availab...

Please sign up or login with your details

Forgot password? Click here to reset