All-relevant feature selection using multidimensional filters with exhaustive search

05/16/2017
by   Krzysztof Mnich, et al.
0

This paper describes a method for identification of the informative variables in the information system with discrete decision variables. It is targeted specifically towards discovery of the variables that are non-informative when considered alone, but are informative when the synergistic interactions between multiple variables are considered. To this end, the mutual entropy of all possible k-tuples of variables with decision variable is computed. Then, for each variable the maximal information gain due to interactions with other variables is obtained. For non-informative variables this quantity conforms to the well known statistical distributions. This allows for discerning truly informative variables from non-informative ones. For demonstration of the approach, the method is applied to several synthetic datasets that involve complex multidimensional interactions between variables. It is capable of identifying most important informative variables, even in the case when the dimensionality of the analysis is smaller than the true dimensionality of the problem. What is more, the high sensitivity of the algorithm allows for detection of the influence of nuisance variables on the response variable.

READ FULL TEXT
research
10/31/2018

MDFS - MultiDimensional Feature Selection

Identification of informative variables in an information system is ofte...
research
06/01/2020

A Combined Approach To Detect Key Variables In Thick Data Analytics

In machine learning one of the strategic tasks is the selection of only ...
research
10/19/2012

Sufficient Dimensionality Reduction with Irrelevant Statistics

The problem of finding a reduced dimensionality representation of catego...
research
03/25/2013

Random Intersection Trees

Finding interactions between variables in large and high-dimensional dat...
research
05/12/2016

Context-dependent feature analysis with random forests

In many cases, feature selection is often more complicated than identify...
research
01/16/2021

Informative core identification in complex networks

In network analysis, the core structure of modeling interest is usually ...
research
09/14/2023

Causal Entropy and Information Gain for Measuring Causal Control

Artificial intelligence models and methods commonly lack causal interpre...

Please sign up or login with your details

Forgot password? Click here to reset