Generating Artificial Outliers in the Absence of Genuine Ones – a Survey

by   Georg Steinbuss, et al.

By definition, outliers are rarely observed in reality, making them difficult to detect or analyse. Artificial outliers approximate such genuine outliers and can, for instance, help with the detection of genuine outliers or with benchmarking outlier-detection algorithms. The literature features different approaches to generate artificial outliers. However, systematic comparison of these approaches remains absent. This surveys and compares these approaches. We start by clarifying the terminology in the field, which varies from publication to publication, and we propose a general problem formulation. Our description of the connection of generating outliers to other research fields like experimental design or generative models frames the field of artificial outliers. Along with offering a concise description, we group the approaches by their general concepts and how they make use of genuine instances. An extensive experimental study reveals the differences between the generation approaches when ultimately being used for outlier detection. This survey shows that the existing approaches already cover a wide range of concepts underlying the generation, but also that the field still has potential for further development. Our experimental study does confirm the expectation that the quality of the generation approaches varies widely, for example, in terms of the data set they are used on. Ultimately, to guide the choice of the generation approach in a specific context, we propose an appropriate general-decision process. In summary, this survey comprises, describes, and connects all relevant work regarding the generation of artificial outliers and may serve as a basis to guide further research in the field.


page 1

page 2

page 3

page 4


Benchmarking Unsupervised Outlier Detection with Realistic Synthetic Data

Benchmarking unsupervised outlier detection is difficult. Outliers are r...

Probabilistic Outlier Detection and Generation

A new method for outlier detection and generation is introduced by lifti...

ODIM: an efficient method to detect outliers via inlier-memorization effect of deep generative models

Identifying whether a given sample is an outlier or not is an important ...

Contextual Outlier Interpretation

Outlier detection plays an essential role in many data-driven applicatio...

On making optimal transport robust to all outliers

Optimal transport (OT) is known to be sensitive against outliers because...

On Classification from Outlier View

Classification is the basis of cognition. Unlike other solutions, this s...

Intelligent decision: towards interpreting the Pe Algorithm

The human intelligence lies in the algorithm, the nature of algorithm li...

Please sign up or login with your details

Forgot password? Click here to reset