Valid Inference Corrected for Outlier Removal

by   Shuxiao Chen, et al.

Ordinary least square (OLS) estimation of a linear regression model is well-known to be highly sensitive to outliers. It is common practice to first identify and remove outliers by looking at the data then to fit OLS and form confidence intervals and p-values on the remaining data as if this were the original data collected. We show in this paper that this "detect-and-forget" approach can lead to invalid inference, and we propose a framework that properly accounts for outlier detection and removal to provide valid confidence intervals and hypothesis tests. Our inferential procedures apply to any outlier removal procedure that can be characterized by a set of quadratic constraints on the response vector, and we show that several of the most commonly used outlier detection procedures are of this form. Our methodology is built upon recent advances in selective inference (Taylor & Tibshirani 2015), which are focused on inference corrected for variable selection. We conduct simulations to corroborate the theoretical results, and we apply our method to two classic data sets considered in the outlier detection literature to illustrate how our inferential results can differ from the traditional detect-and-forget strategy. A companion R package, outference, implements these new procedures with an interface that matches the functions commonly used for inference with lm in R.


Selective Confidence Intervals for Martingale Regression Model

In this paper we consider the problem of constructing confidence interva...

Selective Inference for Sparse Multitask Regression with Applications in Neuroimaging

Multi-task learning is frequently used to model a set of related respons...

Explainable outlier detection through decision tree conditioning

This work describes an outlier detection procedure (named "OutlierTree")...

Conditional Selective Inference for Robust Regression and Outlier Detection using Piecewise-Linear Homotopy Continuation

In practical data analysis under noisy environment, it is common to firs...

Shapley value confidence intervals for variable selection in regression models

Multiple linear regression is a commonly used inferential and predictive...

A Simple Model for Subject Behavior in Subjective Experiments

In a subjective experiment to evaluate the perceptual audiovisual qualit...

Analytical method for detecting outlier evaluators

Epidemiologic and medical studies often rely on evaluators to obtain mea...

Please sign up or login with your details

Forgot password? Click here to reset