Outlier Detection with Globally Optimal Exemplar-Based GMM

Yang, Xing‐Wei; Latecki, Longin Jan; Pokrajac, Dragoljub

doi:10.1137/1.9781611972795.13

Cited by 84 publications

(47 citation statements)

References 29 publications

(47 reference statements)

Supporting

Mentioning

Contrasting

Unclassified

Order By: Relevance

“…K (x m , x n ) = exp(∥x m − x n ∥ 2 /h), and the parameter h and the ridge parameter γ were set at 100 and 0.1, respectively. To measure the performance of each algorithm, we utilized the Detection rate (True positive rate), the False alarm rate (False positive rate) from the work 17 , and the ROC curves, which are defined as follows:…”

Section: Resultsmentioning

confidence: 99%

“…Along this line, He et al 16 proposed FindCBLOF to determine the Cluster-Based Local Outlier Factor (CBLOF) for each data point. Yang et al 17 had introduced a globally optimal exemplarbased GMM to detect the outliers, in which a Gaussian is centered at each point. The outlier factor at each point is calculated by the weighted sum of the mixture proportion with the weights representing the similarities to the other points.…”

Section: Introductionmentioning

confidence: 99%

See 1 more Smart Citation

Outlier Detection Based on Local Kernel Regression for Instance Selection

Peng¹,

Cheung

2014

IJCIS

View full text Add to dashboard Cite

show abstract

Section: Resultsmentioning

confidence: 99%

Section: Introductionmentioning

confidence: 99%

Outlier Detection Based on Local Kernel Regression for Instance Selection

Peng¹,

Cheung

2014

IJCIS

View full text Add to dashboard Cite

show abstract

“…The most straightforward outlier detection method, modelbased method, is to create a model for all samples, and then predict outliers as those having large deviations from the established profiles. For example, the Gaussian mixture model (GMM) [11] fits the whole dataset to a mixed Gaussian distribution and evaluates the parameters through the Expectation-Maximization [29] or a deep estimation network [12]. However, GMM needs to predetermine the appropriate cluster type and number, which are crucial and extremely difficult.…”

Section: Classic Outlier Detection Methodsmentioning

confidence: 99%

“…We compare MO-GAAL with nine representative outlier detection algorithms. They can be divided into seven categories: (i) two density-based methods, LOF [18] and KDEOS [38]; (ii) two density estimators, GMM [11] and Parzen [17]; (iii) a typical distance-based approach, kNN [37]; (iv) an angle-based model, FastABOD [38]; (v) a clusterbased model, k-means; (vi) a popular one-class classification model, OC-SVM [35] and (vii) the Active-Outlier detection model, AO [24]. In addition, AGPO and SO-GAAL are also compared on real-world datasets to demonstrate the necessity of using multiple generators with different objectives.…”

Section: Evaluation Measuresmentioning

confidence: 99%

“…The most straightforward way is to create a model for all samples, and then compute the outlier scores based on the deviations from the established normal profiles. Specific methods include the statistical-based models [11], [12], regression-based models [13], cluster-based models [14], and reconstruction-based models [15], [16], which make different assumptions about the generating mechanism of Xiangnan He is the corresponding author. normal data.…”

Section: Introductionmentioning

confidence: 99%

See 1 more Smart Citation

Generative Adversarial Active Learning for Unsupervised Outlier Detection

Liu

Zhou

et al. 2019

IEEE Trans. Knowl. Data Eng.

203

129

View full text Add to dashboard Cite

Outlier detection is an important topic in machine learning and has been used in a wide range of applications. In this paper, we approach outlier detection as a binary-classification issue by sampling potential outliers from a uniform reference distribution. However, due to the sparsity of data in high-dimensional space, a limited number of potential outliers may fail to provide sufficient information to assist the classifier in describing a boundary that can separate outliers from normal data effectively. To address this, we propose a novel Single-Objective Generative Adversarial Active Learning (SO-GAAL) method for outlier detection, which can directly generate informative potential outliers based on the mini-max game between a generator and a discriminator. Moreover, to prevent the generator from falling into the mode collapsing problem, the stop node of training should be determined when SO-GAAL is able to provide sufficient information. But without any prior information, it is extremely difficult for SO-GAAL. Therefore, we expand the network structure of SO-GAAL from a single generator to multiple generators with different objectives (MO-GAAL), which can generate a reasonable reference distribution for the whole dataset. We empirically compare the proposed approach with several state-of-the-art outlier detection methods on both synthetic and real-world datasets. The results show that MO-GAAL outperforms its competitors in the majority of cases, especially for datasets with various cluster types or high irrelevant variable ratio. The experiment codes are available at: https://github.com/leibinghe/GAAL-based-outlier-detection

show abstract