Spurious Features Everywhere -- Large-Scale Detection of Harmful Spurious Features in ImageNet

IEEE International Conference on Computer Vision (ICCV), 2022

9 December 2022

Matthias Hein

ArXiv (abs)PDF HTML Github (31★)

Main:9 Pages

32 Figures

Bibliography:3 Pages

2 Tables

Appendix:24 Pages

Abstract

Benchmark performance of deep learning classifiers alone is not a reliable predictor for the performance of a deployed model. In particular, if the image classifier has picked up spurious features in the training data, its predictions can fail in unexpected ways. In this paper, we develop a framework that allows us to systematically identify spurious features in large datasets like ImageNet. It is based on our neural PCA components and their visualization. Previous work on spurious features of image classifiers often operates in toy settings or requires costly pixel-wise annotations. In contrast, we validate our results by checking that presence of the harmful spurious feature of a class is sufficient to trigger the prediction of that class. We introduce a novel dataset "Spurious ImageNet" and check how much existing classifiers rely on spurious features.

View on arXiv

Comments on this paper