Wrapper-Based Adversarial Input Screening for Deep Image Classifiers Using Feature Squeezing and Logit-Space Inconsistency
view PDF
view PDF

How to Cite

Hyso, Alketa, and Dezdemona Gjylapi. 2026. “Wrapper-Based Adversarial Input Screening for Deep Image Classifiers Using Feature Squeezing and Logit-Space Inconsistency”. Journal of Innovative Image Processing 8 (3): 870-93. https://doi.org/10.36548/jiip.2026.3.007.

Keywords

Adversarial Example Detection
Feature Squeezing
Logit-Space Inconsistency
Wrapper-Based Screening
Image Classifier Security

Abstract

Adversarial perturbations pose a practical integrity risk to vision-based decision systems by causing image classifiers to misclassify inputs after visually subtle changes. This paper evaluates a wrapper-based adversarial input screening approach that compares a classifier’s output on an original image with its output after benign feature-squeezing transformations. The evaluated transformations include median filtering, non-local means denoising, and bit-depth reduction. Using a frozen-backbone ResNet50 on CIFAR-10, the detector is assessed under untargeted and targeted Fast Gradient Sign Method attacks, followed by stronger Projected Gradient Descent verification. Detection thresholds are calibrated only on clean data using a fixed 5% false positive rate protocol and are validated on held-out clean samples. The results show that softmax-space ℓ1 inconsistency provides moderate, transformation-dependent detection, whereas logit-space ℓ2 inconsistency yields a stronger, more stable screening signal. Median filtering with logit-space ℓ2 achieves near-complete detection at ε = 8/255 and remains the most reliable configuration across the perturbation sensitivity analysis, while non-local means denoising becomes effective mainly for large perturbations. An additional two-stage gate is evaluated for deployment-oriented alert triage; it does not replace the primary detector or override its screening decision, but ranks flagged inputs by risk severity. Further evaluations show that performance decreases on TinyImageNet and that targeted threshold-aware adaptive optimisation can substantially reduce detection recall. The findings support the use of median filtering with logit-space ℓ2 inconsistency as a tool for screening adversarial inputs to image classifiers, but its effectiveness depends on dataset complexity, classifier behaviour, and attack adaptivity.

References

  1. Ren, Kui, Tianhang Zheng, Zhan Qin, and Xue Liu. "Adversarial Attacks and Defenses In Deep Learning." Engineering 6, no. 3 (2020): 346-360.
  2. Huang, Bo, Yi Wang, and Wei Wang. "Model-Agnostic Adversarial Detection by Random Perturbations." In IJCAI, pp. 4689-4696. 2019.
  3. Sharif, Mahmood, Sruti Bhagavatula, Lujo Bauer, and Michael K. Reiter. "Accessorize to A Crime: Real and Stealthy Attacks on State-Of-The-Art Face Recognition." In Proceedings of the 2016 acm sigsac conference on computer and communications security, pp. 1528-1540. 2016.
  4. Eykholt, Kevin, Ivan Evtimov, Earlence Fernandes, Bo Li, Amir Rahmati, Chaowei Xiao, Atul Prakash, Tadayoshi Kohno, and Dawn Song. "Robust Physical-World Attacks on Deep Learning Visual Classification." In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1625-1634. 2018.
  5. Deng, Zhijie, Xiao Yang, Shizhen Xu, Hang Su, and Jun Zhu. "Libre: A Practical Bayesian Approach to Adversarial Detection." In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 972-982. 2021.
  6. Dong, Yinpeng, Chang Liu, Wenzhao Xiang, Hang Su, and Jun Zhu. "Competition on Robust Deep Learning." National Science Review 10, no. 6 (2023): nwad087.
  7. Xu, Weilin, David Evans, and Yanjun Qi. "Feature Squeezing: Detecting Adversarial Examples in Deep Neural Networks." arXiv preprint arXiv:1704.01155 (2017).
  8. Tian, Jinyu, Jiantao Zhou, Yuanman Li, and Jia Duan. "Detecting Adversarial Examples from Sensitivity Inconsistency of Spatial-Transform Domain." In Proceedings of the AAAI conference on artificial intelligence, vol. 35, no. 11, pp. 9877-9885. 2021.
  9. Aldahdooh, Ahmed, Wassim Hamidouche, Sid Ahmed Fezza, and Olivier Déforges. "Adversarial Example Detection for DNN Models: A Review and Experimental Comparison." Artificial Intelligence Review 55, no. 6 (2022): 4403-4462.
  10. Kumar, Siddheshwar, Shashank Srivastava, and Shashwati Banerjea. "Fortifying Vision Models: A Comprehensive Survey of Defences Against Adversarial Examples." Applied Soft Computing (2025): 113874.
  11. Vassilev, Apostol, Alina Oprea, Alie Fordyce, and Hyrum Andersen. "Adversarial Machine Learning: A Taxonomy and Terminology of Attacks And Mitigations." (2024).
  12. Macas, Mayra, Chunming Wu, and Walter Fuertes. "Adversarial Examples: A Survey of Attacks and Defenses in Deep Learning-Enabled Cybersecurity Systems." Expert Systems with Applications 238 (2024): 122223.
  13. Gong, Yuxin, Shen Wang, Xunzhi Jiang, Liyao Yin, and Fanghui Sun. "Adversarial Example Detection Using Semantic Graph Matching." Applied Soft Computing 141 (2023): 110317.
  14. Priya, Ch. E. N. Sai, and Manas Kumar Yogi. "Trustworthy AI Principles to Face Adversarial Machine Learning: A Novel Study." Journal of Artificial Intelligence and Capsule Networks 5, no. 3 (2023): 227-245. https://doi.org/10.36548/jaicn.2023.3.002.
  15. Nartker, Makaela, Zhenglong Zhou, and Chaz Firestone. "When Will AI Misclassify? Intuiting Failures on Natural Images." Journal of vision 23, no. 4 (2023): 4-4.
  16. Sitawarin, Chawin, Arvind Sridhar, and David Wagner. "Improving the Accuracy-Robustness Trade-Off for Dual-Domain Adversarial Training." In ICML 2021 Workshop on Uncertainty and Robust-ness in Deep Learning, vol. 3. 2021.
  17. Madry, Aleksander, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. "Towards Deep Learning Models Resistant to Adversarial Attacks." arXiv preprint arXiv:1706.06083 (2017).
  18. Tack, Jihoon, Sihyun Yu, Jongheon Jeong, Minseon Kim, Sung Ju Hwang, and Jinwoo Shin. "Consistency Regularization for Adversarial Robustness." In Proceedings of the AAAI conference on artificial intelligence, vol. 36, no. 8, pp. 8414-8422. 2022.
  19. Zhao, Shiji, Xizhe Wang, and Xingxing Wei. "Mitigating Accuracy-Robustness Trade-Off via Balanced Multi-Teacher Adversarial Distillation." IEEE transactions on pattern analysis and machine intelligence 46, no. 12 (2024): 9338-9352.
  20. Koenig, Matthias, Annelot W. Bosman, Holger H. Hoos, and Jan N. Van Rijn. "Critically Assessing the State of the Art in Neural Network Verification." Journal of Machine Learning Research 25, no. 12 (2024): 1-53.
  21. B., Shuriya, Kowsalya S., Varatharajan N., and Sivaraju S.S. 2025. “FLIDS: Fuzzy Logic-Based Framework for Interpretable Image Manipulation Detection”. Journal of Trends in Computer Science and Smart Technology 7 (3): 312-330. https://doi.org/10.36548/jtcsst.2025.3.002.
  22. Meng, Dongyu, and Hao Chen. "Magnet: A Two-Pronged Defense Against Adversarial Examples." In Proceedings of the 2017 ACM SIGSAC conference on computer and communications security,135-147. 2017.
  23. Carlini, Nicholas, and David Wagner. "Adversarial Examples Are Not Easily Detected: Bypassing Ten Detection Methods." In Proceedings of the 10th ACM workshop on artificial intelligence and security,3-14. 2017.
  24. Krizhevsky, Alex, Vinod Nair, and Geoffrey Hinton. "CIFAR-10 (Python version)." Data set. University of Toronto, 2009. https://www.cs.toronto.edu/~kriz/cifar.html.