PhishSim: Aiding Phishing Website Detection with a Feature-Free Tool

07/13/2022
by   Rizka Purwanto, et al.
0

In this paper, we propose a feature-free method for detecting phishing websites using the Normalized Compression Distance (NCD), a parameter-free similarity measure which computes the similarity of two websites by compressing them, thus eliminating the need to perform any feature extraction. It also removes any dependence on a specific set of website features. This method examines the HTML of webpages and computes their similarity with known phishing websites, in order to classify them. We use the Furthest Point First algorithm to perform phishing prototype extractions, in order to select instances that are representative of a cluster of phishing webpages. We also introduce the use of an incremental learning algorithm as a framework for continuous and adaptive detection without extracting new features when concept drift occurs. On a large dataset, our proposed method significantly outperforms previous methods in detecting phishing websites, with an AUC score of 98.68 rate (TPR) of around 90 0.58 data in the future, and is feasible to deploy in real systems with a processing time of roughly 0.3 seconds.

READ FULL TEXT
research
07/22/2020

PhishZip: A New Compression-based Algorithm for Detecting Phishing Websites

Phishing has grown significantly in the past few years and is predicted ...
research
09/01/2019

WhiteNet: Phishing Website Detection by Visual Whitelists

Phishing websites aiming at stealing users' information by claiming fake...
research
11/10/2017

Traffic Analysis with Deep Learning

Deep Neural Networks (DNN) has obtained enormous attention with its adva...
research
05/16/2023

A Review of Data-driven Approaches for Malicious Website Detection

The detection of malicious websites has become a critical issue in cyber...
research
11/05/2021

Phish What You Wish

IT professionals have no simple tool to create phishing websites and rai...
research
10/24/2020

Towards Benchmark Datasets for Machine Learning Based Website Phishing Detection: An experimental study

In this paper, we present a general scheme for building reproducible and...
research
09/12/2023

Cookiescanner: An Automated Tool for Detecting and Evaluating GDPR Consent Notices on Websites

The enforcement of the GDPR led to the widespread adoption of consent no...

Please sign up or login with your details

Forgot password? Click here to reset