Hybrid Deep Learning Architecture for Phishing Detection Using Transformer Encoders and GATv2

Authors

  • Raynald Sumarga Universitas Negeri Surabaya
  • Ervin Yohannes Universitas Negeri Surabaya

Abstract

Phishing attacks are becoming increasingly evasive, exploiting valid HTTPS certificates, disguised domains, and web pages that are nearly indistinguishable from legitimate ones, which limits the effectiveness of traditional blacklist- and rule-based detection. This study proposes a hybrid deep learning architecture that represents a URL and its corresponding HTML content within a single graph rather than processing them separately. URL components (domain, subdomain, path, query) and HTML elements are each embedded through dedicated Transformer Encoders, connected via a virtual [DOC] node, and propagated through a three-layer residual GATv2 network; the resulting graph representation is further enriched with eleven handcrafted URL features and nine handcrafted HTML features designed to capture typosquatting and domain-generation patterns. The model is evaluated on a roughly 3,000-sample subset of the PhreshPhish dataset across three train-validation-test split scenarios (90:10, 80:20, 70:30), with an identical, fixed 1,000-sample test set across all three. The model attains its best result under the 90:10 split, with 89.60% accuracy, 89.78% F1-score, and 95.42% AUC-ROC, and remains stable across all three splits (accuracy 85.80%-89.60%, AUC-ROC consistently above 93%), although the 70:30 split stands out with a markedly different error profile (precision 92.20%, recall 80.40%) compared to the other two splits. These results show that implicit, graph-level fusion of URL and HTML modalities is a viable direction for phishing detection, though further refinement is needed to improve consistency across data splits.

 

Keywords— phishing URL detection, graph attention network, transformer encoder, multimodal fusion, hybrid deep learning, feature engineering

Downloads

Download data is not yet available.

Published

2026-09-25

Issue

Section

Articles
Abstract views: 1