An Encoder-only Transformer With Features - As-tokens For Robust Network Intrusion Detection
Abstract
Existing Transformer-based Network Intrusion Detection Systems (NIDS) exhibit three
interrelated weaknesses. First, most architectures process flow features as flat vectors,
preventing the model from discovering pairwise inter-feature dependencies that are informative
for distinguishing attack patterns. Second, a persistent methodological problem across the
NIDS literature is data leakage: when preprocessing transformations are fitted on the full
dataset before partitioning, test-set statistics contaminate the training pipeline, producing
inflated and non-reproducible results. Third, no prior work has jointly addressed these two
problems in a single lightweight architecture while also providing a comprehensive inference-
efficiency and deployment-feasibility analysis. This thesis addresses all three weaknesses
through one unified contribution: an encoder-only Transformer with features-as-tokens —
which treats each flow-level numerical feature as an individual token so that self-attention can
discover inter-feature dependencies — evaluated under a strictly leak-free protocol on the
CICIDS2017 benchmark, with a full ablation study and inference-latency comparison against
tree-based baselines.
The proposed architecture tokenises each scalar feature value into a d-dimensional
embedding, prepends a learned column embedding, and applies a feature gate that suppresses
uninformative inputs prior to tokenisation. Four pre-normalised encoder blocks with multi-
head self-attention and feed-forward sub-layers process the token sequence, and an attention-
pooling head aggregates the sequence into a single classification vector. To eliminate leakage,
all preprocessing operators — including constant-feature removal, class-conditional
imputation, outlier clipping, and quantile normalisation — are fitted exclusively on the training
partition and applied without refitting to validation and test sets. The resulting model contains
only 355,270 trainable parameters.
On the binary intrusion detection task (Normal vs. Attack), the Transformer achieves an
F1-score of 0.9945, AUC-PR of 0.99947, and accuracy of 99.78%, confirming that the features-
as-tokens paradigm is highly effective for binary flow classification. On the 15-class multi-
class task, it achieves 99.72% accuracy and macro F1 of 0.7791; performance degrades only
on attack categories with fewer than 10 training samples (Web Attack SQL Injection and Web
Attack XSS), where insufficient discriminative signal prevents reliable detection. A four-model
equal-weight ensemble — combining the Transformer's softmax outputs with those of Random
Forest, XGBoost, and LightGBM — recovers macro F1 to 0.8848 by compensating for these
rare-class failures. The standalone Transformer achieves an inference latency of 0.0676 ms per
flow (approximately 14,800 flows per second), making it 1.6–7.8× faster than all tree-based
baselines while maintaining competitive accuracy. In this thesis, the term “robust” refers
specifically to three properties of the proposed system: reliable evaluation under a strictly leak-
free protocol that prevents artificially inflated metrics, stable detection performance across
most traffic categories despite severe class imbalance, and deployment-ready inference speed
suitable for real-time high-throughput network environments.
Collections
- Industrial Engineering [3109]
