Anomaly-based intrusion detection systems (IDS) are commonly evaluated under the assumption that network traffic remains stationary over time. However, in real-world deployments, they are constantly faced with concept drift. This drift may make the overall performance of the detectors, e.g., the F1-score, look fine and mask substantial variations in detector behaviour. The goal of this study is not to conclude if concept drift affects overall performance, but to describe how it affects the anomaly detection performance over time and what are the different failure modes of the behaviour. For this purpose, three datasets for intrusion detection, namely CIC-IoT-2023, Edge-industrial IoT (Edge-IIoTset), and UNSW-NB15, with different class distributions are used and evaluated with a common temporal evaluation framework. The results show that concept drift doesn't lead to a consistent decrease in performance. In contrast, it results in specific behaviour changes, such as total blindness, sensitivity inflation and relatively stable detection. In all experimental conditions, the isolation forest (IF) was the most unstable in terms of behaviours. Its recall was still close to zero on Edge-IIoTset dataset while Gaussian mixture model (GMM) and autoencoder (AE) were able to detect the attacks with near-perfect accuracy in the attack-contaminated segments. The results indicate that the overall performance of an IDS may mask failure modes that are operationally relevant and that a time-based behavioural assessment more realistically assesses IDS reliability in a dynamic cyber security context.
📄 Full text (101,845 characters)extracted from the PDF · click to expand
Behavioral instability in anomaly-based intrusion detection systems under concept drift: from blindness to sensitivity inflation
Mohammad M. Rasheed1, Mustafa Muwafak Alobaedy2*
College of Engineering, University of Information Technology and Communications, Baghdad, 10013, Iraq1
Centre for Image and Vision Computing, Multimedia University, Persiaran Multimedia, Cyberjaya, 63100, Selangor, Malaysia2
Received: 28-March-2026; Revised: 19-July-2026; Accepted: 21-July-2026
©2026 Mohammad M. Rasheed, Mustafa Muwafak Alobaedy. This is an open access article distributed under the Creative Commons Attribution (CC BY) License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
Abstract
Anomaly-based intrusion detection systems (IDS) are commonly evaluated under the assumption that network traffic remains stationary over time. However, in real-world deployments, they are constantly faced with concept drift. This drift may make the overall performance of the detectors, e.g., the F1-score, look fine and mask substantial variations in detector behaviour. The goal of this study is not to conclude if concept drift affects overall performance, but to describe how it affects the anomaly detection performance over time and what are the different failure modes of the behaviour. For this purpose, three datasets for intrusion detection, namely CIC-IoT-2023, Edge-industrial IoT (Edge-IIoTset), and UNSW-NB15, with different class distributions are used and evaluated with a common temporal evaluation framework. The results show that concept drift doesn't lead to a consistent decrease in performance. In contrast, it results in specific behaviour changes, such as total blindness, sensitivity inflation and relatively stable detection. In all experimental conditions, the isolation forest (IF) was the most unstable in terms of behaviours. Its recall was still close to zero on Edge-IIoTset dataset while Gaussian mixture model (GMM) and autoencoder (AE) were able to detect the attacks with near-perfect accuracy in the attack-contaminated segments. The results indicate that the overall performance of an IDS may mask failure modes that are operationally relevant and that a time-based behavioural assessment more realistically assesses IDS reliability in a dynamic cyber security context.
Keywords
Concept drift, Intrusion detection system (IDS), Anomaly-based intrusion detection, Isolation forest (IF), Temporal evaluation, Behavioral analysis.
1. Introduction
Intrusion detection systems (IDS) are an integral part of network security, as the internet of things (IoT) has significantly increased the attack surface of today's networks. Machine learning based IDS and anomaly detection-based IDS in particular, have demonstrated good performance in laboratory settings, often with F1-scores of more than 0.95 [1, 2]. Recent extensive reviews have shown that today deep learning and classical anomaly detectors have become the leading research areas for IoT security [2, 3]. However, these high numbers are increasingly recognized as potentially misleading because they are usually derived from the assumption that the statistical distribution of training and deployment data is the same, which is not the case in most operational scenarios [4, 5].
Change in the statistical properties of data over time, called concept drift, is thus one of the fundamental challenges for deployed detectors [6]. The drift is the result of the changing of attack techniques, the evolution of legitimate behaviour, infrastructure changes, and changes in the composition of traffic in the case of network intrusion detection [3, 7]. The impact of drift on security classifiers has been reported in several areas of security including malware detection [8, 9] and network intrusion detection [10, 11] and systems like Transcend [8] and TESSERACT [9] have demonstrated that temporally unaware evaluation greatly overestimates performance. There are many precedents for behavioural detection, a technique that characterizes malicious activity based on its dynamic patterns instead of static signatures, in network defence, including early work on behavioural scanning-worm detection [12] that inspired the current behaviour-centric notion of drift.
Despite the considerable research on drift detection [13, 14] and adaptation via online learning and ensemble retraining [15, 16], a basic lack of understanding exists regarding the effects of concept drift on the operational behaviour of IDS. Previous work mostly focuses on the amount of performance loss, and does not often describe what kind of failure modes result from drifting. A detector which becomes gradually less accurate is very different from a detector which suddenly ceases to detect all attacks, and both differ from one which suddenly raises alarms indiscriminately, but aggregate metrics, like F1-score, hide these differences [17]. This paper directly tackles this issue by analysing the particular failure modes of anomaly-based detectors under drift. Although machine learning methods have been successfully applied to a wide range of sensing and security applications ranging from indoor localisation [18] to intrusion detection, their temporal stability under distribution shift is relatively under-explored.
The goal of this study is not only quantifying the effect of concept drift, but also at describing the change in the anomaly-based IDS decision behavior over time. Specifically, it aims to identify common behavioral failure patterns, determine whether vulnerability to concept drift is associated with the underlying detection paradigm or specific algorithms, and provide a temporal evaluation framework that reveals failure modes not captured by aggregate performance metrics.
Three anomaly detection algorithms from two detection paradigms (isolation-based, and density/reconstruction-based) are applied to three benchmark datasets with different distributional characteristics (CIC-IoT-2023, Edge-IIoTset, and UNSW-NB15), and baseline (α = 0) experiments on all three datasets are used to distinguish injected drift from natural data structure. The four contributions are as follows. (i) Empirical behavioural characterisation: three behaviour patterns emerged: immediate blindness, sensitivity inflation and drift-resistant detection. (ii) Evidence of paradigm-level tendency: the isolation forest (IF; isolation-based) showed behavioural instability (sensitivity inflation on CIC-IoT-2023, UNSW-NB15 and immediate blindness on Edge-IIoTset), while the Gaussian mixture model (GMM; density-based) and the autoencoder (AE; reconstruction-based) remained stable in all environments. (iii) Evidence of the same environment (Edge-IIoTset) yielding diametrically opposite behaviour for different paradigms, confirmed by multi-seed replication with confidence intervals (CI). (iv) Operationalisation of a temporal discrimination score (recall − false positive rate (FPR), equivalent to Youden's J statistic), and behavioural stability indices, as deployment-oriented evaluation aids, complemented by a robustness analysis covering segment count, threshold, drift type, pre-processing, hyperparameters and computational cost.
This paper does not attempt to provide a universal taxonomy of anomaly-detector behaviour under concept drift, nor does it attempt to provide an empirical study of behaviour for every detector in every paradigm in every benchmark environment. Rather, it provides an empirical behavioural analysis of three detectors from two paradigms in three benchmark environments, and argues that aggregate metrics alone can mask practically important failure modes.
The rest of this paper is structured as follows: In Section 2, related works about concept drift in security applications and temporal evaluation are reviewed. The methodology is presented in Section 3, which consists of datasets, algorithms, evaluation framework, stability metrics and robustness-analysis setup. The experimental results are reported in Section 4, along with a robustness and sensitivity analysis. The findings, implications, comparative analysis and limitations of the study are discussed in Section 5. The conclusions and future work are outlined in Section 6.
2. Literature review
This section summarises some of the latest literature pertaining to concept drift and anomaly detection in IDS. A series of studies from 2024–2026 is summarised, each according to its methodology, principal results, advantages, and limitations.
2.1 Concept drift in anomaly-based detection
Chu et al. [14] tackled the concept drift problem in IoT data streams by combining an ensemble of non-parametric tests (Kolmogorov–Smirnov, Wilcoxon rank-sum and Mann–Kendall) with an isolation-forest detector, using a sliding-window distribution comparison for drift localisation. The approach achieved accurate drift-point detection on synthetic streams and Edge-IIoTset, with the benefit of label-free, easily interpretable drift points; its limitation is that it focuses on drift-point localisation rather than the behavioural degradation of the detector itself.
Xu et al. [16] compared batch and streaming learners for binary anomaly detection in IoT traffic under drift, using heterogeneous streams synthesised by mixing datasets and replayed sample-by-sample. They found that batch random forest failed under drift, whereas adaptive random forest achieved an F1-score of 0.990 ± 0.006 at one-third the cost. A realistic streaming protocol with cost accounting is an advantage, but the study is restricted to tree-based supervised streaming models and includes no unsupervised paradigm analysis.
Mallidi and Ramisetty [3] conducted a systematic, PRISMA-style review of machine learning/deep learning based IoT-IDS design centred on feature selection and data balancing. They concluded that these two factors dominate reported performance. The review is comprehensive and up-to-date, but it only discusses, without empirically evaluating, temporal robustness against drift.
2.2 Temporal evaluation methodologies
Pendlebury et al. [9] proposed TESSERACT to reveal temporal and spatial experimental bias through time-aware splitting based on the area-under-time metric, showing that realistic time-ordered evaluation significantly reduces apparent performance. The protocol is reusable and constraint-based, but it was framed for supervised malware detection rather than unsupervised network-anomaly behaviour. Barbero et al. [19] later improved this line with TRANSCENDENT, introducing conformal-evaluation guarantees for drifting malware classifiers.
Mohale and Obagbuwa et al. [17] evaluated explainable-artificial intelligence pipelines for machine-learning-based IDS across various feature-selection and model configurations, using black-box explanations for benchmark traffic. They reported better interpretability with maintained accuracy, while confirming the sensitivity of metrics to evaluation design. The work offers analyst trust and transparency, but it provides only a snapshot evaluation and does not track behaviour over time under drift.
Yang et al. [11] presented contrastive AE for drifting detection and explanation, which combines contrastive learning with distance-based scoring (measured by the distance to known classes in the latent space) to detect and explain drifting samples. It achieved effective out-of-distribution drift detection together with explanations, offering actionable interpretability; however, it targets drift detection rather than detector failure-mode characterisation.
Arp et al. [5] catalogued methodological pitfalls of machine learning in security, such as temporal snooping and base-rate neglect, through a systematic survey with reproductions, and identified ten common pitfalls together with their solutions. The guidance is broad and field-shaping, but it is prescriptive rather than an empirical behavioural study under drift.
2.3 Hybrid detectors and drift adaptation
Work over the last few years (2025–2026) has increasingly focused on the reconstruction, isolation and density aspects together. Haque and Soliman [20] combined a transformer AE with IF and XGBoost, using reconstruction error with adaptive thresholding validated by IF, and reported 95% accuracy with high intrusion recall. The hybrid offers strength and adaptive sensitivity, but it is snapshot-only with no temporal drift analysis. Many such systems build on classical drift detectors, for example ADWIN [21], which established a baseline using an adaptive-windowing approach.
Bagui et al. [22] proposed an IF drift monitor to trigger Random-Forest retraining on traffic big data, using distribution monitoring with scheduled retraining, and reported constant accuracy under changing traffic conditions. The benefit is an end-to-end working process; the limitation is its dependence on continual retraining and labels, the very overhead that this static-calibration study aims to characterise. A 2026 AE–GMM hybrid [23] combines reconstruction with density estimation, while Bachar et al. [24] combine AE with IF for IoT anomalies; both achieve high accuracy but acknowledge that AE can learn dataset-specific regularities that may hinder robustness when the distribution changes. These limitations are further supported by earlier drift studies that placed AE on top of supervised nets but needed periodic retraining [10], showed that drift handling depends heavily on diversity assumptions [25], or advanced stream- and online-learning IDS but tested adaptation speed instead of the failure modes of static detectors [26, 27].
Unlike these adaptation- and accuracy-oriented studies, the present work does not propose any new detector nor any new adaptation mechanism. It rather describes the evolution of the decision behaviour of pre-existing unsupervised detectors under drift, across paradigms and environments, under a deliberately static calibration. There are three gaps that appear throughout the literature surveyed. First, most studies focus on drift adaptation; however, they only report the overall accuracy or F1-score of static detectors and do not characterise their failure modes. Second, hybrid AE/IF/GMM systems are tested on a single snapshot, so it is not known how they behave under sustained drift. Third, when drift is studied, it is typically one type of drift on one dataset, and rarely includes CI, sensitivity analyses, or cross-paradigm, cross-environment comparisons. The present study aims to fill these gaps with a paradigm-spanning, multi-dataset, unified behavioural analysis including explicit robustness and uncertainty quantification.
3. Methodology
The experimental methodology in this study is presented in this section. The general design of the pipeline is introduced first, followed by the introduction of the datasets, the anomaly detection algorithms, the temporal segmentation and drift injection algorithms, the adaptive evaluation framework, and the proposed behavioural stability metrics.
3.1 Pipeline overview
An integrated experimental pipeline was developed to identify the influence of environmental change on detector behaviour. The important design principles are as follows. First, all detection algorithms are only trained with early benign traffic, so that they learn a model of normal behaviour and are not exposed to attack patterns. Second, identical preprocessing, feature selection, and calibration of threshold values is used for all the datasets and all algorithms, which allows direct comparison across environments and algorithms. Third, a fixed decision threshold (99th percentile of validation scores) is used throughout the evaluation, making no effort to use adaptive thresholding, to ensure that behavioural changes can be investigated in a controlled static calibration environment. This design choice enhances interpretability; however, it also implies that some observed failure modes should be interpreted as the combined effect of distributional shift and non-adaptive thresholding rather than being attributed solely to the underlying model. Figure 1 presents a block diagram of the complete experimental workflow.
Figure 1: Block diagram of the system architecture and working process, showing the data, processing, and output layers and the interactions between the three anomaly detectors and the behavioural stability indices.
The data processing pipeline consists of the following steps: (1) data loading and stratified sampling; (2) temporal ordering and partitioning of the dataset into 20 equal-sized segments; (3) optional drift injection by progressively adding Gaussian noise to the segments after the midpoint; (4) extraction of numerical features followed by feature standardization; (5) training the anomaly detection model using benign samples from the early segments; (6) threshold calibration on a held-out benign validation set; (7) temporal evaluation of all 20 segments using the proposed behavioural assessment framework; and (8) computation of the behavioural stability indices. For the IF and GMM, anomaly scores are obtained using the models' native scoring functions (i.e., the decision function and negative log-likelihood, respectively). For the AE, anomaly scores are computed as the mean squared reconstruction error. The overall process is outlined in Algorithm 1. The symbols in the algorithm are defined as follows: D is the dataset; K=20 is the total number of temporal segments; K_train = 8 is the number of initial temporal segments used for training the model; α is the drift-strength factor (α=0 for the baseline experiments and α=0.5 for the primary experiments); and the random seed is set to 42 for reproducibility.
For each temporal segment, the segment type is determined based on its class composition (mixed, attack-only, or benign-only). Subsequently, recall, FPR, F1-score, and detection stability (DS) are computed whenever applicable. Finally, the behavioural stability indices, namely the blindness index (BI), alarm fatigue index (AFI), stability variance (SV), and mean DS, are calculated.
Algorithm 1: Temporal Behavioural Evaluation Pipeline
Input: Dataset D = {(x_i, y_i)}_{i=1}^N, number of segments K = 20, training segments K_train = 8, drift strength α = 0.5, threshold percentile q = 0.99, random seed = 42
Output: Behavioural metrics {(Recall_k, FPR_k, F1_k, DS_k)}_{k=0}^{K-1}, stability indices (BI, AFI, SV, DS_bar)
Temporal Segmentation: Partition D into K ordered segments {S_0, S_1, ..., S_{K-1}} of equal size n = floor(N/K)
Drift Injection: For each segment S_k where k > K/2:
For each sample x in S_k: x_tilde_j = x_j + eta_j, where eta_j ~ N(0, (α · (k - K/2)/(K/2) · sigma_j^train)^2)
Feature Extraction: Extract numeric features, compute standardization parameters (mu^train, sigma^train) from benign samples of {S_0, ..., S_{K_train - 1}}
Training Data Preparation:
B = {x_i : y_i = benign, x_i in {S_0, ..., S_{K_train - 1}}}
Split B into B_train (80%) and B_val (20%)
Standardize: x_j = (x_j - mu_j^train) / sigma_j^train for all samples
Model Training: Fit anomaly detector M on B_train
Threshold Calibration: tau = Q_q({A_M(x) : x in B_val})
Temporal Evaluation: For each segment k = 0, 1, ..., K-1:
Compute predictions: y_hat_i = 1[A_M(x_i) > tau] for all x_i in S_k
Determine segment type based on class composition:
If mixed: compute Recall_k, FPR_k, F1_k, DS_k
If attack-only: compute Recall_k only
If benign-only: compute FPR_k only
Stability Indices: Compute BI, AFI, SV, DS_bar as defined in Section 3.7
Return {(Recall_k, FPR_k, F1_k, DS_k)}, (BI, AFI, SV, DS_bar)
3.2 Datasets
Three benchmark datasets with fundamentally different distributional characteristics were chosen to maximise the variety of the observable behavioural patterns.
CIC-IoT-2023 [28] is a full IoT network traffic dataset created by the Canadian Institute for Cybersecurity. 33 attack types within 7 categories collected from a testbed of 105 IoT devices. The dataset contains 46 pre-extracted network flow features. After the stratified sampling and the temporal ordering, about 200,000 samples remained, grouped into 20 segments, each containing 10,000 samples. The dataset has a steady high rate of attack (around 97%) in all segments of time, representing a situation in which the class distribution is stable while the distribution of features can change. This attack rate is significantly higher than in real-world network traffic; this discrepancy has some impact on the interpretability of aggregate metrics such as F1-score, as discussed in Section 5.3.
Edge-IIoTset [29] is a cybersecurity dataset for IoT and IIoT environments, generated from a multi-layer testbed architecture. The original dataset contains 1176 extracted features, from which the dataset authors selected 61 features after removing highly correlated ones. After preprocessing, about 182,000 samples were kept, and were grouped into 20 segments of about 9,123 samples each. This dataset has a high degree of temporal class imbalance: most segments (0-1 and 5-19) contain only attack traffic (TrueAttackRate nearly equal to 1.0), one segment (3) is benign-only (TrueAttackRate = 0.0), and only two segments (2 and 4) contain a mix of both classes (TrueAttackRate nearly equal to 0.25 and 0.83, respectively).
UNSW-NB15 [30] is a network intrusion dataset generated by the Australian Centre for Cyber Security with the IXIA PerfectStorm tool, to generate realistic modern network traffic. The dataset includes nine attack categories and normal traffic, and 49 features were extracted from network flow attributes. The number of samples after stratified sampling was about 62,780, with 20 segments containing about 3,139 samples each. The dataset has a moderate attack rate (around 68%), which is an intermediate situation between the extreme class imbalance of CIC-IoT-2023 and a balanced distribution.
For all three datasets all numeric columns were kept after removing the label column and any timestamp columns. Non-finite values (infinity) were replaced with NaN and remaining NaN values were imputed with zero. No feature selection or correlation filtering has been used as the goal was to assess the stability of behaviour without artificial optimisation. All features were normalised using the mean and the standard deviation (SD) calculated only from the benign training samples. The impact of correlation-based feature filtering and of the missing-value imputation strategy is examined separately as a robustness analysis in Section 4.5, where both are shown not to alter the reported behavioural patterns.
The contrasting characteristics of these three datasets, one with a stable high attack rate, one with extreme temporal class imbalance, and one with a moderate, realistic class balance, allow the study of behavioural instability under varying environmental conditions.
3.3 Anomaly detection algorithms
Three anomaly detection algorithms spanning two fundamentally different detection paradigms were selected: the isolation-based paradigm and the density/reconstruction-based paradigm. The aim of this study is behavioural characterisation across representative anomaly-detection paradigms rather than exhaustive algorithm benchmarking. Accordingly, well-established algorithms were chosen to represent each paradigm: the IF for the isolation-based paradigm, and the GMM and AE for the density- and reconstruction-based paradigm, respectively. Including two independent algorithms (GMM and AE) within the latter paradigm allows the analysis to provide preliminary evidence of paradigm-level tendencies rather than algorithm-specific artefacts: where both behave identically, the behaviour is more plausibly associated with the paradigm than with one algorithm. This design isolates the effect of the underlying detection principle, and the robustness analyses in Section 4.5 further confirm that the observed behaviour reflects the paradigm rather than a particular parameter choice.
The IF [31] finds the anomalies by measuring the ease with which a sample can be isolated using random recursive partitioning of the feature space. For a given sample, the anomaly score is based on the average path length over an ensemble of T isolation trees: the anomaly score is defined as in Equation 1, where the normalising factor is given by Equation 2.
s(x, n) = 2^(- E[h(x)] / c(n)) (1)
Where E[h(x)] denotes the expected path length of sample x averaged over all isolation trees, n is the number of training samples, and c(n) represents the average path length of an unsuccessful search in a binary search tree. Anomaly scores approaching 1.0 indicate anomalous samples, whereas scores close to 0.5 indicate normal samples. The IF model was implemented using the Scikit-learn IsolationForest class with 200 estimators (n_estimators (T) = 200), contamination = "auto", and random_state = 42. All remaining parameters were kept at their default values, including max_samples = "auto" and max_features = 1.0.
c(n) = 2H(n - 1) - 2(n - 1)/n (2)
The GMM [32] is employed to model the probability density of normal network traffic using a single Gaussian component with diagonal covariance (n_components = 1, covariance_type = 'diag', reg_covar = 1e-3). The diagonal covariance assumption treats the features as conditionally independent, providing computational efficiency; however, it does not capture correlations among features. The model estimates the parameters of the Gaussian distribution, namely the mean vector and the diagonal covariance matrix. The anomaly score for each sample is computed as its negative log-likelihood, as defined in Equation 3. A common detection threshold is then determined for all anomaly detectors using the threshold calibration procedure described in Section 3.5 (Equation 6).
A(x) = -log p(x | mu, Sigma) = d/2 · log(2*pi) + 1/2 · sum_{j=1}^{d} log(sigma_j^2) + 1/2 · sum_{j=1}^{d} (x_j - mu_j)^2 / sigma_j^2 (3)
Higher values of A(x) indicate that a sample lies in a low-density region of the learned distribution and is therefore more likely to be anomalous. To ensure numerical stability during covariance estimation, a covariance floor of epsilon = 10^{-3} is applied such that sigma_j^2 = max(sigma_j^2, epsilon).
The AE is an undercomplete neural network trained exclusively on benign samples to learn a compact representation of normal network behaviour. The architecture consists of an encoder with layers d → 64 → 32 and a symmetric decoder with layers 32 → 64 → d. Rectified linear unit activation functions are used in the hidden layers, whereas a linear activation function is employed in the output layer. The model is trained using the Adam optimizer and the mean squared reconstruction error loss function for 10 epochs with a batch size of 512.
The reconstructed output is expressed as shown in Equation 4.
x_hat = g(f(x)) (4)
where f(·) denotes the encoder function, g(·) denotes the decoder function, and x_hat represents the reconstructed input. The anomaly score is computed as the mean squared reconstruction error between the input and its reconstruction, as defined in Equation 5:
A_AE(x) = 1/d · sum_{j=1}^{d} (x_j - x_hat_j)^2 (5)
Samples that deviate substantially from the learned normal manifold produce larger reconstruction errors and are therefore classified as anomalies. The AE model was implemented using TensorFlow/Keras, with the random seed fixed at 42 to ensure reproducibility.
3.4 Temporal segmentation and drift injection
These three algorithms are two different paradigms of anomaly detection, IF and density/reconstruction-based detection (GMM, and AE). This distinction is crucial for the current study since it allows the study of whether vulnerability to concept drift is specific to a particular algorithm or is inherent to the detection paradigm.
The data sets were split into 20 segments of the same length, with the same temporal order as in the original data sets. This temporal segmentation allows to analyze the behaviour changes and the performance of the detectors in detail along the data stream.
A controlled drift injection strategy was used to study the impact of concept drift. After the middle of the data stream (Segment 10), Gaussian noise was added to the feature values, with the amplitude of the noise increasing over time. For a sample x in segment k (k > 10), the drifted feature value is defined as shown in Equation 6.
x_tilde_j = x_j + eta_j, eta_j ~ N(0, (α · (k - 10)/10 · sigma_j^train)^2) (6)
Where α is the drift-strength factor, and sigma_j^train denotes the SD of the jth feature computed from the benign training data. The maximum drift-strength factor for the primary experiments was 0.5. The injected noise is thus linearly ramped up from 0% in Segment 10 to 50% in the last segment of the training set, of each feature's standard deviation. This controlled perturbation mimics a concept drift over time that enables a systematic evaluation of the behaviour of the anomaly detectors under progressive concept drift. In this study the main drift generation mechanism is the injection of gaussian noise. In addition, to check the robustness and generalizability of the proposed behavioural analysis framework, three other drift scenarios are also tested: uniform drift, spike drift and covariate shift; and the results are presented in Section 4.5, which ensures that the results observed are not limited to Gaussian perturbations.
3.5 Training and threshold calibration
In order to differentiate the effect of concept drift injected in the data from the natural variations in the data distribution, a baseline condition (α = 0) was tested across all three datasets. This baseline allows a direct comparison of the detector behaviour with and without artificial drift and therefore the observed changes in behaviour to be attributed to the nature of the data or the injected perturbations.
Samples from the first eight temporal segments were also taken from the benign samples, to create the training set. The 80:20 split was taken with 80% of the samples being used for model training and 20% being used for threshold calibration. The numerical features were all z-scored using only the benign training data to avoid information leakage. Then, a common decision threshold was found for each anomaly detector based on the benign validation set, as described in Equation 7:
tau = Q_{0.99}({A(x_i)}_{i=1}^{N_v}) (7)
Where A(x_i) denotes the anomaly score of the ith benign validation sample, N_v is the number of validation samples, and Q_{0.99}(·) represents the 99th-percentile function. Consequently, a sample x is classified as anomalous if A(x) > tau.
Choosing the threshold at the 99th percentile allows to define a consistent false-positive baseline for all the anomaly detection algorithms and datasets, since about 1% of the benign samples are misclassified as anomalies. The calibrated threshold is not adapted during the testing and is kept constant over time. This means that degradation in detection performance observed can be due not only to the impact of concept drift on the learned model but also to the interplay between changing anomaly score distributions and the static anomaly score threshold calibration policy. This design is specifically chosen to separate the effect of concept drift in the presence of a fixed decision boundary, as it is a realistic deployment scenario where anomaly detectors are trained once and used without frequent re-training.
3.6 Adaptive evaluation framework
The experiments were repeated on the Edge-IIoTset dataset with five different random seeds (42, 123, 456, 789 and 2024) to evaluate the robustness of the reported results. This analysis is especially relevant for the IF algorithm, which includes stochastic operations, like random feature selection and random split-point generation. The results shown are the mean ± SD for the five experimental runs.
One of the main methodological difficulties in temporal evaluation is dealing with segments with extreme class imbalance. Traditional evaluation approaches like the F1-score depend on having both attack and benign samples, and are undefined or misleading if one class is missing. An adaptive temporal evaluation framework is used to overcome this limitation. The Recall and FPR are calculated using Equation 8 and Precision and F1-score are calculated using Equation 9.
Recall_k = TP_k / (TP_k + FN_k), FPR_k = FP_k / (FP_k + TN_k) (8)
Precision_k = TP_k / (TP_k + FP_k), F1_k = 2 * Precision_k * Recall_k / (Precision_k + Recall_k) (9)
where TP_k, FP_k, FN_k, and TN_k denote the numbers of true positives, false positives, false negatives, and true negatives, respectively, in the kth temporal segment.
For mixed segments, containing both attack and benign samples, recall, FPR, Precision, and F1-score are computed. For attack-only segments, where no benign samples are present, only recall is reported because FPR and precision are undefined. Likewise, for benign-only segments, only the FPR is computed because recall cannot be evaluated.
3.7 Behavioural monitoring metrics
In addition to the conventional performance metrics, three behavioural stability indices and one discrimination metric are employed to characterize detector behaviour over time. Let S_R denote the set of temporal segments in which recall is defined, and let S_F denote the set of segments in which the FPR can be computed.
The BI quantifies the detector's inability to identify malicious traffic and is defined as shown in Equation 10.
BI = 1 - (1/|S_R|) * sum_{k in S_R} Recall_k (10)
A value approaching 1 indicates that the detector misses most attacks, whereas 0 represents perfect attack detection. The AFI measures the average false-positive burden imposed on security analysts and is defined as shown in Equation 11.
AFI = (1/|S_F|) * sum_{k in S_F} FPR_k (11)
Higher AFI values indicate a larger proportion of false alarms, thereby increasing the operational burden in practical deployments.
The SV quantifies the temporal consistency of detection performance and is defined as shown in Equation 12.
SV = (1/|S_R|) * sum_{k in S_R} (Recall_k - Recall_bar)^2 (12)
where Recall_bar = (1/|S_R|) * sum_{k in S_R} Recall_k. In other words, low variance denotes consistent behaviour (whether consistently good or consistently poor), whereas high variance denotes behavioural instability over time.
DS measures the capability of the detector to differentiate attacks from normal traffic per segment. This metric is mathematically equivalent to Youden's J statistic [33], which was originally proposed for the evaluation of diagnostic tests (Equation 13).
DS_k = J_k = Recall_k - FPR_k = Sensitivity_k + Specificity_k - 1 (13)
While Youden's J has been used extensively in medical diagnostics and ROC analysis, its systematic use as a temporal monitoring metric for IDS behavioural analysis has not been previously explored. It is adopted here in particular because, unlike F1-score, it is not subject to the problem of class imbalance, which is a desirable property for temporal IDS evaluation where the class proportions may vary from one segment to another.
The aggregate DS is computed as shown in Equation 14:
DS_bar = (1/|S_M|) * sum_{k in S_M} DS_k (14)
Where S_M = S_R ∩ S_F is the set of mixed segments. A value of 1.0 means perfect discrimination, 0.0 means random-level performance, and negative values indicate systematic misclassification.
3.8 Robustness analysis
To verify that the observed behavioural patterns are not specific to the selected experimental configuration, an extensive robustness analysis comprising twelve controlled experiments was conducted.
The analyses include:
1. Segment-count sensitivity: K={10, 15, 20, 25, 30}
2. Threshold sensitivity (95th, 97th, 99th and 99.5th) percentile thresholds.
3. Multi-seed replication: five random seeds (42, 123, 456, 789, and 2024) with 95% CI reported for all three datasets.
4. Alternative types of drift: Gaussian, uniform, spike and covariate-shift drift.
5. Treatment of missing values: Zero vs. Median imputation.
6. Perform feature correlation analysis removing features with high correlation (0.95).
7. Hyperparameter sensitivity: determine the sensitivity of the results to the number of estimators used for IF and Gaussian components used for GMM.
8. Computational profiling: Profiling of training times, inference latency and peak memory usage.
9. Comparison of different architectures of autoencoder: evaluation of variants of the encoder-decoder architecture.
10. Statistical significance testing: two-sided Wilcoxon signed-rank tests with Bonferroni correction and Cliff's delta effect sizes for the evaluation of differences between the paradigms of detection.
11. Analysis of anomaly-scores during threshold calibration: quantitative evaluation of anomaly-score distributions.
12. Scalability analysis: profiling of training and inference cost for different sizes of training data.
Together, these analyses provide a comprehensive assessment of the robustness, statistical reliability, computational efficiency, and generalizability of the proposed behavioural evaluation framework under diverse experimental conditions.
4. Results
This section presents experimental results for each of the datasets and an analysis of the results across datasets. All results are obtained from the unified pipeline outlined in Section 3, and the same preprocessing, training and evaluation protocol is applied to all three datasets.
4.1 CIC-IoT-2023: paradigm-dependent drift response
The effects of injected concept drift on the CIC-IoT-2023 dataset reveal markedly different behaviours between the isolation-based and density/reconstruction-based detection paradigms. IF exhibited a progressive increase in both detection sensitivity and false alarms following drift injection. The mean recall increased from 0.149 during the pre-drift period (Segments 0–9) to 0.288 during the drift period (Segments 10–19), while the mean FPR increased from 0.017 to 0.035. Similarly, the F1-score increased from 0.259 in the pre-drift segments to approximately 0.55 in the final segment. The DS increased modestly from 0.132 to 0.253, indicating that the improvement in recall exceeded the corresponding increase in FPR. The overall performance indices were mean recall = 0.218, mean FPR = 0.026, BI = 0.782, and SV = 0.0072.
Although the aggregate performance metrics suggest improved detection, this behaviour reflects sensitivity inflation rather than genuine robustness. Under the fixed decision threshold, progressive drift shifts the anomaly-score distribution, causing an increasing number of attack samples and eventually more benign samples to exceed the calibration threshold. Consequently, recall and FPR increase simultaneously without any model adaptation. This effect is reflected by the substantial increase in behavioural instability, with SV increasing by approximately 650-fold (from 0.000011 under the baseline condition to 0.0072 with drift) and the FPR rising from 0.017 to 0.063 in the final segment. Importantly, this behaviour is not universally beneficial. On datasets with different score distributions, such as Edge-IIoTset, the same fixed-threshold mechanism produces the opposite outcome, leading to near-complete detection failure. Therefore, from an operational perspective, the progressively changing alarm behaviour of IF under drift reduces deployment reliability, as detector sensitivity becomes unpredictable without periodic recalibration.
In contrast, the GMM demonstrated strong resistance to injected drift. There was no significant change in recall over time with a mean of 0.508 and a small amount of temporal variation (SV = 0.000028). The FPR stayed low at around 0.019 throughout the temporal segments and the DS was around 0.490 throughout the temporal segments. The overall summary indices were: Mean recall = 0.508, Mean FPR = 0.019, BI = 0.492 and SV = 0.000028, suggesting that the detectors' behaviour remained very stable despite the progressive perturbation.
The same drift-resistant behaviour was observed in the AE. The mean recall was 0.401 and there was little temporal variation (SV = 0.000015). The FPR was consistently low (around 0.018) with the DS remaining fairly stable (around 0.383) during the evaluation period. The mean recall was 0.401, mean FPR was 0.018, BI was 0.599, and SV was 0.000015, which demonstrated that the reconstruction-based detection had good discrimination even under the injected drift.
Figure 2: Temporal trajectories of detection metrics on CIC-IoT-2023. Shaded area indicates drift injection zone (segments 10–19).
These observations are confirmed by the baseline experiments. In the baseline condition (α=0), IF was at a stable level of 0.149 and there was very little behavioural variation (SV = 0.000011). Following drift injection (α = 0.5), Recall progressively increased to 0.384 in the final segment, accompanied by a substantial increase in temporal instability (SV = 0.0072). In contrast, GMM produced nearly identical performance under both baseline and drift conditions, confirming that its behaviour is largely insensitive to the injected perturbation. These results support the hypothesis that drift vulnerability is primarily associated with the isolation-based detection paradigm, whereas density- and reconstruction-based approaches exhibit substantially greater behavioural robustness under progressive concept drift.
Baseline comparison (α = 0 vs. α = 0.5): To distinguish the effects of injected concept drift from the inherent characteristics of the data, the entire experimental pipeline was repeated under a baseline condition (α = 0), in which no artificial drift was introduced. Table 1 compares the summary performance and behavioural stability indices obtained under the baseline and drift-injected (α = 0.5) conditions.
Table 1 CIC-IoT-2023 behavioural comparison: with and without drift injection
Condition | Algorithm | Mean Recall | Mean FPR | BI | AFI | SV | Mean DS
α = 0 (no drift) | IF | 0.149 | 0.017 | 0.851 | 0.017 | 0.000011 | 0.132
α = 0 (no drift) | GMM | 0.508 | 0.018 | 0.492 | 0.018 | 0.000030 | 0.490
α = 0.5 (with drift) | IF | 0.218 | 0.026 | 0.782 | 0.026 | 0.007206 | 0.192
α = 0.5 (with drift) | GMM | 0.508 | 0.019 | 0.492 | 0.019 | 0.000028 | 0.490
4.2 Edge-IIoTset: behavioural divergence
The Edge-IIoTset results exhibit the most pronounced contrast between the two anomaly detection paradigms. Under the same dataset, evaluation pipeline, and threshold calibration strategy, the isolation-based and density/reconstruction-based approaches produced fundamentally different behaviours.
This divergence provides the strongest evidence supporting the paradigm-level interpretation proposed in this study. Unlike the CIC-IoT-2023 and UNSW-NB15 datasets, where behavioural differences emerged following injected concept drift, the divergence on Edge-IIoTset was observed under the baseline evaluation without artificial drift. Specifically, IF exhibited almost complete detection failure (recall ≈ 0), whereas both the GMM and the AE achieved near-perfect detection performance (recall ≈ 1.0). Consequently, the observed behavioural gap cannot be attributed to the drift injection procedure but instead reflects the fundamentally different ways in which the detection paradigms model the underlying distribution of benign traffic. This discovery reinforces the result that the behaviour of the detector is not solely dependent on the algorithm but also on the paradigm of anomaly detection that is used.
The IF was totally blind in their behaviour during the evaluation. The detection rate for all temporal segments (including the mixed segments, Segment 2 with 25% attack samples and Segment 4 with 83% attack samples) was essentially zero, meaning that the detector was unable to detect attacks even when the attacks were the majority of the segment. The FPR was still quite low (around 0.01), but the few positive predictions that were produced by the detector were almost entirely false alarms and not true alarm indications of an attack. As a result, the overall behavioural indices were BI = 1.0000, AFI ≈ 0.010 and DS = -0.01, meaning that the discrimination performance was below random guessing. By comparison, the GMM had very stable and accurate detection. Recall was 1.0 in all attack containing segments and FPR was consistently low at about 0.011 in the mixed segments and 0.013 in the benign-only segment.
The corresponding behavioural indices were BI = 0.0000, AFI = 0.012, and DS = 0.989, indicating excellent discrimination capability and stable temporal behaviour.
The AE closely replicated the performance of GMM. Recall remained 1.0 throughout all attack-containing segments, while the FPR was approximately 0.011 in the mixed segments and 0.009 in the benign-only segment. The resulting behavioural indices (BI = 0.0000, AFI = 0.011, and DS = 0.988) further confirm the robustness of the reconstruction-based detection paradigm on this dataset.
Figure 3: Behavioural divergence on Edge-IIoTset: IF (blind) vs. GMM (perfect detection).
Per-segment performance statistics are shown in Table 2.
Table 2 Temporal evaluation results for Edge-IIoTset (all segments with computable metrics)
Segment | True attack rate | Segment type | IF Recall | IF FPR | GMM recall | GMM FPR | AE Recall | AE FPR
0 | 1.000 | Attack-only | 0.000 | n/a | 1.000 | n/a | 1.000 | n/a
1 | 1.000 | Attack-only | 0.000 | n/a | 1.000 | n/a | 1.000 | n/a
2 | 0.246 | Mixed | 0.000 | 0.011 | 1.000 | 0.011 | 1.000 | 0.011
3 | 0.000 | Benign-only | n/a | 0.011 | n/a | 0.013 | n/a | 0.009
4 | 0.831 | Mixed | 0.000 | 0.010 | 1.000 | 0.011 | 1.000 | 0.013
5–19 | 1.000 | Attack-only | 0.000 | n/a | 1.000 | n/a | 1.000 | n/a
Baseline Comparison (α = 0 vs. α = 0.5): Table 3 presents a comparison of the summary performance and behavioural stability indices under the baseline (α = 0) and drift-injected (α = 0.5) conditions. The results are almost the same both with and without it. This verifies that the behavioural difference seen in Edge-IIoTset is not caused by the drift injected but is a natural difference in the data.
Table 3 Edge-IIoTset behavioural comparison: with and without drift injection
Condition | Algorithm | Mean Recall | Mean FPR | BI | AFI | Mean DS
α = 0 (no drift) | IF | 0.000 | 0.011 | 1.000 | 0.011 | −0.010
α = 0 (no drift) | GMM | 1.000 | 0.012 | 0.000 | 0.012 | 0.989
α = 0.5 (with drift) | IF | 0.000 | 0.011 | 1.000 | 0.011 | −0.010
α = 0.5 (with drift) | GMM | 1.000 | 0.012 | 0.000 | 0.012 | 0.989
To test the robustness and reproducibility of the proposed framework, five different random seeds (42, 123, 456, 789, and 2024) were used. Table 4 summarizes the aggregated results, reported as the mean ± SD across all runs.
Table 4 Multi-seed replication for Edge-IIoTset (5 seeds, mean ± std)
Algorithm | Mean Recall | Mean FPR | BI | AFI
IF | 0.0001 ± 0.0001 | 0.0100 ± 0.0014 | 0.9999 ± 0.0001 | 0.0100 ± 0.0014
GMM | 1.0000 ± 0.0000 | 0.0101 ± 0.0019 | 0.0000 ± 0.0000 | 0.0101 ± 0.0019
Nearly zero SDs suggest that the behavioural difference is reproducible and not a result of an unstable configuration.
4.3 UNSW-NB15: confirming paradigm-dependent behaviour
The UNSW-NB15 dataset provides an important intermediate validation environment between CIC-IoT-2023 and Edge-IIoTset, with an attack proportion of approximately 68%. Consequently, it serves as an effective benchmark for assessing whether the observed behavioural patterns generalize across datasets with different class distributions.
Similar to the behaviour observed on CIC-IoT-2023, the IF exhibited sensitivity inflation under injected concept drift. The mean recall increased from 0.190 during the pre-drift period (Segments 0–9) to 0.316 during the drift period (Segments 10–19), while the mean FPR increased from 0.009 to 0.029. The increase in recall also increased the DS from 0.181 to 0.289, which meant that recall grew faster than FPR. The overall summary indices were mean recall = 0.253, mean FPR = 0.018, BI = 0.747, and SV = 0.0062. The results observed in CIC-IoT-2023 suggest that the improvement in IF is mainly due to the change in the distribution of the anomaly scores, which is induced by the drift, under a fixed decision threshold, and not to the robustness of the model.
Again, the GMM showed drift resistant behaviour. The mean recall was not altered (0.430) and there was no significant variation over time (SV = 0.000138). The FPR stayed at about 0.010 and the DS was almost constant at 0.419 during the evaluation. The mean recall, mean FPR, BI and SV were 0.430, 0.010, 0.570 and 0.000138 respectively, which further showed the robustness of the density-based detection paradigm.
The AE outperformed the other models in terms of overall detection performance and had the lowest rate of concept drift. The mean of the recall was 0.575, and there was only a small amount of time variation (SV = 0.000409). The FPR was always found to be quite small around 0.010 and mean DS 0.565. The FPR did not increase with recall in the last few segments, although it did increase slightly (from 0.554 to 0.632), the magnitude of this change was significantly less than that observed for IF. This behaviour is due to the stable performance of the detector and not the sensitivity inflation due to drift.
Figure 4: Temporal trajectories of detection metrics on UNSW-NB15.
The trends observed are very similar to that of CIC-IoT-2023 (Figure 2). Specifically, the recall and FPR of IF shows a progressive increase after the injected drift, while those of GMM and AE remain almost constant during the evaluation time. Among the three algorithms, AE achieves the highest overall Recall while preserving stable discrimination.
Baseline Comparison (α = 0 vs. α = 0.5): The baseline experiments further confirm that the behavioural changes observed for IF are primarily caused by the injected concept drift rather than the intrinsic characteristics of the dataset. Under the baseline condition (α = 0), the mean recall was 0.167, increasing to 0.253 after drift injection (α = 0.5), while the SV increased by approximately 100-fold. In contrast, the GMM exhibited only marginal differences between the two conditions, with the mean recall changing from 0.426 to 0.430, demonstrating its insensitivity to the injected perturbation. A quantitative comparison of the baseline and drift-injected conditions is presented in Table 5.
Table 5 UNSW-NB15 behavioural comparison: with and without drift injection
Condition | Algorithm | Mean Recall | Mean FPR | BI | AFI | SV | Mean DS
α = 0 (no drift) | IF | 0.167 | 0.009 | 0.833 | 0.009 | 0.000062 | 0.158
α = 0 (no drift) | GMM | 0.426 | 0.010 | 0.574 | 0.010 | 0.000090 | 0.415
α = 0.5 (with drift) | IF | 0.253 | 0.018 | 0.747 | 0.018 | 0.006211 | 0.235
α = 0.5 (with drift) | GMM | 0.430 | 0.010 | 0.570 | 0.010 | 0.000138 | 0.419
4.4 Cross-dataset comparison
Table 6 summarizes the behavioural characteristics of all algorithm–dataset combinations, from which several important observations can be drawn.
First, the observed behaviour is strongly associated with the underlying anomaly detection paradigm. The IF, representing the isolation-based paradigm, consistently exhibited either sensitivity inflation (CIC-IoT-2023 and UNSW-NB15) or behavioural blindness (Edge-IIoTset). In contrast, both the GMM and the AE, representing the density- and reconstruction-based paradigms, demonstrated stable detection performance across all three datasets. The consistency of these findings across two independent algorithms provides compelling evidence that vulnerability to concept drift is primarily a paradigm-level characteristic rather than a property of an individual algorithm.
Second, the density- and reconstruction-based paradigms demonstrated advantages in both temporal stability and discriminative capability. Across the three datasets, the Mean DS ranged from −0.010 to 0.235 for IF, compared with 0.419 to 0.989 for GMM and 0.383 to 0.988 for AE. Likewise, GMM and AE always had significantly lower SV than IF. These results show that the density- and reconstruction-based methods not only have more stable performance in the presence of concept drift, but also perform better in discriminating between the benign and malicious traffic. Consequently, the drift-resistant paradigms improve both robustness and detection performance, rather than requiring a trade-off between the two.
Third, the DS metric provides the clearest overall ranking of detector performance across all experimental settings: GMM (Edge-IIoTset, 0.989) ≈ AE (Edge-IIoTset, 0.988) > AE (UNSW-NB15, 0.565) > GMM (CIC-IoT-2023, 0.490) > GMM (UNSW-NB15, 0.419) > AE (CIC-IoT-2023, 0.383) > IF (UNSW-NB15, 0.235) > IF (CIC-IoT-2023, 0.192) > IF (Edge-IIoTset, −0.010).
This ranking consistently favours the density- and reconstruction-based paradigms while highlighting the limitations of the isolation-based approach. Fourth, the sensitivity inflation observed in IF was consistently reproduced on two independent datasets (CIC-IoT-2023 and UNSW-NB15), despite their substantially different attack proportions (approximately 97% and 68%, respectively). Furthermore, the baseline experiments (α = 0) confirmed that this behaviour was induced by the injected concept drift rather than by the intrinsic characteristics of the datasets, providing additional evidence for the robustness of the proposed experimental design.
Finally, the relative performance of the two drift-resistant detectors was found to be dataset dependent. The AE achieved the highest detection performance on UNSW-NB15 (mean recall = 0.575, Mean DS = 0.565), outperforming GMM (mean DS = 0.419). Conversely, GMM achieved superior discrimination on CIC-IoT-2023 (Mean DS = 0.490 versus 0.383 for AE), whereas both detectors exhibited virtually identical performance on Edge-IIoTset (0.989 versus 0.988). These findings indicate that the AE should be regarded as competitive with, rather than universally superior to, GMM. Overall, both the reconstruction-based and density-based paradigms provide an effective combination of high detection accuracy, strong temporal stability, and robustness to concept drift, with their relative performance depending on the characteristics of the dataset.
Figure 5: Cross-dataset behavioural stability indices for all algorithm–dataset combinations.
Figure 6: Baseline comparison (α = 0 vs. α = 0.5) across all three datasets.
Table 6 Cross-dataset behavioural pattern summary
Dataset | IF | GMM | AE
CIC-IoT-2023 | Sensitivity Inflation | Drift-Resistant | Drift-Resistant
Edge-IIoTset | Immediate Blindness | Stable Detection | Stable Detection
UNSW-NB15 | Sensitivity Inflation | Drift-Resistant | Drift-Resistant
4.5 Robustness and sensitivity analysis
Segment-count sensitivity analysis: To evaluate whether the reported behavioural patterns depend on the temporal segmentation strategy, the complete evaluation pipeline was repeated using K=10, 15, 25 and 30 temporal segments. Figure 7 presents the Mean DS for the IF and GMM across the different segmentation levels. On CIC-IoT-2023 and Edge-IIoTset, the Mean DS remained nearly unchanged over the entire range of segment counts, while GMM also exhibited stable performance on UNSW-NB15. These results indicate that the behavioural patterns reported in Section 4 are robust to the choice of temporal segmentation and are not an artifact of using 20 segments. The value K = 20 was therefore retained for the main experiments because it provides an appropriate balance between temporal resolution and statistical reliability, enabling the onset of behavioural changes to be identified while ensuring that each segment contains sufficient samples for stable performance estimation.
Figure 7: Segment-count sensitivity: Mean DS across K = 10–30 segments. The dashed line marks the chosen K = 20.
Threshold sensitivity analysis: To evaluate the influence of the decision threshold on detector performance, the fixed 99th-percentile threshold was varied to the 95th, 97th, and 99.5th percentiles. As illustrated in Figure 8, the absolute Mean DS values changed with the threshold selection; however, the relative ranking of the algorithms remained unchanged across all threshold settings. Specifically, the GMM consistently outperformed the IF on both CIC-IoT-2023 and UNSW-NB15, while the near-complete behavioural blindness of IF on Edge-IIoTset persisted under every threshold configuration, with the Mean DS remaining at or below zero. These findings demonstrate that the principal conclusions of the study are robust to the choice of decision threshold, indicating that the observed paradigm-level differences are not an artifact of threshold calibration.
Figure 8: Threshold sensitivity: Mean DS at the 95th–99.5th percentiles. The dashed line marks the chosen 99th percentile.
Multi-Seed Robustness with Confidence Intervals: The multi-seed analysis was extended from Edge-IIoTset to all three datasets, and 95% CIs were computed over five random seeds (Figure 9). The CI are narrow for the drift-resistant configurations. For example, GMM on Edge-IIoTset achieved BI = 0.000 ± 0.000 and Mean DS = 0.990 ± 0.002, demonstrating excellent reproducibility. In contrast, IF exhibited wider CI on CIC-IoT-2023 (Mean DS = 0.083, 95% CI [−0.002, 0.168]), reflecting the stochastic nature of its axis-aligned partitioning. Nevertheless, its elevated BI remained clearly distinguishable from that of GMM.
The single-seed results reported in Sections 4.1–4.3 (e.g., an IF recall of 0.218 on CIC-IoT-2023 under injected drift) correspond to a representative experiment conducted with a fixed random seed, whereas the results presented here represent the mean performance across five independent runs. The differences between the single-seed and multi-seed results arise from the inherent stochasticity of the IF algorithm, which is explicitly quantified by the reported CI. In contrast, the density- and reconstruction-based detectors exhibit negligible variability across different random seeds.
An important statistical observation is that, for CIC-IoT-2023, the 95% CI for the IF Mean DS includes zero ([−0.002, 0.168]), indicating that its discrimination performance is not consistently distinguishable from random performance across random seeds under the selected calibration. This finding is consistent with the central conclusion of the study, as the density- and reconstruction-based detectors produce narrow, well-separated CI that remain substantially above zero, whereas the isolation-based detector exhibits both lower performance and greater statistical variability. Although the absolute metric values vary across random seeds, the qualitative behavioural pattern characterized by higher and more variable BI for IF relative to the density- and reconstruction-based detectors remains consistent. Accordingly, the paradigm-level conclusions are robust to random initialization and experimental replication.
Figure 9: Multi-seed robustness with 95% CIs across all three datasets (five seeds). Note that differences from the single-seed results in Sections 4.1–4.3 are due to the averaging over five independent runs, while the larger intervals for IF are due to its stochastic partitioning.
To verify that the findings are not specific to Gaussian noise, three additional drift mechanisms were tested: uniform noise, sparse high-magnitude spikes, and systematic covariate shift. Table 7 shows that the paradigm-dependent pattern holds across all four drift types. GMM remains drift-resistant on CIC-IoT-2023 (Mean DS ≈ 0.49 for every type), on Edge-IIoTset (Mean DS ≈ 0.989 for every type), and on UNSW-NB15 (Mean DS ≈ 0.42). IF remains blind on Edge-IIoTset (Mean DS ≈ −0.010). On UNSW-NB15, covariate shift raises IF Mean DS from 0.234 under Gaussian drift to 0.407, indicating that isolation-based detection is more responsive to systematic mean shifts than to zero-mean or spike perturbations; nevertheless, GMM remains slightly higher under every drift type.
Table 7 Mean DS under four drift types
Dataset | Algorithm | Gaussian | Uniform | Spike | Covariate
CIC-IoT-2023 | IF | 0.192 | 0.178 | 0.139 | 0.240
CIC-IoT-2023 | GMM | 0.490 | 0.490 | 0.490 | 0.489
Edge-IIoTset | IF | -0.010 | -0.010 | -0.010 | -0.010
Edge-IIoTset | GMM | 0.989 | 0.989 | 0.989 | 0.989
UNSW-NB15 | IF | 0.234 | 0.220 | 0.183 | 0.407
UNSW-NB15 | GMM | 0.420 | 0.417 | 0.417 | 0.431
Preprocessing robustness: Two preprocessing decisions were examined. First, replacing zero-imputation with median-imputation produced identical Mean DS values to three decimal places on all three datasets, confirming that the imputation choice does not introduce artificial patterns. Second, removing highly correlated features (Pearson |r| > 0.95) reduced the feature count from 46 to 37 on CIC-IoT-2023, from 43 to 38 on Edge-IIoTset, and from 41 to 36 on UNSW-NB15. The paradigm-dependent ordering is preserved: GMM remains stronger than IF in every dataset. Some IF values shift modestly after decorrelation (for example, 0.192→0.234 on CIC-IoT-2023 and −0.010→0.047 on Edge-IIoTset), but these shifts do not reverse the reported behavioural conclusions as shown in Table 8.
Table 8 Preprocessing robustness: imputation strategy and feature-correlation removal (Mean DS)
Dataset | Algorithm | Zero-imp | Median-imp | Features (all→decorr.) | Mean DS (all→decorr.)
CIC-IoT-2023 | IF | 0.192 | 0.192 | 46→37 | 0.192→0.234
CIC-IoT-2023 | GMM | 0.490 | 0.490 | 46→37 | 0.490→0.488
Edge-IIoTset | IF | -0.010 | -0.010 | 43→38 | -0.010→0.047
Edge-IIoTset | GMM | 0.989 | 0.989 | 43→38 | 0.989→0.988
UNSW-NB15 | IF | 0.234 | 0.234 | 41→36 | 0.234→0.193
UNSW-NB15 | GMM | 0.420 | 0.420 | 41→36 | 0.420→0.416
The fixed hyperparameters were swept so that the comparison would not depend on the particular values chosen. IF was re-run with 50, 100, 200 and 300 estimators, GMM with 1, 2, 3 and 5 mixture components, and the AE with the three architectures listed in Table 11. Table 9 shows that none of these settings reverses the paradigm-level ordering on which the conclusions rest. On Edge-IIoTset, GMM stays drift-resistant at every component count while IF stays blind, with Mean DS of 0.988 to 0.990 and −0.010 respectively. GMM also outperforms IF across the whole swept range on CIC-IoT-2023 and UNSW-NB15. No formal grid or Bayesian search against a validation objective was carried out, because the study describes how detectors behave over time under one fixed and deployment-representative configuration rather than seeking the best achievable score for any single algorithm.
Table 9 Hyperparameter sensitivity (range of Mean DS across swept values)
Dataset | Algorithm (sweep) | Mean DS range
CIC-IoT-2023 | IF (n_est 50–300) | 0.123 – 0.220
CIC-IoT-2023 | GMM (n_comp 1–5) | 0.370 – 0.508
Edge-IIoTset | IF (n_est 50–300) | -0.012 – -0.009
Edge-IIoTset | GMM (n_comp 1–5) | 0.989 – 0.990
UNSW-NB15 | IF (n_est 50–300) | 0.192 – 0.234
UNSW-NB15 | GMM (n_comp 1–5) | 0.259 – 0.420
Computational efficiency analysis: Training time, inference latency, and peak memory consumption were profiled for all algorithms, and the results are summarized in Table 10. The GMM demonstrated substantially higher computational efficiency than the IF. Specifically, the GMM required approximately two orders of magnitude less training time, taking 0.023 s compared with 3.47 s for IF on the CIC-IoT-2023 dataset. Similarly, the per-sample inference latency of GMM ranged from 0.6–0.9 μs, whereas IF required 9.6–14.3 μs. The peak memory usage for both algorithms was also low, and was less than 8 MB in all experiments. The results show that the drift-resistant density-based paradigm not only guarantees better behavioural stability, but also is also much more efficient in terms of computation, thus making it well suited for resource-limited IoT environments. All experiments were conducted on a standard Google Colab CPU runtime (Intel Xeon processor at 2.20 GHz with 13 GB RAM) using a single-threaded configuration where applicable. The GMM employed a single Gaussian component with diagonal covariance, which contributed to its exceptionally low training time.
Table 10 Computational profiling (training time, per-sample inference, peak memory)
Dataset | Algorithm | Train (s) | Inference (µs/sample) | Peak memory (MB)
CIC-IoT-2023 | IF | 3.472 | 9.64 | 0.88
CIC-IoT-2023 | GMM | 0.023 | 0.89 | 1.06
Edge-IIoTset | IF | 4.161 | 14.32 | 1.43
Edge-IIoTset | GMM | 0.092 | 0.77 | 7.59
UNSW-NB15 | IF | 3.506 | 11.44 | 0.93
UNSW-NB15 | GMM | 0.045 | 0.62 | 3.34
AE architecture justification: The selected encoder–decoder architecture (d → 64 → 32 → 64 → d) was compared with a shallower configuration (d → 32 → d) and a deeper configuration (d → 128 → 64 → 32 → ... → d). As shown in Table 11, the selected architecture achieved the best overall performance on Edge-IIoTset (Mean DS = 0.991) and UNSW-NB15 (Mean DS = 0.619). Although the deeper architecture attained a slightly higher Mean DS on CIC-IoT-2023 (0.416 versus 0.392), it also exhibited a higher SV, indicating reduced temporal consistency. The shallow architecture showed the weakest performance, particularly on Edge-IIoTset, where its Mean DS was 0.829, compared with 0.991 for the selected architecture. Consequently, the proposed encoder–decoder configuration was retained because it provides the most favourable balance between detection performance, temporal stability, and computational complexity, rather than being uniformly superior across all datasets.
Table 11 AE architecture comparison
Dataset | Architecture | Mean DS | BI | SV
CIC-IoT-2023 | Shallow (32) | 0.395 | 0.552 | 0.007860
CIC-IoT-2023 | Selected (64,32) | 0.392 | 0.560 | 0.006356
CIC-IoT-2023 | Deep (128,64,32) | 0.416 | 0.505 | 0.017618
Edge-IIoTset | Shallow (32) | 0.829 | 0.079 | 0.011952
Edge-IIoTset | Selected (64,32) | 0.991 | 0.000 | 0.000000
Edge-IIoTset | Deep (128,64,32) | 0.988 | 0.000 | 0.000000
UNSW-NB15 | Shallow (32) | 0.574 | 0.396 | 0.006085
UNSW-NB15 | Selected (64,32) | 0.619 | 0.312 | 0.012973
UNSW-NB15 | Deep (128,64,32) | 0.565 | 0.317 | 0.020043
Note: The DS values in this architecture comparison differ marginally from the main-results values (Sections 4.1–4.3) because the models were retrained for this controlled comparison; the small differences do not affect the ranking of architectures. The AE's main-run summary values are subject to the same effect: independent retraining runs vary by about ±0.01 in Mean DS (for example, 0.383 versus 0.393 on CIC-IoT-2023) owing to residual non-determinism in network training, without changing the drift-resistant pattern.
Statistical significance: A formal statistical analysis was conducted to confirm whether the differences in the per-segment DS values between the anomaly detection paradigms were statistically significant. Accordingly, two-sided Wilcoxon signed-rank tests were performed on the paired per-segment DS values, and Cliff's delta (δ) was calculated to quantify the effect size. The results are summarized in Table 12.
The statistical findings are consistent with the multi-seed 95% CIs presented in Figure 9, which are clearly separated across the detection paradigms. For example, on CIC-IoT-2023, the Mean DS interval for IF is [−0.002, 0.168], compared with [0.456, 0.493] for GMM. Similarly, on UNSW-NB15, the corresponding intervals are [0.200, 0.279] for IF and [0.406, 0.443] for GMM, providing further evidence of the superior and more consistent performance of the density-based detector.
In Table 12, the sign of Cliff's delta (δ) follows the order of the listed comparison. Thus, δ = −1.0 for the IF vs. GMM and IF vs. AE comparisons indicates that IF achieved lower DS values than both GMM and AE in every mixed temporal segment. For Edge-IIoTset, only two mixed segments were available, making the Wilcoxon signed-rank test inapplicable; therefore, the corresponding entries are reported as n/a. Nevertheless, the effect-size analysis remains consistent with the overall findings, showing that IF performs worse than both GMM and AE, whereas GMM and AE exhibit virtually identical performance (δ = 0.0), indicating a negligible effect size.
Table 12 Statistical significance of paradigm differences in per-segment DS (two-sided Wilcoxon signed-rank, Bonferroni-corrected; Cliff's δ effect size)
Dataset | Comparison | n (mixed) | Wilcoxon p | Significant (α=0.0167) | Cliff's δ
CIC-IoT-2023 | IF vs GMM | 20 | < 0.0001 | Yes | -1.0 (large)
CIC-IoT-2023 | IF vs AE | 20 | < 0.0001 | Yes | -1.0 (large)
CIC-IoT-2023 | GMM vs AE | 20 | < 0.0001 | Yes | +1.0 (large)
Edge-IIoTset | IF vs GMM | 2 | n/a | n/a | -1.0 (large)
Edge-IIoTset | IF vs AE | 2 | n/a | n/a | -1.0 (large)
Edge-IIoTset | GMM vs AE | 2 | n/a | n/a | +0.0 (negligible)
UNSW-NB15 | IF vs GMM | 20 | < 0.0001 | Yes | -1.0 (large)
UNSW-NB15 | IF vs AE | 20 | < 0.0001 | Yes | -1.0 (large)
UNSW-NB15 | GMM vs AE | 20 | < 0.0001 | Yes | -1.0 (large)
Note: Wilcoxon is marked n/a on Edge-IIoTset because the dataset has only two mixed segments. The Cliff's δ signs follow the pair order in Table 12; negative values for IF vs GMM/AE indicate that IF has lower per-segment DS values.
Score-distribution analysis: The mechanism described in Section 5.2 can be verified through a direct analysis of the anomaly score distributions. During threshold calibration, the anomaly scores of benign and attack samples were collected from the first mixed segment of each dataset and compared with the fixed 99th-percentile threshold (τ). Table 13 summarizes three score-separation statistics: (i) the area under the receiver operating characteristic curve (AUROC) computed from the raw anomaly scores, (ii) the normalized mean score gap between attack and benign samples, and (iii) the proportion of attack samples with anomaly scores exceeding τ.
The three characteristic behavioural modes emerge exactly as expected. On Edge-IIoTset, the IF exhibits almost no score separation (AUC = 0.517, normalized gap = −0.007), with no attack samples exceeding the calibrated threshold. This represents the behavioural blindness observed in Section 4.2. In contrast, both the GMM and the AE achieve near-perfect score separation on the same dataset (AUC = 1.000 and 0.998, respectively), with all attack samples exceeding the decision threshold.
On CIC-IoT-2023 and UNSW-NB15, IF achieves relatively high score separation (AUC = 0.897 and 0.963, respectively), yet only a small proportion of attack samples exceed the calibrated threshold (0.149 and 0.198, respectively). As concept drift progressively shifts the anomaly score distribution, increasing numbers of both attack and benign samples cross the fixed threshold, producing the sensitivity inflation observed in Sections 4.1 and 4.3. In contrast, the density- and reconstruction-based detectors maintain a relatively stable proportion of attack samples above the threshold (0.405–0.575 on the drifted datasets and 1.000 on Edge-IIoTset), consistent with their drift-resistant behaviour.
As an internal consistency check, the proportion of attack samples exceeding the calibration threshold closely matches the first-segment recall values reported in Sections 4.1–4.3. For example, on CIC-IoT-2023, the corresponding values are 0.149 for IF and 0.513 for GMM, confirming the consistency between the score-distribution analysis and the temporal behavioural evaluation.
Table 13 Score-separation statistics at calibration time (first mixed segment): AUC of raw anomaly scores, normalised mean gap between attack and benign scores, and fraction of attack scores above the fixed threshold τ
Dataset | Algorithm | Score AUC | Normalised gap | Fraction of attacks > τ
CIC-IoT-2023 | IF | 0.897 | 1.984 | 0.149
CIC-IoT-2023 | GMM | 0.924 | 0.071 | 0.513
CIC-IoT-2023 | AE | 0.921 | 0.043 | 0.405
Edge-IIoTset | IF | 0.517 | -0.007 | 0.000
Edge-IIoTset | GMM | 1.000 | 0.191 | 1.000
Edge-IIoTset | AE | 0.998 | 0.231 | 1.000
UNSW-NB15 | IF | 0.963 | 1.844 | 0.198
UNSW-NB15 | GMM | 0.973 | 0.184 | 0.418
UNSW-NB15 | AE | 0.991 | 0.126 | 0.575
Scalability Across Data Volumes: To evaluate scalability, the training and inference costs of the three anomaly detection algorithms were measured using four subsets of the CIC-IoT-2023 dataset containing 25,000, 50,000, 100,000, and 200,000 samples. The evaluation protocol was identical to that of the main experimental pipeline, where each model was trained using the benign samples from the first eight temporal segments at the corresponding data volume. All reported values are the median of three independent runs performed using Google Colab CPU runtime. The results in Table 14 show that the computational ranking of the algorithms does not change with the increase of the data volume. The IF took 0.474–0.585 s for training while the GMM took only 0.003–0.006 s for training. The highest training cost was the AE, which had to spend 1.862–2.600 s because of the iterative optimization of the neural networks. This trend was also seen for inference latency. As the size of the dataset increased, the per-sample inference time of the IF model was between 10.60 and 12.02 μs, the per-sample inference time of the GMM model was between 0.61 and 0.65 μs, and the per-sample inference time of the AE model was between 12.34 and 14.09 μs.
These results are similar to the computational profiling results presented in Table 10, where GMM was demonstrated to be one to two orders of magnitude more computationally efficient than the tree-based (IF) and neural network-based (AE) detectors. The absolute execution times are different due to the different amount of training data used and experimental setup, but the relative computation order of the three algorithms is the same. Moreover, the slight rise in training and inference costs with the growth of the data volume demonstrates good scalability.
Even at the largest evaluated volume (200,000 samples), none of the algorithms exhibited computational overhead that would limit their applicability to real-time or streaming IoT environments, while GMM remained the most efficient and computationally scalable approach.
Table 14 Scalability of training time and per-sample inference latency across data volumes on CIC-IoT-2023 (median of three repetitions)
Volume | Algorithm | Training time (s) | Inference (µs/sample)
25,000 | IF | 0.585 | 12.02
25,000 | GMM | 0.003 | 0.61
25,000 | AE | 2.263 | 14.07
50,000 | IF | 0.421 | 10.80
50,000 | GMM | 0.004 | 0.63
50,000 | AE | 1.862 | 14.09
100,000 | IF | 0.455 | 10.60
100,000 | GMM | 0.004 | 0.63
100,000 | AE | 2.562 | 12.34
200,000 | IF | 0.474 | 10.85
200,000 | GMM | 0.006 | 0.65
200,000 | AE | 2.600 | 13.48
5. Discussion
The results reveal important differences in how the evaluated detectors respond to concept drift. The following discussion relates these observations to IDS evaluation and practical deployment.
5.1 Behavioural patterns and their operational implications
Based on the experimental results, it is not always true that concept drift leads to a single uniform pattern of performance degradation. Instead, three qualitatively different behavioural patterns were found, and they have different operational consequences. Immediate blindness (IF on Edge-IIoTset) is the most operationally dangerous pattern. The detector produces no true detections, yet continues to operate silently, giving little external indication that a failure has occurred. In a deployment setting, this would leave the network effectively unprotected while providing operators with a false sense of security. In this condition, low alarm rates, which are typically considered desirable, become indistinguishable from the absence of detection. The detector still produces a few anomaly alarms (FPR ≈ 0.01); however, these alarms are virtually unrelated to attack samples and thus the detector is not completely silent, but rather blind to attacks. Sensitivity inflation (IF on CIC-IoT-2023 and UNSW-NB15) is characterised by a progressive increase in both recall and FPR as the distribution shifts. A plausible explanation is that, as drift accumulates, the partition boundaries learned by IF become misaligned, so that more samples from both classes cross the fixed decision threshold. The α = 0 baselines support this interpretation: with no injected drift, IF remains temporally flat on both datasets. This pattern is particularly problematic because aggregate measures such as the F1-score can appear to improve even as discrimination quality deteriorates.
Drift-resistant detection (GMM and AE in all three environments) indicates that behavioural collapse due to drift is not certain in the evaluated environments. Both density-based and reconstruction-based detection methods proved to be comparably stable across the tested environments. One possible and plausible explanation is that the presence of anomaly boundaries anchored to the learned benign statistics is less influenced by the moderate additive perturbation employed in the drift injection. In practical terms, temporal stability is therefore an important dimension of detector quality in its own right, that is, separate from absolute detection performance.
These findings demonstrate that the operational evaluation of IDSs should address three complementary aspects: (i) whether the detector continues to identify attacks effectively, (ii) whether it maintains a low false alarm rate, and (iii) whether its detection behaviour remains stable over time. No single aggregate performance metric can adequately capture all three dimensions. Consequently, a comprehensive evaluation requires complementary behavioural stability measures in addition to conventional classification metrics.
5.2 Paradigm-dependent drift vulnerability
A methodological caveat is warranted before interpreting these results. The present design uses one representative algorithm for the isolation-based paradigm (IF) and two for the density/reconstruction-based paradigm (GMM and AE). Strictly speaking, a single representative per paradigm does not permit a formal statistical claim that the behaviour is determined by the paradigm rather than by the individual algorithm. The term "paradigm-dependent" is therefore used in this paper to denote an observed paradigm-level tendency: the two density/reconstruction algorithms behave consistently with each other and differently from the isolation-based algorithm across all three datasets and all robustness settings. While this consistency between two independent algorithms strengthens the interpretation, confirming it as a general paradigm property would require additional algorithms from each family, which is identified as future work. The findings should therefore be interpreted as evidence from representative algorithms rather than definitive proof for all algorithms within each paradigm.
A central finding of this study is that drift vulnerability seems to be strongly associated with the detection paradigm and not a particular individual algorithm. In all three of the datasets, IF showed either sensitivity inflation or blindness, whereas both GMM and AE exhibited drift-resistant behaviour. The fact that this same general pattern of stability is reproduced by two algorithmically distinct approaches is a strength in interpretation of the effect as paradigm linked rather than one specific to the model alone.
This contrast is largest on Edge-IIoTset, where the same environment results in total blindness in the isolation-based method, and near perfect detection in the density- and reconstruction-based methods. The baseline comparison and multi-seed replication provide evidence that this conclusion reflects a robust, structural relationship with the evaluated data environment.
One possible explanation can be found in how the paradigms draw the boundaries of anomalies. IF is based on recursive random partitioning. If the attack samples form dense regions which match benign partitions, they may not be isolated as anomalous. As drift causes a shift in the feature distributions these partitions can become more and more misaligned, and this results in the progressive threshold-crossing behaviour that appears as sensitivity inflation. By contrast, GMM and AE define anomalies with respect to a learned representation of benign traffic, either in the form of density estimation or fidelity reconstruction. Because these boundaries are based on the statistical characteristics of benign data, moderate additive noise may not change the relative positions of benign and attack data samples with respect to the decision boundary to a large extent. This explanation is still suggestive, not definitive, as the present study does not involve a causal feature-level analysis.
The practical implication is that IDS evaluation should not be based on one algorithm evaluated against one dataset. Doing so only offers limited assurance with regard to deployment behaviour. A more reliable strategy is matrix-like evaluation where more than one paradigm is tested in more than one environment, and the resulting behaviour patterns are compared explicitly.
A quantitative account of why the same algorithm behaves differently across datasets follows from the geometry of the benign training region relative to the attack distribution. For IF, the expected path length, and hence the anomaly score, depends on how separable benign points are under random axis-aligned splits. In CIC-IoT-2023 and UNSW-NB15, the benign traffic is located in a relatively small region. As the drift increases, more and more points move beyond the fixed threshold, causing the sensitivity inflation (recall increases from 0.149 in the pre-drift segments to 0.384 in the final segment on CIC-IoT-2023, and from 0.190 to 0.397 on UNSW-NB15, with FPR rising accordingly). In Edge-IIoTset, on the other hand, the distributions of the benign and attack classes are heavily overlapping at the selected contamination level, and isolation assigns very similar scores to both classes: the 99th percentile threshold is then above almost all attack scores, resulting in recall ≈ 0 (immediate blindness) irrespective of drift. The key is thus the distance between the benign score distribution and the attack score distribution at the time of calibration, where a positive but small distance means drift will cause the sensitivity to be inflated, and where a near-zero distance means the detector will be blind from the start. Both failure modes are avoided by density- and reconstruction-based detectors, whose scores are based on the likelihood (GMM) or reconstruction error (AE) of the benign manifold, and which are only marginally affected by the additive drift used here, which explains their near-constant SV (below 0.0005) on all three datasets.
5.3 The limitations of F1-score under high attack rates
The CIC-IoT-2023 results show how the F1-score can be a poor proxy for operational quality when the attack rate is very high. For IF, F1-score increases significantly over the evaluation time, making it seem that it is working better. However, the corresponding increase in FPR means that the detector is also becoming less selective, and this makes the apparent gain less useful.
The mechanism can be formalized as follows. Let π denote the attack rate. Precision can then be expressed as a function of recall and the FPR, as shown in Equation 15.
Precision = (π · Recall) / (π · Recall + (1 - π) · FPR) (15)
When π is very large, the part of the denominator only involving attack dominates. Even large increases in FPR may have only a limited effect on precision as the weight of the contribution from false positives is scaled by the small factor (1 - π). The sensitivity of F1-score with respect to change in FPR can therefore be described by Equation 16.
dF1/dFPR ∝ -(1 - π) / [π · Recall + (1 - π) · FPR]^2 (16)
As π approaches 1 the factor (1 − π) approaches 0 so F1-score becomes less and less sensitive to changes in false alarms. This is the reason that F1-score can be a misleading measure for the degradation of operations in highly imbalanced settings. UNSW-NB15 provides a useful comparative benchmark because its more balanced class distribution makes the F1-score more sensitive to simultaneous changes in recall and the FPR. Although this does not make the F1-score an ideal metric for temporal behavioural evaluation, it demonstrates that its limitations are dataset-dependent rather than universal. Specifically, the shortcomings of the F1-score become considerably more pronounced under conditions of extreme class imbalance, where substantial behavioural changes may be obscured despite significant variations in detection performance.
The DS addresses this limitation directly by placing attack detection and false alarms in a symmetric manner at the segment level. What is needed here is neither mathematical novelty nor a new family of methods, but a temporally interpretable view of the model's operational usefulness. For this reason, temporal IDS evaluations should explicitly report recall and FPR for segments, in addition to a discrimination-oriented measure, instead of just aggregate F1-score.
5.4 Relation to established drift and stability metrics
The proposed behavioural indices are intentionally simple, interpretable, and complementary to existing drift-monitoring approaches rather than replacements. (DS = recall − FPR) corresponds to Youden's J statistic [33] computed for each temporal segment; its novelty lies in its temporal application rather than the underlying formulation. Existing drift detection method [13] and ADWIN [21], identify changes in the error stream of supervised models but provide limited insight into the resulting detector behaviour. Similarly, distribution-based measures, including the population stability index (PSI), Kullback–Leibler (KL) divergence, and the Kolmogorov–Smirnov (KS) statistic used in EMNCD [14], quantify changes in the input distribution but do not indicate whether the detector remains operationally effective. In contrast, conventional performance metrics such as AUROC, AUPRC, and F1-score summarize discrimination at a single operating point and, as demonstrated in Section 5.3, may obscure substantial behavioural changes under severe class imbalance. The proposed indices address this gap by explicitly characterizing detector behaviour over time. (BI = 1 − mean recall) quantifies missed detections, AFI = mean FPR measures false-alarm accumulation, SV captures temporal fluctuations in recall, and DS provides a threshold-dependent measure that jointly reflects detection capability and false-alarm behaviour across temporal segments.
These indices have a number of operational benefits. They are not supervised error stream based like DDM and ADWIN, and can be used for unsupervised anomaly detectors. Furthermore, they assess the detector's response to the distribution shift, rather than its size, allowing the operators to detect changes to the distribution from a shift that would impact detection performance but not from a benign shift. The indices also localise the onset of behavioural degradation, rather than averaging over the evaluation period, as the indices are computed for each temporal segment. BI and DS need to have labelled data for offline evaluation, while AFI and the temporal alarm-rate trajectory can be monitored in an ongoing manner without labels in the deployment phase, which makes them suitable for continuous monitoring.
The primary contribution of this study is not a direct comparison with accuracy of existing systems since the goal of this study is to characterize the behavioural response to concept drift and not introduce a new concept drift detection algorithm. Instead, the key contribution lies in providing a cross-paradigm, cross-dataset behavioural evaluation framework that reveals detector stability, blindness, and sensitivity inflation under evolving data distributions.
5.5 Implications for IDS deployment
These findings have a number of practical implications for the deployment of IDS in dynamic environments. First, the immediate-blindness pattern indicates that continuous monitoring must not rely on alarm volume alone. A detector that is producing very few alarms may be representing a quiet environment, but it may also be representing complete failure. Verification mechanisms should therefore include periodic checks that the detector is still producing meaningful true positives. Second, the repeated sensitivity inflation observed in IF implies that the isolation-based detectors may need regular recalibration when deployed in environments subject to distributional shift. The fixed-threshold design used here was deliberately chosen to isolate the static-calibration setting, but in practice adaptive thresholding or scheduled recalibration may alleviate the progressive loss of discrimination. Third, the drift resistance exhibited by GMM and AE hints that the use of density and reconstruction-based methods may provide gains in deployment stability. Meanwhile, there is no correlation between stability and absolute performance. Some operating conditions may require a more stable detector with a moderate recall. Stability and detection strength should, therefore, be weighed when making deployment decisions.
Fourth, there is the complementary paradigms in parallel, which means there is a strong interaction between algorithm and environment that could lead to an increase in resilience. A combination of isolation and density/reconstruction detectors could minimize the chance that one change in the environment will render the whole detection system useless. Overall, the findings indicate that as well as being accurate, the models need to be diverse if they are to be useful in dynamic security applications.
5.6 Limitations
However, there are some caveats to the extensive evaluation. The study only explores three anomaly detection algorithms, which belong to two detection paradigms. While similar behavior was seen in several sets of data and drift cases, it would be good to validate the paradigm level conclusions with further testing of other supervised, semi-supervised, ensemble, and more advanced density and reconstruction-based models. Second, although four types of controlled drift (Gaussian drift, uniform drift, spike drift, covariate shift drift) were studied, prior-probability shift and adversarial drift are not studied. Thirdly, the Edge-IIoTset dataset has only two mixed-class temporal segments, thus restricting the statistical power of mixed-segment analyses in spite of multi-seed validation using CI. In addition, the experiments had a fixed decision threshold, and did not retrain to investigate the behavioural effects of concept drift.
Although threshold-sensitivity analysis confirmed that the relative ranking of the algorithms remained unchanged, practical IDSs commonly employ adaptive thresholding, retraining, or continual learning strategies, which should be investigated in future work. Finally, the proposed DS metric is mathematically equivalent to Youden's J statistic; therefore, the contribution of this work lies not in introducing a new metric but in its systematic temporal application, together with complementary behavioural indices, for analysing detector stability, blindness, and sensitivity inflation under evolving data distributions. A complete list of abbreviations is listed in Appendix I.
6. Conclusion and future work
This paper investigated the behavioural instability of anomaly-based IDSs under concept drift using a unified experimental pipeline across three datasets and three anomaly detection algorithms. The results indicate that vulnerability to concept drift exhibits a paradigm-level tendency in the evaluated environments. In particular, isolation-based detection is consistently susceptible to sensitivity inflation and blindness, whereas density-based and reconstruction-based detection remain comparatively stable across the tested environments. The findings also demonstrate that aggregate performance measures, such as the F1-score, can obscure operationally significant behavioural changes, particularly under high attack rates. In contrast, a temporal evaluation framework based on recall, FPR, and the DS provides a more deployment-relevant assessment of IDS reliability by revealing behavioural dynamics that are not captured by aggregate metrics alone. More broadly, the results suggest that evaluating IDSs in dynamic environments should extend beyond determining whether overall performance declines. It should also examine how detector behaviour evolves over time, whether failure emerges gradually or abruptly, and whether the detector remains operationally reliable throughout deployment.
Future work should extend this behavioural evaluation framework to additional anomaly detection algorithms, more diverse and realistic concept drift scenarios, adaptive learning strategies, and real-world deployment environments.
Acknowledgment
The authors would like to express their sincere gratitude to the University of Information Technology and Communications, Baghdad, Iraq, and Multimedia University, Malaysia, for their valuable support.
Conflicts of interest
The authors have no conflicts of interest to declare.
Data availability
The datasets used in this study—CIC-IoT-2023, Edge-IIoTset, and UNSW-NB15—are publicly available from their respective repositories: CIC-IoT-2023 (https://www.unb.ca/cic/datasets/iotdataset-2023.html), Edge-IIoTset (https://www.kaggle.com/datasets/mohamedamineferrag/edgeiiotset-cyber-security-dataset-of-iot-iiot), and UNSW-NB15 (https://research.unsw.edu.au/projects/unsw-nb15-dataset). The complete experimental code, including the Google Colab notebooks required to reproduce all tables and figures presented in this study, is publicly available on GitHub at https://github.com/mohammadrasheed-max/ids-behavioural-drift.
Author's contribution statement
Mohammad M. Rasheed: Conceptualization, investigation, data curation, writing original draft, writing review and editing. Mustafa Muwafak Alobaedy: Study conception, design, supervision, investigation, draft manuscript preparation.
References
[1] Thakkar A, Lohiya R. A review on challenges and future research directions for machine learning-based intrusion detection system. Archives of Computational Methods in Engineering. 2023; 30(7):4245-69.
[2] Morshedi R, Matinkhah SM. A comprehensive review of deep learning techniques for anomaly detection in IoT networks: methods, challenges, and datasets. Engineering Reports. 2025; 7(9):1-29.
[3] Mallidi SK, Ramisetty RR. Optimizing intrusion detection for IoT: a systematic review of machine learning and deep learning approaches with feature selection and data balancing. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery. 2025; 15(2): e70008.
[4] Sommer R, Paxson V. Outside the closed world: on using machine learning for network intrusion detection. In symposium on security and privacy 2010 (pp. 305-16). IEEE.
[5] Arp D, Quiring E, Pendlebury F, Warnecke A, Pierazzi F, Wressnegger C, et al. Dos and don'ts of machine learning in computer security. In 31st USENIX security symposium 2022 (pp. 3971-88).
[6] Gama J, Žliobaitė I, Bifet A, Pechenizkiy M, Bouchachia A. A survey on concept drift adaptation. ACM Computing Surveys. 2014; 46(4):1-37.
[7] Shyaa MA, Ibrahim NF, Zainol Z, Abdullah R, Anbar M, Alzubaidi L. Evolving cybersecurity frontiers: a comprehensive survey on concept drift and feature dynamics aware machine and deep learning in intrusion detection systems. Engineering Applications of Artificial Intelligence. 2024; 137:1-34.
[8] Jordaney R, Sharad K, Dash SK, Wang Z, Papini D, Nouretdinov I, et al. Transcend: detecting concept drift in malware classification models. In 26th USENIX security symposium 2017 (pp. 625-42).
[9] Pendlebury F, Pierazzi F, Jordaney R, Kinder J, Cavallaro L. {TESSERACT}: eliminating experimental bias in malware classification across space and time. In 28th USENIX security symposium 2019 (pp. 729-46).
[10] Andresini G, Pendlebury F, Pierazzi F, Loglisci C, Appice A, Cavallaro L. Insomnia: towards concept-drift robustness in network intrusion detection. In proceedings of the 14th ACM workshop on artificial intelligence and security 2021 (pp. 111-22). ACM.
[11] Yang S, Zheng X, Li J, Xu J, Wang X, Ngai EC. Recda: concept drift adaptation with representation enhancement for network intrusion detection. In proceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining 2024 (pp. 3818-28). ACM.
[12] Rasheed MM, Faaeq MK. Behavioral detection of scanning worm in cyber defense. In proceedings of the future technologies conference 2018 (pp. 214-25). Cham: Springer International Publishing.
[13] Gama J, Medas P, Castillo G, Rodrigues P. Learning with drift detection. In Brazilian symposium on artificial intelligence 2004 (pp. 286-95). Berlin, Heidelberg: Springer Berlin Heidelberg.
[14] Chu R, Jin P, Qiao H, Feng Q. Intrusion detection in the IoT data streams using concept drift localization. AIMS Mathematics. 2023; 9(1):1535-61.
[15] Shyaa MA, Zainol Z, Abdullah R, Anbar M, Alzubaidi L, Santamaría J. Enhanced intrusion detection with data stream classification and concept drift guided by the incremental learning genetic programming combiner. Sensors. 2023; 23(7):1-34.
[16] Xu L, Ding X, Peng H, Zhao D, Li X. ADTCD: an adaptive anomaly detection approach toward concept drift in IoT. IEEE Internet of Things Journal. 2023; 10(18):15931-42.
[17] Mohale VZ, Obagbuwa IC. Evaluating machine learning-based intrusion detection systems with explainable AI: enhancing transparency and interpretability. Frontiers in Computer Science. 2025; 7:1-23.
[18] Hashim A, Rasheed M, Abdullah S. Analysis of bluetooth low energy based indoor localization system using machine learning algorithms. Journal of Engineering Science and Technology. 2021; 16(4):2816-24.
[19] Barbero F, Pendlebury F, Pierazzi F, Cavallaro L. Transcending transcend: revisiting malware classification in the presence of concept drift. In IEEE symposium on security and privacy (SP) 2022 (pp. 805-23). IEEE.
[20] Haque A, Soliman H. A transformer-based autoencoder with isolation forest and XGBoost for malfunction and intrusion detection in wireless sensor networks for forest fire prediction. Future Internet. 2025; 17(4):1-12.
[21] Bifet A, Gavalda R. Learning from time-changing data with adaptive windowing. In proceedings of the SIAM international conference on data mining 2007 (pp. 443-8). Society for Industrial and Applied Mathematics.
[22] Bagui SS, Khan MP, Valmyr C, Bagui SC, Mink D. Model retraining upon concept drift detection in network traffic big data. Future Internet. 2025; 17(8):1-24.
[23] Hussein SA, Répás SR. A hybrid intrusion detection framework using deep autoencoder and machine learning models. AI. 2026; 7(2):1-31.
[24] Bachar M, Khiat A, El GK. Hybrid autoencoder and isolation forest for IoT anomaly detection with a novel model. Engineering, Technology & Applied Science Research. 2026; 16(1):31123-9.
[25] Seth S, Chahal KK, Singh G. Concept drift–based intrusion detection for evolving data stream classification in ids: approaches and comparative study. The Computer Journal. 2024; 67(7):2529-47.
[26] Horchulhack P, Viegas EK, Lopez MA. A stream learning intrusion detection system for concept drifting network traffic. In 6th cyber security in networking conference (CSNet) 2022 (pp. 1-7). IEEE.
[27] Wahab OA. Intrusion detection in the IoT under data and concept drifts: online deep learning approach. IEEE Internet of Things Journal. 2022; 9(20):19706-16.
[28] Neto EC, Dadkhah S, Ferreira R, Zohourian A, Lu R, Ghorbani AA. CICIoT2023: a real-time dataset and benchmark for large-scale attacks in IoT environment. Sensors. 2023; 23(13):1-26.
[29] Ferrag MA, Friha O, Hamouda D, Maglaras L, Janicke H. Edge-IIoTset: a new comprehensive realistic cyber security dataset of IoT and IIoT applications for centralized and federated learning. IEEE Access. 2022; 10:40281-306.
[30] Moustafa N, Slay J. UNSW-NB15: a comprehensive data set for network intrusion detection systems (UNSW-NB15 network data set). In military communications and information systems conference (MilCIS) 2015 (pp. 1-6). IEEE.
[31] Liu FT, Ting KM, Zhou ZH. Isolation forest. In eighth international conference on data mining 2008 (pp. 413-22). IEEE.
[32] Reynolds DA. Gaussian mixture models. Encyclopedia of Biometrics. 2009; 741(659-63):3.
[33] Youden WJ. Index for rating diagnostic tests. Cancer. 1950; 3(1):32-5.
Appendix I
S. No. | Abbreviation | Description
1 | AE | Autoencoder
2 | AFI | Alarm Fatigue Index
3 | AUC | Area Under the Curve
4 | AUPRC | Area Under the Precision–Recall Curve
5 | AUROC | Area Under the Receiver Operating Characteristic Curve
6 | BI | Blindness Index
7 | CI | Confidence Interval
8 | CIC | Canadian Institute for Cybersecurity
9 | DS | Detection Stability
10 | FPR | False Positive Rate
11 | GMM | Gaussian Mixture Model
12 | IDS | Intrusion Detection System
13 | IF | Isolation Forest
14 | IIoT | Industrial Internet of Things
15 | IoT | Internet of Things
16 | SD | Standard Deviation
17 | SV | Stability Variance
Automatically extracted. Refer to the original PDF for figures, tables, and formatting.