AI/ML · Course project · 3 min read

Network Intrusion Detection Pipeline

Near-perfect scores are a reason to dig deeper, not to celebrate.

As a course project, I built a network intrusion-detection pipeline that compares four model families on benign versus DoS/DDoS flow data, and flagged its own near-perfect scores as a validation risk rather than a result to celebrate.

My role
I built the consolidation notebook, the cleanup and correlation-pruning steps, trained the four model families, and ran the diagnostics.
Context
Course project, network security
Period
2025
Status
Course project · 2025

Course project · 2025

  • ~63features retainedpruned from a larger set by removing pairs correlated above 0.95
  • 80/20train / test splitstratified across benign and attack traffic
  • 1.00best ROC-AUC (RF & GB)a near-perfect score flagged as possible leakage or overfitting, not a result to trust as-is
  • 0.67logistic regression ROC-AUCthe weaker linear baseline the ensembles are compared against
  • Python
  • pandas
  • scikit-learn
  • TensorFlow/Keras
  • Matplotlib
  • Seaborn

Roles: AI/ML Engineer.

The problem

This was a course project where I built a pipeline to tell benign network traffic apart from denial-of-service and distributed denial-of-service attacks, using flow-level data. Flow records carry identifiers, timestamps, and many measurements that describe the same underlying behaviour in different units, so a model trained on the raw columns can look accurate for the wrong reasons. The pipeline had to consolidate several capture files into one dataset, strip out what wasn’t a real predictive signal, and compare more than one kind of model before trusting any score it produced.

My role

I built the dataset-consolidation notebook that merged the benign and attack flow captures, wrote the cleanup and correlation-pruning steps, trained the model families, and produced the confusion-matrix and ROC diagnostics the project’s conclusions rest on. The evaluation choices — which features to prune, how to split the data, and which diagnostics to trust over a single headline score — were also mine to make and defend.

What I built

The pipeline moves in one direction: raw flow files in, a labelled dataset out, then a shrinking and more trustworthy feature set, then several models trained on the same split so their scores are comparable. Nothing downstream is allowed to see a row that a later stage will still reshape, which is what keeps the final comparison honest.

  • Dataset consolidation

    Combined the benign and DoS/DDoS flow captures into one labelled dataset.

  • Cleanup

    Dropped identifier and metadata columns and any zero-variance features.

  • Correlation pruning

    Removed features correlated above 0.95, keeping a smaller set of informative columns.

  • Model comparison

    Trained and compared Logistic Regression, Random Forest, Gradient Boosting, and a Keras ANN on the same stratified split.

  • Diagnostics

    Read confusion matrices and ROC curves instead of trusting a single accuracy number.

How it works

Each stage only does one job — combine, clean, prune, split, model, or diagnose — so a change in one step doesn’t quietly change what an earlier step already decided.

Raw flow captures are consolidated, cleaned, pruned by correlation, and compared across four model families before the near-perfect scores are put through diagnostics. Relationships: Flow CSVs to Consolidated dataset; Consolidated dataset to Cleanup; Cleanup to Correlation pruning; Correlation pruning to Split & scale; Split & scale to Models; Models to Diagnostics.

Tech stack

Diagnostics
  • Matplotlib: Confusion-matrix and ROC plots
  • Seaborn: Correlation and distribution plots

Outcome

On the held-out test split, Logistic Regression reached a ROC-AUC of about 0.67 — a reasonable linear baseline, and no more. Random Forest and Gradient Boosting both reached close to 1.00, and the Keras ANN reached about 0.99.

ROC-AUC by model family
ROC-AUC by model family
LabelValue
Logistic Regression0.67
Random Forest1
Gradient Boosting1
ANN0.99

Scores that close to a perfect 1.00, on flow data built from a handful of capture files, are not a result to celebrate. They read as a possible sign of leakage or overfitting: some feature may still be encoding the label indirectly, or the benign and attack traffic may simply be too separable in this particular capture window to say how the model would behave on new traffic. The project’s own conclusion was to treat the near-perfect tree-ensemble scores as a flag, and to follow them with time-based or cross-capture validation — testing on a different time period or a different capture source — before drawing any conclusion about real-world performance.

What I learned

Extremely high offline scores should trigger stronger validation, not immediate confidence.

Separating dataset construction from modelling also made the pipeline easier to audit and rerun: when a score looked suspicious, I could tell whether the problem sat in the data or in the model, instead of guessing.

What I’d do next

The next step is the one the diagnostics already called for: rerun the split by time or by capture source instead of a random shuffle, and see whether the tree-ensemble scores survive.