What do you do, in one sentence?
I build privacy-aware NLP and grounded AI systems, from dataset and evaluation design through APIs, mobile apps and production support.
About Mohamed Hafez
I’m a Toronto-based engineer who enjoys taking complex ideas from research and architecture through implementation. My work spans AI/ML, software systems, Android products, and technical teaching.
Current status
Education
I completed my M.A. in Information Systems & Technology at York University, and it was officially awarded in 2026.
Straight answers
Short answers to what recruiters and hiring managers ask most. The portfolio assistant quotes these exactly and labels them as my answers.
I build privacy-aware NLP and grounded AI systems, from dataset and evaluation design through APIs, mobile apps and production support.
The common thread is end-to-end ownership: I take a problem from the question through data, architecture, build, evaluation and documentation, and then explain it. The strongest story is not that I’ve worked in many fields; it’s that I repeatedly connect them into complete, understandable, reliable systems. My lead track is machine learning and NLP, with document AI and privacy as the focus. Backend, Android and teaching are the supporting range.
Machine learning and NLP engineering roles, especially where documents, personal data or evaluation are central: document AI, PII extraction, LLM evaluation and grounded RAG systems. I also take on Java backend and OpenText Documentum work, Android roles that involve ML or OCR, and technical teaching.
I currently work as a Founding AI Engineer at Eklan, and I’m open to new roles. My notice period is one week.
It depends on the scope, so I answer per track. Machine learning and AI: about four years, from DotPy (June 2021 to August 2023) and graduate research at York (September 2024 to April 2026), plus my current work at Eklan. Professional software engineering: two to three years. Android: five years of client work, from July 2019 to June 2024. These tracks overlap, and together they cover more than six years. I have not held a formal people-management role; at DotPy I planned work, coordinated delivery and mentored junior colleagues, with no direct reports.
I deposited my thesis in July 2026, had the SPRINT-PP paper accepted to IEEE CASCON 2026, built Northstar and this portfolio, and joined Eklan as a Founding AI Engineer.
I moved to Toronto in September 2024 for my master’s at York University. I hold an open Post-Graduation Work Permit valid through June 2029, so I can work for any Canadian employer without sponsorship.
I build end-to-end NLP pipelines, from dataset construction and validation through fine-tuning, inference and evaluation. At York I built ASC-PIE, 333,109 examples under one 19-type PII schema, and benchmarked encoder, encoder-decoder and decoder-only models; RoBERTa-large reached 0.991 strict F1. I also developed SPRINT-PP, which stops a model forgetting earlier PII types without storing personal data. Before that I delivered more than ten applied ML projects at DotPy and integrated models into APIs and Android apps.
I hold both halves of the document-intelligence job: research on extracting PII from unstructured text, and more than a year building Documentum workflows, DQL validation and Java PDF pipelines for banking clients, where every document carries personal data. I know where extraction models fail and where document workflows break.
I evaluate and build LLM systems. In my thesis, 3-shot key-value prompting lifted Qwen2.5-7B from 0.400 fine-tuned to 0.778 strict F1, at 1.76 seconds per query against 26.8 seconds for JSON prompting, so protocol design mattered as much as the model. I also built Northstar, a grounded RAG service that cites document, page and chunk and refuses when retrieval evidence is weak.
My thesis is an evaluation framework: one schema, one parser and one scorer across three model families, with output validity as a first-class metric next to strict and normalized F1, latency, robustness and continual-learning metrics.
I built Java and Python services, REST endpoints and iText PDF utilities on Tomcat inside regulated Documentum workflows, with structured logging, runbooks and root-cause analysis in production. My ML and mobile background helps me integrate services across systems.
I implemented OpenText Documentum workflows and lifecycles for document intake, review, routing, archival and retention for banking clients, with DQL validation of metadata and retention, Java and iText PDF services, and production support.
I have taught since 2019: teaching assistant at York for Systems Architecture and Data Visualization, more than 200 hours of private instruction for students from six institutions, and more than 15 workshops for over 100 attendees, with problem sets, mock exams and auto-graded labs.
Much of my work has sat between the model and the user: client proofs of concept and REST contracts at BASS, model-serving contracts with mobile and backend teams at DotPy, and years of explaining hard material to learners. I can scope, build and explain.
Mind’s Eye, my undergraduate capstone: assistive smart glasses for people who are blind or visually impaired and people living with Alzheimer’s. In a five-person team I delivered more than 80% of the system: the Android app, the ESP32-CAM link over Bluetooth, the cloud vision integrations, and OCR with spoken feedback. No service recognised Egyptian banknotes, so I built that recogniser myself. It earned an A+. What I took from it is that assistive AI only works when capture, latency, errors and spoken feedback are designed as one system.
My thesis built ASC-PIE. I unified five PII data sources, including my own synthetic augmentation, into a 333,109-example corpus under one 19-type schema and one scoring harness, so encoder, encoder-decoder and decoder-only models could be compared fairly. RoBERTa-large reached 0.991 strict F1, 3-shot key-value prompting lifted Qwen2.5-7B from 0.400 to 0.778, and corpus design alone moved strict F1 from 0.415 to 0.992. Then I worked on catastrophic forgetting: naive updates wiped earlier PII types to zero, so I built SPRINT-PP, which keeps prototypes instead of personal data. It reached 83.5% average accuracy, the best of the four strategies I compared, with 25.6% less training time than distillation.
No. It is replay-free: it stores class prototypes and classifier-head snapshots, never raw text or personal data. That is an operational privacy property, not a formal guarantee such as differential privacy.
The corpus builder deduplicates within each split, removes any validation or test example that also appears in training, and runs a second leakage pass after balancing. The splits are fixed and disjoint.
I measure the things that can go wrong separately: structural validity of the output, correctness under strict and normalized matching, per-category performance, robustness to realistic perturbations, and latency. For retrieval systems I add faithfulness and context metrics, and a refusal path for weak evidence.
At DotPy each project moved from scoping and data preparation through feature engineering, model selection and evaluation to delivery as a notebook, an API service or a Dockerized prototype, with documentation and handover. A later example is the dive planner, a client proof of concept: four classifiers compared offline, with the chosen model served to the Android app through a POST endpoint.
With suspicion. In my network intrusion detection project, random forest and gradient boosting reached near-perfect ROC-AUC. I flagged likely separable features or leakage and recommended time-based or cross-capture validation before trusting the result.
Several things, and each one changed the design. Naive sequential fine-tuning destroyed earlier PII classes, which led to SPRINT-PP. Fine-tuned decoder models in the 7 to 8 billion parameter range underperformed prompting, which led to the in-context protocol study. Mind’s Eye responses were slow on cheap servers, and the fix we identified was resizing images and using faster services. My early per-dataset baselines were not comparable, which is why ASC-PIE has one schema and one scorer.
My title was AI & ML Specialist. I planned work, coordinated delivery and mentored junior colleagues, with no direct reports.
No. I completed a CEH training programme at Inspire Academy in 2022; I did not take the certification exam.
No. It is 99.5% offline test accuracy for a decision tree. The app is a planning aid, not a safety device.
Contact
Open to new roles, one-week notice period · Toronto, Ontario, Canada