About Mohamed Hafez

Engineering across AI, software, mobile, and teaching

I’m a Toronto-based engineer who enjoys taking complex ideas from research and architecture through implementation. My work spans AI/ML, software systems, Android products, and technical teaching.

Current status

Practical details

Location
Toronto, Ontario, Canada.
Availability
Currently working as Founding AI Engineer at Eklan. Open to new roles, with a one-week notice period.
Work authorization
Open PGWP valid through June 2029.
Sponsorship
Legally authorized to work in Canada; no sponsorship required.
Work arrangement
Open to onsite, hybrid, and remote roles in Canada.
Relocation
Open to relocation within Canada and the GTA.
Target level
Targeting intermediate-level opportunities.
Employment preference
Full-time is preferred; contract opportunities are open.

Education

York University

I completed my M.A. in Information Systems & Technology at York University, and it was officially awarded in 2026.

Straight answers

Questions I’m often asked

Short answers to what recruiters and hiring managers ask most. The portfolio assistant quotes these exactly and labels them as my answers.

About me

What do you do, in one sentence?

I build privacy-aware NLP and grounded AI systems, from dataset and evaluation design through APIs, mobile apps and production support.

You have worked in ML, backend, Android and teaching. Why so many tracks?

The common thread is end-to-end ownership: I take a problem from the question through data, architecture, build, evaluation and documentation, and then explain it. The strongest story is not that I’ve worked in many fields; it’s that I repeatedly connect them into complete, understandable, reliable systems. My lead track is machine learning and NLP, with document AI and privacy as the focus. Backend, Android and teaching are the supporting range.

What kind of role are you looking for?

Machine learning and NLP engineering roles, especially where documents, personal data or evaluation are central: document AI, PII extraction, LLM evaluation and grounded RAG systems. I also take on Java backend and OpenText Documentum work, Android roles that involve ML or OCR, and technical teaching.

Are you working now, and when could you start?

I currently work as a Founding AI Engineer at Eklan, and I’m open to new roles. My notice period is one week.

How many years of experience do you have?

It depends on the scope, so I answer per track. Machine learning and AI: about four years, from DotPy (June 2021 to August 2023) and graduate research at York (September 2024 to April 2026), plus my current work at Eklan. Professional software engineering: two to three years. Android: five years of client work, from July 2019 to June 2024. These tracks overlap, and together they cover more than six years. I have not held a formal people-management role; at DotPy I planned work, coordinated delivery and mentored junior colleagues, with no direct reports.

What have you done since finishing your degree in April 2026?

I deposited my thesis in July 2026, had the SPRINT-PP paper accepted to IEEE CASCON 2026, built Northstar and this portfolio, and joined Eklan as a Founding AI Engineer.

Why did you move to Canada?

I moved to Toronto in September 2024 for my master’s at York University. I hold an open Post-Graduation Work Permit valid through June 2029, so I can work for any Canadian employer without sponsorship.

My background, by role

What is your background for a machine learning or NLP engineering role?

I build end-to-end NLP pipelines, from dataset construction and validation through fine-tuning, inference and evaluation. At York I built ASC-PIE, 333,109 examples under one 19-type PII schema, and benchmarked encoder, encoder-decoder and decoder-only models; RoBERTa-large reached 0.991 strict F1. I also developed SPRINT-PP, which stops a model forgetting earlier PII types without storing personal data. Before that I delivered more than ten applied ML projects at DotPy and integrated models into APIs and Android apps.

What is your background for document AI or privacy work?

I hold both halves of the document-intelligence job: research on extracting PII from unstructured text, and more than a year building Documentum workflows, DQL validation and Java PDF pipelines for banking clients, where every document carries personal data. I know where extraction models fail and where document workflows break.

What is your background for an AI or LLM engineering role?

I evaluate and build LLM systems. In my thesis, 3-shot key-value prompting lifted Qwen2.5-7B from 0.400 fine-tuned to 0.778 strict F1, at 1.76 seconds per query against 26.8 seconds for JSON prompting, so protocol design mattered as much as the model. I also built Northstar, a grounded RAG service that cites document, page and chunk and refuses when retrieval evidence is weak.

What is your background for model or LLM evaluation work?

My thesis is an evaluation framework: one schema, one parser and one scorer across three model families, with output validity as a first-class metric next to strict and normalized F1, latency, robustness and continual-learning metrics.

What is your background for a software or backend engineering role?

I built Java and Python services, REST endpoints and iText PDF utilities on Tomcat inside regulated Documentum workflows, with structured logging, runbooks and root-cause analysis in production. My ML and mobile background helps me integrate services across systems.

What is your OpenText Documentum and ECM experience?

I implemented OpenText Documentum workflows and lifecycles for document intake, review, routing, archival and retention for banking clients, with DQL validation of metadata and retention, Java and iText PDF services, and production support.

What is your teaching background?

I have taught since 2019: teaching assistant at York for Systems Architecture and Data Visualization, more than 200 hours of private instruction for students from six institutions, and more than 15 workshops for over 100 attendees, with problem sets, mock exams and auto-graded labs.

Have you worked directly with clients or customers?

Much of my work has sat between the model and the user: client proofs of concept and REST contracts at BASS, model-serving contracts with mobile and backend teams at DotPy, and years of explaining hard material to learners. I can scope, build and explain.

Projects and how I work

Which project are you most proud of?

Mind’s Eye, my undergraduate capstone: assistive smart glasses for people who are blind or visually impaired and people living with Alzheimer’s. In a five-person team I delivered more than 80% of the system: the Android app, the ESP32-CAM link over Bluetooth, the cloud vision integrations, and OCR with spoken feedback. No service recognised Egyptian banknotes, so I built that recogniser myself. It earned an A+. What I took from it is that assistive AI only works when capture, latency, errors and spoken feedback are designed as one system.

Tell me about your thesis.

My thesis built ASC-PIE. I unified five PII data sources, including my own synthetic augmentation, into a 333,109-example corpus under one 19-type schema and one scoring harness, so encoder, encoder-decoder and decoder-only models could be compared fairly. RoBERTa-large reached 0.991 strict F1, 3-shot key-value prompting lifted Qwen2.5-7B from 0.400 to 0.778, and corpus design alone moved strict F1 from 0.415 to 0.992. Then I worked on catastrophic forgetting: naive updates wiped earlier PII types to zero, so I built SPRINT-PP, which keeps prototypes instead of personal data. It reached 83.5% average accuracy, the best of the four strategies I compared, with 25.6% less training time than distillation.

Is SPRINT-PP a formal privacy guarantee?

No. It is replay-free: it stores class prototypes and classifier-head snapshots, never raw text or personal data. That is an operational privacy property, not a formal guarantee such as differential privacy.

How did you prevent data leakage in the benchmark?

The corpus builder deduplicates within each split, removes any validation or test example that also appears in training, and runs a second leakage pass after balancing. The splits are fixed and disjoint.

How do you evaluate an LLM feature?

I measure the things that can go wrong separately: structural validity of the output, correctness under strict and normalized matching, per-category performance, robustness to realistic perturbations, and latency. For retrieval systems I add faithfulness and context metrics, and a refusal path for weak evidence.

Walk me through how you deliver a model.

At DotPy each project moved from scoping and data preparation through feature engineering, model selection and evaluation to delivery as a notebook, an API service or a Dockerized prototype, with documentation and handover. A later example is the dive planner, a client proof of concept: four classifiers compared offline, with the chosen model served to the Android app through a POST endpoint.

How do you handle results that look too good?

With suspicion. In my network intrusion detection project, random forest and gradient boosting reached near-perfect ROC-AUC. I flagged likely separable features or leakage and recommended time-based or cross-capture validation before trusting the result.

Tell me about something that did not work.

Several things, and each one changed the design. Naive sequential fine-tuning destroyed earlier PII classes, which led to SPRINT-PP. Fine-tuned decoder models in the 7 to 8 billion parameter range underperformed prompting, which led to the in-context protocol study. Mind’s Eye responses were slow on cheap servers, and the fix we identified was resizing images and using faster services. My early per-dataset baselines were not comparable, which is why ASC-PIE has one schema and one scorer.

Straight answers

Were you a manager at DotPy?

My title was AI & ML Specialist. I planned work, coordinated delivery and mentored junior colleagues, with no direct reports.

Are you a Certified Ethical Hacker?

No. I completed a CEH training programme at Inspire Academy in 2022; I did not take the certification exam.

The dive planner reports 99.5% accuracy. Is that a safety claim?

No. It is 99.5% offline test accuracy for a decision tree. The app is a planning aid, not a safety device.

Contact

Let’s connect

Open to new roles, one-week notice period · Toronto, Ontario, Canada