Software · AI/ML · 3 min read

This Portfolio

The site you’re reading: four focuses, one grounded assistant that refuses rather than guesses.

I built this site: a four-focus engineering portfolio with typed, validated content and a grounded AI assistant that refuses rather than guesses.

My role
I designed and built the site end to end: the content architecture, the role switcher, the case-study system, and the grounded assistant.
Context
Personal project, live
Status
Live · evolving

Live · evolving

  • Astro
  • TypeScript
  • Cloudflare Workers
  • Workers AI
  • Zod
  • Playwright

Roles: Software Engineer · AI/ML Engineer.

The problem

A background that spans AI/ML, software, Android, and teaching can read as unfocused if it’s presented as one flat list. A hiring manager scanning for machine-learning work and a hiring manager scanning for Android work shouldn’t have to dig through the same undivided page to find what they need. I wanted a site that could be four different portfolios depending on who was reading it, without maintaining four separate sites.

Who it was for

Anyone evaluating my work from a specific angle: a hiring manager, a collaborator, or a student comparing notes on a project. The role switcher exists so each of them can land on the version of the site organized around what they actually came to look at.

My role

I designed and built the site end to end: the content architecture, the role switcher, the case-study system, and the grounded assistant.

What I built

This site: a four-focus portfolio built on typed, validated content, with a small assistant that answers questions about the work instead of just displaying it.

  • Four-focus role switcher

    A visitor picks AI/ML, software, Android, or teaching, and the same content reorders around that focus instead of showing one generic profile.

  • Typed, validated content

    Every project, role, and case study is defined as a typed, Zod-validated record, so a bad or missing field fails the build instead of shipping.

  • Pre-rendered pages

    Pages are pre-rendered at build time, so the site works fully before any client-side script runs.

  • A grounded assistant

    An in-page assistant answers questions about the work by retrieving from the content published on the site, then refuses rather than guessing when nothing supports an answer.

The assistant is deliberately conservative. A question comes in behind Turnstile and two rate limiters, then goes through lexical retrieval over the published, validated content. If retrieval doesn’t turn up anything that supports an answer, the pipeline refuses before it ever reaches generation. When it does have support, Workers AI (Llama 3.1 8B) writes the answer, its citations are checked against what was actually retrieved, and if that check fails the assistant falls back to an extractive answer built directly from the retrieved text instead of a generated one. Conversation history stays in the browser session only.

Analytics on how the assistant is used is not live today; the pipeline is built to support it later without storing raw transcripts, but nothing is currently collecting or reporting that data.

How it works

The static site pre-renders from the same published content the assistant retrieves from, so the two halves of the project share one source of facts but run on independent paths.

Typed, validated content pre-renders the site, while a separate assistant path retrieves over the same published content, refuses before generation, and validates citations before answering. Relationships: Published content to Astro pages; Question to Turnstile + rate limits; Turnstile + rate limits to Retrieval; Retrieval to Refusal gate; Refusal gate to Workers AI; Workers AI to Citation check; Citation check to Answer; Citation check to Extractive fallback.

Decisions that mattered

  • I chose a zero-budget-first Cloudflare stack over a paid hosting and inference setup because the assistant and the site both had to be sustainable to run indefinitely on a personal budget, not just cheap to demo once.

  • I chose keeping the static site fully useful if the assistant fails over building the assistant as a required layer over the content because an AI feature that visitors depend on to read the portfolio at all is a fragile design; the pages have to stand on their own.

  • I chose no raw-transcript analytics over logging full visitor questions and answers because understanding how the assistant is used shouldn't require storing what people typed.

Tech stack

Site
  • Astro: Pre-rendered pages and content architecture
  • TypeScript: Site and content typing
  • Zod: Runtime content validation
  • React: Interactive islands
  • MDX: Case-study content
Assistant runtime
  • Cloudflare Workers: Assistant request handling
  • Workers AI: Answer generation (Llama 3.1 8B)
  • Turnstile: Bot and abuse protection
Quality
  • Vitest: Unit tests
  • Playwright: End-to-end tests

Outcome

The site runs today as the live, evolving thing you’re reading: a pre-rendered portfolio that stays fully readable without JavaScript, plus an assistant layer that answers from the published content and says “I don’t have enough to answer that” rather than filling the gap.

What I learned

An assistant is only as trustworthy as its willingness to refuse.

Building the refusal path before the generation path changed how I thought about the whole feature: the interesting engineering problem wasn’t getting a model to answer, it was building a system that knew when it shouldn’t.

What I’d do next

The retrieval layer works on the published content that exists today; the natural next step is broadening what’s published rather than changing how retrieval or refusal work. I’d also like to add the usage analytics the pipeline was designed for, on aggregate counts rather than raw transcripts, once I’m confident the privacy trade-offs are right.

The whole site, including the assistant and its tests, is public.