Skip to content

AI engineering · built in the open

A portfolio you can investigate

The assistant answers professional questions, assesses job requirements and inspects public project code. Its answers come with source evidence, and its execution can be inspected directly.

Try the agent → · Read the implementation

From question to evidence-backed answer

1 · Understand

Validate browser messages, check Turnstile, enforce a shared request budget, and include the previous exchange when retrieving context.

2 · Retrieve

Embed the query with Workers AI. Rank the small versioned corpus in memory using cosine similarity and category boosts. Job assessments inspect the published profile evidence.

3 · Inspect

For project questions, the model can read allowlisted public repository files. Each result links to the exact commit inspected. Tool calls and model rounds are bounded.

4 · Respond

Stream cited answers, or validate a structured role assessment against exact source quotations. Display observable timings, usage and versions in optional engineering details.

What the system establishes

A role assessment distinguishes direct evidence, transferable experience and requirements that are not established. CV statements are labeled separately from public code. A missing skill in the evidence is not proof that someone lacks it.

Structured output is validated beyond its JSON shape: each quoted passage must exist in the cited source. That catches fabricated quotations; evaluating whether the evidence actually supports the claim still needs a semantic rubric and human review.

Measured evaluation results

Loading the latest measured results…

The versioned evaluation dataset cover professional questions, requirement assessments, follow-ups and unsupported claims. The owner can run evaluations on demand in GitHub Actions. Full reports are retained in private workflow artifacts; this page publishes aggregate results. The model grader and evaluation rubric are publicly inspectable.

Read the measured engineering case study for the first real-provider run, observed failures and the changes they drove.

Design decisions and limits