# Portfolio conversation evaluation The dataset is a small curated regression suite, not a representative benchmark. Role scenarios paraphrase recurring requirements from the employer postings in `docs/ai-engineering-skills-roadmap.md`; they are not endorsements or copied ads. ## Checks The runner calls the deployed Worker, including real embedding and model APIs, retrieval, tools, source validation and shared rate limits. It verifies complete SSE execution, version identifiers, model/tool budgets, source citation validity, expected source coverage and explicitly unknown skills in role assessments. Exact wording and a particular tool sequence are not required. Optional model grading separately checks factual grounding, answering the question and honest evidence gaps. It receives the actual cited source text, the answer, structured assessment and scenario rubric. Judge model, rubric version, tokens and reported cost are stored separately. A scenario passes only if its workflow checks and enabled semantic grading pass. The table exposes both results. The judge is an automated heuristic, **not human-calibrated**. False positives and false negatives require review. No independent/human quality validation is claimed. Human calibration is a future step: label a varied sample, include intentionally unsupported answers, measure judge agreement, document disagreements and revise the rubric on a held-out set. It cannot be completed honestly by an automated agent. Multi-turn scenarios can also run with latest-message-only retrieval. Baseline rows preserve the same conversation and prompts except the retrieval query; they are excluded from the main pass-rate and do not fail the job. This is an ablation of conversation context, not a recreation of a historical production release. A single stochastic comparison does not establish a statistically significant gain. ## Reproduce Run Actions → Evaluate portfolio agent → Run workflow (main). It uses the existing Worker model credentials and a separate evaluation secret. Local execution: `RAG_EVAL_SECRET=... npm run evaluate --prefix worker` (Node 22+). A reported-cost guard defaults to $0.50 and stops between scenarios; a single in-flight call can cross it and unreported costs cannot be counted. The provider key spending limit remains the hard cap. Set `EVAL_MAX_REPORTED_COST_USD` to change the guard. Set `EVAL_SUITE=smoke`, `SEMANTIC_JUDGE=false` or `EVAL_BASELINE=false` to limit a run. `RAG_API_URL` defaults to the production Worker; never point a run at an untrusted URL. Full reports contain only curated evaluation conversations, not visitor sessions. They are downloadable run artifacts for 30 days. The public scorecard stores only scenario names/outcomes/timings and version metadata. Corpus generation remains an ephemeral build with no automatic Git write-back. The optional model comparison runs the same two assessment scenarios against Haiku 4.5, GPT-5.4 Mini, Gemini 3.8 Flash and Qwen3.7 Plus, with identical evidence, validators and one bounded correction attempt. The grader stays Haiku 4.5 for every candidate; self-grading bias for the Haiku baseline and lack of human calibration are limitations. Results are private artifacts and do not replace the deployed-model scorecard. Two cases per model are a screening sample, not a reliable estimate of general accuracy. Reasoning effort and actual token/cost usage must be considered when comparing prices.