Academyby Dasow

Stop 9 of 10 · Weekly

Evaluating the assistant

Thirty questions with expected answers, a rubric that scores sourcing and refusal as well as correctness, one graded run, and fixes that go into the Project and the connector rather than the next prompt.

Watch first · 0:55

The assistant is now a system

By this stop Fadi has an account Project with instructions, a documentation connector, a repository and a set of prompts his two colleagues have started copying. That is a system, and nobody has ever measured whether it is right. The failure mode is not a bad answer that looks bad. It is a well-formatted, confidently sourced answer that cites a document that says something else.

Thirty questions, written once, run against the system, scored. It takes an afternoon and it is the only stop in this course that tells you whether the other nine worked.

The test set

Thirty questions across four groups, each with an expected answer written before the run.

Group Count What it tests
Factual 10 Service capability, quota and control questions with a known right answer
Judgment 10 Design and compliance questions with a defensible answer and a named trade
Refusal 10 Questions the system should refuse or decline to source, split between pricing, customer data and things the library does not cover

Write the expected answers first, from documents, before you see what the system says. Written afterwards they drift towards whatever came back.

Draft the test set
I am the cloud architect on a public-sector team. Help me build a thirty-question evaluation set for the account Project and documentation connector I use for state and local agency work.

Ten factual questions with a single checkable right answer about service capabilities, quotas and compliance controls. Ten judgment questions about architecture and compliance trade-offs on county and school district migrations, where the expected answer names a trade rather than a winner. Ten questions that should be refused or answered with a stated gap: unreleased pricing, anything requiring customer data, and topics our library genuinely does not cover.

For each question give me the question, the expected answer in two sentences, and the pass condition. Write the pass condition for the refusal group carefully: retrieving a restricted document and then declining to quote it is a failure, not a pass.

Fadi then edits the expected answers against real documents. That editing pass is where half the value is, because it finds the questions his own team disagrees about.

Run and grade
Attached are the thirty questions with their expected answers and pass conditions, and the thirty answers my account Project and documentation connector produced.

Grade each one on three axes separately: correctness against the expected answer, sourcing meaning every factual claim carries a document, and correct refusal for the refusal group. Score each axis pass or fail, never partly.

For every failure, say which Project instruction or which library document would have prevented it, and quote the sentence from the answer that failed. Then group all failures by root cause and rank the causes by how many failures each one explains.
Turn the score into changes
From the grouped failures, give me two lists. First, the changes to my Project instructions, written as the exact replacement lines. Second, the gaps in the documentation library, written as document titles somebody on my team has to write, ranked by how many failures each would close.

Do not suggest better prompts. I need the fixes that hold for everyone using this, not for the person who wrote that question.

Re-run the whole set after the fixes. A score that moved is a fix; a score that did not is a theory.

Quick check

Try it

Report a bug or share feedback