Run, Grade, Repair, and Record
Execute the same cases, grade independent outcomes, diagnose failures, and preserve run evidence.
What you will complete
A passing run log or an honest blocked result after focused repair cycles.
Make the operating decision
Run the same case and grade every criterion pass or fail. Diagnose whether the problem comes from instructions, deterministic code, source data, permissions, or an unresolved product choice. Preserve the failed output and version before editing the artifact.
Make one focused correction, rerun affected cases, and append the new result. Do not weaken the expected result to manufacture a pass. After three unsuccessful repair cycles, stop and request the missing decision or reduce scope. A blocked build with clear evidence is more reliable than a green label built on untested assumptions.
Inspect the artifact itself. Confirm source row IDs, missing values, output location, and lack of external effects. If an independent fresh-context test is permitted, give it the normal inputs rather than the desired answer or suspected bug.
Worked fictional example: Harbourlight Home Services
Harbourlight's first run omits unknown-owner rows because the prompt treats blank as no exception. The team records the failure, corrects the explicit unknown rule, and reruns the same case plus the held-back case. The run log retains both versions and verdicts.
Harbourlight Home Services and every estimate record in this course are fictional. They demonstrate the method and do not represent a Flowgrammer client, a deployed system, or measured savings.
Complete workbook section 7
Use the evidence available for your own bounded job. Write unknown when the evidence is missing, and record the person or action that can resolve it.
- Run the representative case and grade every outcome.
- Classify each failure by layer.
- Make one narrow repair and rerun affected cases.
- Stop after three failed cycles and document the unresolved decision.
Critical gate before continuing
- Failed evidence remains available.
- Expected results do not move to match output.
- The held-back case passes independently.
- No untested criterion is marked pass.
If a gate fails, repair the current section, narrow the scope, leave the route manual, or record a blocked or stop decision. Continuing is not the only successful learner action.
Common failure modes
- Expanding beyond a passing run log or an honest blocked result after focused repair cycles. before the current artifact can be graded.
- Turning a missing value, unavailable source, or blocked integration into a confident conclusion.
- Treating a prompt instruction as proof that the effective tool or permission boundary works.
- Marking a manual, simulated, or untested route as live.
Check your application
1. What should change in one repair cycle?
Explained answer: One focused artifact or configuration issue. A focused change makes cause and effect visible.
2. When should the learner stop repairing automatically?
Explained answer: After three failed cycles or when a product decision is missing. A stopping rule prevents loops from hiding an unresolved requirement.
3. What is valid when a required permission is unavailable?
Explained answer: A blocked result with the missing access and next decision. Blocked is an honest operating state and preserves the authorization boundary.
Return to Build a Business AI Agent with AGENTS.md, Skills and Evals