Define One Agent Job and Outcome
Choose one repeatable business job, define the useful output, and turn quality into binary evidence.
What you will complete
A job brief, v0 boundary, and three to six independent pass/fail outcomes.
Make the operating decision
Describe the trigger, permitted inputs, primary output, person who uses it, and decision it improves. Avoid broad goals such as helping with sales or managing operations. A testable job ends with an inspectable artifact: a brief, decision record, validated file, or bounded action. Record what is outside v0 so later enthusiasm cannot silently enlarge the build.
Turn useful, accurate, and safe into separate binary checks. One criterion can require every estimate exception to include its source row. Another can require missing owners to remain missing. A third can require customer contact to remain a draft-only prohibited action. Independent checks show which part failed and prevent an attractive format from hiding a factual or permission defect.
Choose the evidence before the run. Use one representative real or sanitized case for the first build and hold back at least one case. If no real case is available, use labelled fiction for structure and keep business fit unverified. A passing fixture is not a substitute for a real-input result.
Worked fictional example: Harbourlight Home Services
Harbourlight wants a Friday estimate-exception brief. The trigger is a sanitized weekly CSV export. The output lists open estimates missing an owner or dated next action, cites row IDs, and records unknown values without guessing. The operations manager uses the brief to assign manual follow-up. Customer messages and CRM changes are excluded.
Harbourlight Home Services and every estimate record in this course are fictional. They demonstrate the method and do not represent a Flowgrammer client, a deployed system, or measured savings.
Complete workbook section 1
Use the evidence available for your own bounded job. Write unknown when the evidence is missing, and record the person or action that can resolve it.
- Name one recurring job and its trigger.
- Identify the allowed inputs, primary output, user, and decision.
- Write three to six independent binary outcomes.
- Place every attractive extra in v1 or v2 with a reason.
Critical gate before continuing
- The job can be completed end to end without a second workflow.
- Every outcome has observable evidence.
- External effects are excluded or explicitly gated.
- A held-back input is reserved.
If a gate fails, repair the current section, narrow the scope, leave the route manual, or record a blocked or stop decision. Continuing is not the only successful learner action.
Common failure modes
- Expanding beyond a job brief, v0 boundary, and three to six independent pass/fail outcomes. before the current artifact can be graded.
- Turning a missing value, unavailable source, or blocked integration into a confident conclusion.
- Treating a prompt instruction as proof that the effective tool or permission boundary works.
- Marking a manual, simulated, or untested route as live.
Check your application
1. Which outcome is easiest to grade?
Explained answer: Every exception includes its source row ID. A source-row requirement is observable and independent of taste.
2. What belongs in v0?
Explained answer: Only features required for one useful end-to-end result. A small v0 creates evidence before the learner adds operating complexity.
3. Why hold back an eval case?
Explained answer: To test behaviour that was not tuned to the first example. A held-back case provides a more honest check of general behaviour.
Return to Build a Business AI Agent with AGENTS.md, Skills and Evals