The AI Decision Test — When to Automate
A five-minute scorecard for operators — decide whether AI should draft, assist, or stay out of a task entirely before you waste tokens or create liability.
Why you need a decision test first
Teams burn budget two ways: using AI on work that needs a human judgment call, and not using AI on repetitive draft work a checklist could govern. This test takes five minutes and stops both mistakes.
Run it before you prompt, before you buy a tool, and before you wire AI into a client workflow.
The four zones
| Zone | Meaning | Example |
|---|---|---|
| Green | AI drafts; human spot-checks | Email variants, meeting summaries |
| Yellow | AI assists; human owns outcome | Code review, financial summaries |
| Red | Human only; AI may prep research | Legal advice, medical, hiring final call |
| Gray | Pilot with logging and rollback | Customer-facing chatbots, auto-replies |
Scorecard — rate each 1–5
1 = strongly disagree · 5 = strongly agree
- Reversible — A wrong answer is cheap to fix (no money, safety, or reputation at risk).
- Source-grounded — Truth can be checked against docs, data, or specs you can paste in.
- Pattern-rich — You've done this task ten+ times and know what "good" looks like.
- Time-heavy, judgment-light — Mostly formatting, summarizing, or first drafts.
- Low personalization — One-size-fits-most is acceptable for v1.
- Audit trail OK — You can log prompts and outputs without violating privacy.
- Stakeholder expects AI — Audience knows content may be AI-assisted (or does not care).
Scoring:
- 28–35: Green — automate or batch with light review
- 20–27: Yellow — AI first draft, mandatory human gate
- Below 20: Red — do not ship AI output without expert review
- Any single question scored 1 on reversibility or source-grounded: downgrade one zone
Decision tree (copy-paste)
Answer yes/no:
1. Could a wrong answer cost money, safety, or legal exposure? → YES = stop or Red zone
2. Do you have source material to paste in? → NO = research first, don't guess
3. Can a human review in under 10 minutes? → NO = split task or stay Yellow
4. Is this a recurring task (weekly+)? → YES = worth templating + automation
5. Will the output go external without review? → YES = Yellow minimum, often Red
Task templates by zone
Green — ship with checklist
Task: [e.g. weekly status email from bullet notes]
Output: [format]
Sources: [paste notes]
Review: scan for names, numbers, dates only
Yellow — two-pass workflow
Pass 1 — Plan: outline steps and risks; do not execute
Pass 2 — Execute after I approve
Pass 3 — I will verify: [specific checks]
Red — research only
Summarize public sources on [topic]. Label every claim [SOURCED] or [INFERENCE].
Do not recommend a course of action.
I will decide.
Red-flag list (automatic Red zone)
- Medical, legal, or compliance binding language
- Credit, insurance, or individualized financial advice
- Hiring/fire decisions on real candidates
- Publishing unchecked facts about identifiable people
- Access to production secrets you cannot paste safely
- Tasks where the model cannot see the system of record
Pilot checklist (Gray → Green)
- Define success metric (time saved, error rate, CSAT)
- Log 20 runs before full rollout
- Assign human reviewer with veto power
- Document rollback (how to revert bad sends)
- Re-run decision test after 30 days
One-page worksheet
Task name: _______________
Requester: _______________
Score (7 questions): ___ / 35
Zone: Green / Yellow / Red / Gray
Reviewer: _______________
Sources attached: Y / N
Reversible if wrong: Y / N
Ship date: _______________
Notes: _______________
Score more workflows faster with guides in the MyGearHut free library. The Gear Drop sends one operator-ready AI note per week — optional, unsubscribe anytime.