← All notes
September 30, 2026Engineering Leadership9 min read

Requirements for an AI Automation: Template

Requirements for an AI automation are a written spec for a system that will sometimes be wrong: what it does, how often it may be wrong, what happens to the items it is unsure about, and who catches the rest. Without those four, “automate X” is a wish, not a ticket. This post is for engineering managers, product owners and anyone who has to sign off an AI feature.

I’m Daniel Awde, an Engineering Manager. In why your team’s estimates are wrong I gave this one paragraph: an automation that is mostly correct is not mostly done. Below is the full template I would hand a team, and a worked example.

Updated for 2026.

What the requirement has to contain

A model call is the small part of an AI automation. The rest is decisions that someone has to make before the work starts, or the team makes them silently while building.

Six things must be written down.

  1. The unit of work. One input, one output. “A support ticket in, a queue name out.”
  2. What correct means, and who decides. A person or a rule that is the ground truth, not the model’s own opinion.
  3. A target, measured on a sample. A number you will accept, and the labelled examples it is measured on.
  4. A confidence threshold. Below it, a human decides.
  5. The human path. Who reviews, where, how fast, and what they can change.
  6. A record. Every decision logged well enough to audit and to improve.

If a ticket omits any of these, the team is guessing at the size of the unknown, which is where estimates go wrong.

The template

Copy this into the ticket. Every field needs an answer, or an explicit “not in scope” with a reason.

Automation:         [one sentence: input -> output]
Ground truth:       [who or what decides the right answer]
Evaluation set:     [where the labelled examples live, who labelled them,
                     how they were sampled, languages covered]
Target:             [metric] >= [value] on the evaluation set, per language
Cost of errors:     [which mistake is worse, and by how much, in words]
Confidence rule:    [score source] < [threshold] -> human review
Threshold set by:   [how it was chosen from the evaluation set, by whom]
Human path:         [queue / owner / response time / what they see]
Override:           [how a reviewer corrects the output; where it is stored]
Failure handling:   [timeout / malformed output / unknown category /
                     unsupported language -> what happens to the item]
Logging:            [fields kept, what is NOT kept, retention]
Post-launch check:  [sample audit: who, how often, what triggers a rollback]
Rollout stages:     [shadow -> assist -> automate, with exit criteria]
Out of scope:       [what this automation will not do]

Two fields cause most arguments, so they get their own sections.

Confidence thresholds and human review

A threshold is a decision about cost, not a property of the model. Raising it sends more items to people and fewer mistakes through. Lowering it does the opposite. The right value depends on the cost of each kind of error, which is why “Cost of errors” is a field and not an afterthought. A misrouted ticket that someone reroutes in a minute is a different problem from a wrong answer sent to a customer.

Know where the score comes from. A model’s stated confidence, a token probability and a classifier score are different quantities, and none of them equals “the chance this is right” until you check. Choose the threshold from the labelled evaluation set: look at what fraction of items above each candidate value were actually correct, and pick the value that meets your target. Then write down the evaluation set and the date, so the next person can redo it.

Set the threshold per language and per category if they behave differently. An average hides a weak spot. If the automation handles Arabic and English, the target is written per language, and the evaluation set has enough of each to tell you something. A single blended figure can pass while one language fails.

The human path is part of the feature. Specify the parts a reviewer needs:

  • what they see: the input, the model’s output and the score,
  • what they can do: accept, correct, or reject and route elsewhere,
  • how long an item may wait before it breaches whatever service level the business has,
  • who owns the queue when the reviewer is away.

An escalation queue nobody is assigned to is a silent failure. Assign it before launch.

Corrections are data. Store the reviewer’s override next to the original output. That gives you the next evaluation set for free, and it shows you whether the threshold still holds.

Worked example: ticket triage

This example is illustrative. The values in square brackets are placeholders that your own evaluation set has to fill. None of them are figures from a real project.

Automation:         Inbound support ticket (subject + body) -> one queue
                    from a fixed list of [N] queues
Ground truth:       The queue a senior agent would choose, from the last
                    [period] of tickets, labelled by [named role]
Evaluation set:     [M] labelled tickets, sampled across all queues,
                    with Arabic, English and mixed-language tickets
                    each represented
Target:             Correct queue on >= [X]% of auto-routed tickets, per language
Cost of errors:     Wrong queue: reroute delay. Wrong queue for a
                    [priority class]: unacceptable, always human
Confidence rule:    score < [T], or priority class in [list] -> human
Threshold set by:   Highest [T] at which auto-routed accuracy meets [X]%
                    on the evaluation set, chosen by [role], dated
Human path:         Triage queue owned by [role]; every item shows the
                    ticket, the suggested queue and the score
Override:           Reviewer picks the correct queue; stored with the
                    model output
Failure handling:   Timeout, empty output or unknown queue -> human queue,
                    never a default queue
Logging:            Ticket ID, suggested queue, score, final queue,
                    reviewer. Not the ticket body
Post-launch check:  [role] audits a random sample of auto-routed
                    tickets every [interval]; accuracy under [X]% -> rollback
Rollout stages:     Shadow (suggestions logged, not used) -> assist
                    (suggestion shown to agents) -> automate above [T]
Out of scope:       Replying to the customer, closing tickets

Notice what the template forced. It made someone name the queue list, the priority classes that must never be automated, and the failure behaviour, all before code exists. That is the actual shape of the work, and it is what lets a team estimate it.

Defining “done” for an AI feature

“The model returns a queue” is not done. A workable definition has three parts.

Measured. The target is met on the evaluation set, per language, with the run recorded: date, model version, prompt or configuration, threshold. Rerunning it is a command, not a project.

Contained. Every failure path in the template has been triggered on purpose and behaves as written. Kill the model endpoint. Send malformed output. Send a language you don’t support. Each ends in the human queue, not a silent default.

Operated. Someone owns the review queue, someone owns the sample audit, and the rollback is a switch that has been tested.

Then roll out in stages. Shadow mode runs the automation and logs its suggestions while people keep working as before, so you measure on real traffic with no exposure. Assist shows the suggestion to the human. Automate lets the system act above the threshold. Each stage has an exit criterion in the ticket. Skipping to the last stage is how a ‘mostly correct’ system becomes an incident.

I would rather ship the smallest automation that meets its template than a broad one that half meets it. Everything the automation is not allowed to do belongs in “Out of scope”, written by the same person who wrote the target.

More posts on delivery are listed on the blog, and the AI work I lead is summarised under work.

Frequently asked questions

What should requirements for an AI automation include?

At minimum: the unit of work (one input, one output), who defines a correct answer, a target measured on a labelled sample, a confidence threshold, the human review path, failure handling, logging, and a rollout plan. Add an explicit out-of-scope line. Each item removes a decision the team would otherwise make silently during the build.

How do I choose a confidence threshold?

Choose it from a labelled evaluation set, not from the model’s defaults. For each candidate value, measure how many items above it were actually correct, then pick the value that meets your accuracy target at an acceptable review volume. Record the set, the value and the date. Recheck after any model or prompt change.

What does human-in-the-loop mean in a requirement?

It means the ticket names a person or role who reviews items the system is unsure about or must not decide alone. The requirement specifies what they see, what they can change, how long an item may wait, and who covers when they are away. A review step nobody owns doesn’t count.

Is a 90 percent accurate automation finished?

No. It needs a review path for the remaining errors, a way to measure them after launch, and a decision on which errors are unacceptable. Accuracy on a test set says little about the cost of the misses. The ticket should state which mistakes matter most and how they are caught.

What should I log for an AI automation?

Log enough to audit and improve: an item identifier, the output, the score, the final decision, who decided, and the timestamp. Keep sensitive content out of logs unless there is a stated reason and a retention period. Reviewer corrections should be stored beside the original output, because they become your next evaluation data.

Related reading: Why your team’s estimates are wrong · All posts · Engineering leadership category

By Daniel Awde, Engineering Manager. Updated for 2026. The worked example is illustrative; bracketed values are placeholders, not results from any project.

Related reading

Join the conversation

Your email address will not be published.