Production AI Field guide 01
What happens after an AI proof of concept works?
The next step is a controlled workflow people can depend on. Here are the decisions to make before you hand it over.
Your AI demo worked.
Now send it the same invoice twice.
Does it recognise the duplicate, create another record, or ask someone to check? That question brings a successful proof of concept into contact with daily operations.
A prototype gives you evidence that an approach can work under the conditions you tested. Before a team relies on it, you need to define how it handles exceptions, what it can change, and who takes over when it gets stuck. You also need a way to tell whether the work is worth continuing.
Start with one complete workflow
Choose a narrow task with a clear beginning and end. For an invoice workflow, that might mean receiving a document and creating an approved draft accounting entry. Payment stays outside the pilot.
Illustrative example
This invoice workflow is a design example, not a client case study or a claim about delivered results.
- 01 · RECEIVEKeep the original invoice.
- 02 · CHECKExtract fields and validate records.
- 03 · REVIEWRoute exceptions for approval.
- 04 · RECORDCreate one approved draft entry.
Walk through the whole path with the person who does the work today. Include the spreadsheet they consult, the supplier they call, and the corrections they make before an entry is ready. Each of those details can change the implementation.
You may only need AI for reading an inconsistent document. Validation rules, permissions, and approvals can remain explicit. Anthropic's guide to building effective agents also recommends starting with the simplest approach that meets the task.
1. Decide where exceptions go
An unreadable attachment needs a different response from a missing purchase order. A changed supplier bank account needs a different review from a spelling mismatch.
List the cases that should stop processing. For each, record the reason, the evidence a reviewer needs, and the person responsible. Give the reviewer a way to correct, reject, or return the item without losing the original.
Test those cases deliberately. Keep a small evaluation set containing ordinary invoices and difficult examples, with the expected outcome written down. Re-run it when the model, prompt, or integration changes. A confident-looking answer is not evidence that the fields are correct.
Anthropic's guide to agent evaluations explains how repeatable tasks and explicit success criteria help teams catch regressions as a system changes.
Before moving on
Send an unreadable document through the workflow. Confirm that it reaches the right queue with an explanation and that a person can resolve it.
2. Define what can happen without approval
Reading an invoice, preparing an entry, changing supplier details, and releasing money are separate actions. Grant access for the action the pilot needs, and enforce that boundary in the connected systems.
For this example, the workflow can prepare a draft. A reviewer checks the source document and approves that specific draft before it is recorded. If the draft changes after approval, it needs review again.
Treat instructions inside invoices or emails as untrusted document content. A sentence in an attachment must not grant the workflow new permissions or change the approval policy. Log which person approved which action, while limiting access to sensitive document data.
These boundaries follow OWASP's guidance on excessive agency: minimise permissions, require approval for high-impact actions, and enforce authorisation in downstream systems.
Before moving on
Attempt an action outside the pilot's scope. Confirm that the system blocks it even if the model requests it.
3. Make retries safe
Suppose the accounting system accepts an entry but the connection drops before it confirms success. Your workflow sees a timeout. Sending the same request again could create a duplicate.
Give each intended write a stable operation identifier and use the destination's duplicate-prevention mechanism where available. Keep the same identifier when retrying that operation. AWS explains this pattern in Making retries safe with idempotent APIs.
If the destination cannot guarantee that behaviour, reconcile its records before another write. Route an uncertain result to a person. Separately, check for the same invoice arriving through different emails; duplicate documents and repeated API calls are different problems.
Before moving on
Test both a duplicate submission and a timeout after a successful write. Confirm that the workflow creates one intended entry and leaves an understandable record of what happened.
4. Name the person who owns Monday morning
Someone needs to notice a growing review queue, an expired connection, or a sudden increase in corrections. Name a business owner for the outcome and a technical owner for the system. In a small team, one person may hold both roles, but the responsibilities still need to be clear.
Agree when alerts need a response and who covers absences. The handover should include access, a short operating guide, a change log, and a tested way to pause new work.
Google's SRE chapter on monitoring explains why alerts should identify clear problems that people can act on, with enough context to investigate them.
Pausing the workflow does not undo entries it already created. Define how to identify affected records, reconcile them, and make authorised corrections. Keep the existing manual process available during the pilot.
Before moving on
Ask the operator to pause the workflow and resolve a failed item using the guide. Watch where they need help, then improve the handover.
5. Measure the work that remains
Record the current process before rollout: handling time, correction rate, queue age, and completed volume. Compare similar work over an agreed pilot period. A faster extraction step tells you little if reviewing its output takes longer.
Count human review, exception handling, software, model usage, maintenance, and the initial implementation effort. Track failure and correction rates alongside speed so that a throughput increase does not hide more rework.
A useful time measure
Net staff time released = previous handling time − review time − exception handling − ongoing administration.
Freed capacity has value, but it becomes a cash saving only when spending changes. Be explicit about whether the goal is less overtime, more completed work, fewer corrections, or a shorter turnaround. Avoid counting the same benefit twice.
Before moving on
Agree what result would justify expanding the pilot, what would require another iteration, and what would make you stop. Choose those thresholds before seeing the results.
A practical launch checklist
Bring the business owner and technical owner together. For each item, ask to see the evidence:
- The workflow has a defined start, end, and list of excluded actions.
- Ordinary inputs and known exceptions have been tested against expected outcomes.
- Permissions and approvals are enforced outside the model.
- Duplicate documents and uncertain write results have a safe resolution path.
- An operator can inspect failures, pause work, and use the manual fallback.
- A baseline, pilot review date, and expansion criteria are agreed.
An unanswered item gives you the next piece of work. Start with a limited rollout, inspect what happens, and expand when the evidence supports it.
Find the next step for your workflow
If you're deciding where AI belongs in your operations, start with the free Operations Audit. It gives you an instant score and a recommended first fix, without requiring an email to see your results.
Already have a prototype? Tell us what it does and where it gets stuck. We can discuss the integration, review, and ownership work needed to take it further.
Start the free Operations AuditSources & further reading
The references below explain the engineering principles behind this guide. The invoice example, time measure, and launch checklist are AureliaEdge's practical synthesis.
- Anthropic — Building effective agents
Choosing simple workflows, adding autonomy only when needed, and testing agent behaviour. Used in “Start with one complete workflow”.
- Anthropic — Demystifying evals for AI agents
Repeatable evaluation tasks, success criteria, and regression testing. Used in “Decide where exceptions go”.
- OWASP — LLM06:2025 Excessive Agency
Least privilege, human approval, and authorisation outside the model. Used in “Define what can happen without approval”.
- AWS Builders' Library — Making retries safe with idempotent APIs
Stable request identifiers and avoiding repeated side effects after a timeout. Used in “Make retries safe”.
- Google SRE — Monitoring Distributed Systems
Actionable alerts, service signals, and monitoring that helps people investigate failures. Used in “Name the person who owns Monday morning”.
