Start with a recurring task that has accessible inputs, a clear definition of done, and mistakes your team can catch before they reach a customer. Compare candidates using the same evidence, then pilot one small handoff with a named owner and a measured baseline. The scorecard below is a suggested decision method, not a promise of savings.
What makes a task ready for AI automation?
Write the task as a handoff: when something arrives, someone uses specific information to produce a result for another person. “Improve operations” is too broad. “Turn incoming project requests into draft intake records for the coordinator to review” is a task you can inspect.
Choose work your staff can already explain. Ask the person doing it to show a routine case and a difficult one. If they disagree with their manager about the correct outcome, resolve that disagreement before buying software. Otherwise, a tool demonstration can hide an unfinished business decision.
AI belongs in the parts that involve interpreting language or preparing a draft. Moving an approved record between systems may only require ordinary automation. Keep those needs separate in your task description so a vendor can explain exactly where a model adds value.
How do you build a useful shortlist?
Ask each team member where they repeatedly copy information, summarize messages, search for instructions, or prepare similar drafts. Record actual examples instead of collecting general complaints about being busy. Include the downstream person who receives each result, because their cleanup work belongs in the assessment.
For every candidate, capture:
- The event that starts the work and the person accountable for it.
- The records, messages, or documents needed to complete it.
- The output and where it must be stored.
- The normal decision rules and common exceptions.
- The time spent doing, checking, and correcting the work.
- The consequence of a wrong, missing, or duplicate result.
You now have a shortlist of processes rather than a list of products. If you need help describing the difference between purchasing a tool and commissioning a connected system, read what an AI automation agency does.
How should you score the candidates?
Use a simple worksheet with the labels ready, needs work, and blocked. Avoid adding everything into a single score that lets a serious operational problem disappear behind an attractive time estimate. The labels are a suggested workshop format, not an industry benchmark.
Repeatability: Can the team describe the usual path and identify exceptions? Mark a task ready when the inputs and expected output are reasonably consistent. Mark it needs work when instructions depend on information people have never written down.
Input access: Can the proposed system retrieve the correct version of each required record through an approved connection? Mark it blocked when nobody owns the source or the team cannot establish which record is authoritative. A polished draft based on outdated information still creates work.
Reviewability: Can a reviewer check the result against the source without repeating the entire task? A draft intake summary may be easy to compare with an email. A recommendation requiring extensive investigation may offer little relief if staff must reconstruct the reasoning every time.
Consequence: Can you keep the initial output internal and reversible? Favor tasks where the pilot can produce a draft or a suggested change. Put customer commitments, record deletion, and other consequential actions behind explicit decisions about ownership and approval.
Operational value: Does the task happen often enough to matter, and does its delay hold up other work? Use your own records. Do not assume a tedious task is automatically the largest bottleneck, or that the most frequent task is the most valuable one.
What would a first project look like in practice?
Consider a hypothetical service business comparing intake summaries, automatic customer replies, and weekly report preparation. Intake summaries might be ready if messages are accessible and the coordinator already checks every new request. Automatic replies might need more work if the team has no agreed rules for unusual customer requests.
Report preparation could be blocked if staff disagree about which spreadsheet contains the current figures. That does not mean reporting is a poor long-term candidate. It means data ownership is a prerequisite, while intake summaries may offer a smaller first project with fewer dependencies.
For intake, define an output with named fields: request type, requested timing, missing information, and a link to the original message. Require unknown information to remain unknown. The coordinator should be able to accept, edit, or reject the draft without changing how incoming requests are tracked.
Voice work deserves the same assessment. If AI calling is a candidate, RizzDial supports AI voice agent calls and CRM connections. First define the calling task, required record updates, and human handoff; a product category alone does not establish that calling should be your first automation project.
How do you measure the work before changing it?
Observe a representative batch of real work before the pilot. Include routine cases, incomplete submissions, and exceptions. Record active handling time separately from waiting time. A process may feel slow because it sits in a queue, even when the hands-on work takes little effort.
Use the same definition of completion for the manual process and the pilot. If the manual task ends with an accepted intake record, the automated task must also end there. A generated draft is an intermediate output, not a completed record, while it still needs review.
Track these measures alongside the task volume:
- Human handling time, including review and correction.
- Elapsed time from arrival to accepted completion.
- Cases returned because information is missing or wrong.
- Exceptions that require the original manual process.
- Maintenance time spent updating instructions and connections.
Keep the observations in a shared document. Small samples can help identify failure modes, but they should not become sweeping claims about future performance. Repeat the comparison when the type of work or the process changes materially.
What should the pilot be allowed to do?
Write a short boundary statement before connecting systems. For example: “The pilot may read the intake inbox and create draft records. It may not send customer messages, change accepted records, or choose an appointment.” This gives staff a concrete way to spot behavior outside the agreed task.
Separate preparation from approval. Microsoft's Power Automate approvals documentation describes workflows that request human sign-off and wait for a response. That illustrates a useful design option: an automated process can pause at a decision point while a person remains accountable for the outcome.
Choose someone who reviews exceptions and someone who can pause the system. They may be the same person in a small business. Document the manual fallback and make sure the team knows which queue to use if the connection fails or a draft cannot be trusted.
When should you expand, revise, or stop?
Agree on acceptance criteria before seeing the pilot results. Require the outputs to meet the team's existing quality standard, with enough reduction in total effort or delay to justify ongoing ownership. State which errors would immediately pause the pilot regardless of its average performance.
Expand only when the evidence supports the next step. That might mean covering another request type while keeping human review. It does not have to mean removing approvals. Change one boundary at a time so the team can understand what caused a new problem.
Revise when failures have a specific fix, such as a missing source document or an unclear field definition. Stop when the task needs as much correction as manual completion, when input access remains unresolved, or when nobody can own maintenance. Stopping a weak pilot is a useful decision.
What should you bring to an implementation conversation?
Bring the task worksheet, sample inputs with unnecessary personal information removed, the baseline, the acceptance criteria, and the proposed permissions. Ask the implementer to demonstrate the entire handoff, including a failed run and a reviewer rejecting the output.
Ask who maintains the connection, how changes are tested, and how staff resume manual work. The guide to hiring an AI consultant provides more context for evaluating implementation help. You can also contact James Hill with the process you want to assess and the result your team needs.
What do owners ask before picking a first task?
Should we automate the task that takes the most time?
Include it in the shortlist, but check readiness and consequences first. A large task with unclear instructions may need process work before automation. A smaller, well-defined handoff can be easier to evaluate.
Do we need to replace our existing tools?
Start by mapping the systems you already use and the connections the task requires. Ask an implementer to demonstrate those connections before considering a replacement. Tool changes should follow a documented requirement.
Can the first project remain under human review?
Yes. Draft preparation with review can be the intended operating model. Judge it on accepted work and total staff effort, including review, instead of treating unattended operation as the only successful outcome.