Work in progress — we're building something great. Some pages and features are coming soon.
Implementation11 min read·

How to implement an AI employee: a step-by-step playbook

Implementation quality, not model quality, determines whether an AI employee delivers. Two businesses using identical technology routinely get completely different results, and the difference is almost always process discipline.

This playbook covers the six phases we run, what each one produces, and the failure mode each phase exists to prevent.

Phase 1 — Workflow audit

Map the process as it is actually performed, not as the procedure document describes it. Sit with the people doing the work, count volumes, time the steps and — critically — collect the exceptions they handle so instinctively that they forget to mention them.

Output: a documented process map, volume and handling-time baseline, a list of systems touched, and a shortlist of automatable candidates ranked by impact and risk.

Failure mode prevented: automating an imagined process. Nothing derails a deployment faster than discovering in week six that thirty per cent of cases follow an undocumented path.

Phase 2 — Scoping and guardrails

Decide explicitly what the AI employee owns, what it may decide alone, what needs approval and what it must never touch. Write the escalation matrix now, with the team, not later under pressure.

Keep the initial scope narrower than feels comfortable. A worker that handles three case types brilliantly builds more organisational trust than one that handles fifteen adequately.

Output: a scope document, escalation matrix, approval thresholds and the success metrics you will be judged on.

  • Which case types are in scope, explicitly listed
  • Autonomy level per case type: read-only, draft, or full
  • Financial and irreversibility thresholds requiring human approval
  • Named process owner and review cadence

Phase 3 — Knowledge and integration

Gather documentation, policies, pricing rules, tone-of-voice guidance and historical cases. Resolve contradictions before ingestion; the worker will otherwise pick one version and defend it confidently.

Integrate with the systems in scope using least-privilege access, and build the write actions behind explicit validation. Test each tool independently before wiring them into the worker.

Output: a grounded knowledge base with clear ownership, working integrations, and a test suite of real historical cases with known correct outcomes.

Phase 4 — Shadow mode

Run the AI employee on live work with every output reviewed by a human before it goes anywhere. Two to three weeks is usually enough, and it is the highest-value phase in the entire project.

Track accuracy per case type, not overall. An eighty-five per cent average can hide a case type running at forty per cent, and that is exactly the one that will damage trust at go-live.

Log every correction with a reason. These corrections are the training material that turns a competent generic assistant into something that sounds like your best agent.

Phase 5 — Controlled go-live

Release autonomy per case type as each clears your accuracy threshold, rather than flipping a single switch. Out-of-hours-first is a useful soft launch: customers get a genuine improvement over silence, and the stakes are lower.

Monitor daily for the first fortnight — resolution rate, escalation rate, satisfaction and any pattern in corrections. Keep the ability to revert any case type to draft mode instantly, and tell the team it exists.

Communicate to customers plainly that AI is involved and that a human is always available. Transparency costs nothing and prevents the one complaint that generates disproportionate noise.

Phase 6 — Scale and maintain

After a month of stable operation, review the logs and widen scope where the evidence supports it. Most initial guardrails turn out to be more conservative than necessary.

Maintenance is not optional. Schedule a monthly review covering accuracy trends, new case types appearing in escalations, and knowledge updates triggered by policy or pricing changes. Assign it to a named owner with time allocated.

The second and third AI employees are dramatically cheaper than the first, because the integrations, governance and organisational trust already exist. That compounding is the actual strategic argument for starting now rather than waiting.

The five most common mistakes

First, skipping the audit because the process seems obvious. Second, defining success as deflection rather than resolution. Third, launching without system access, producing an expensive FAQ. Fourth, excluding the team from design and then wondering why adoption stalls. Fifth, treating go-live as the end of the project rather than the start of the operating phase.

Each of these is avoidable, and each is more common than any technical failure we encounter.

Frequently asked questions

How long does implementation take?

Two to six weeks for a single well-scoped process, with shadow mode accounting for a significant share of that time.

Who needs to be involved from our side?

A process owner with real knowledge of the work, someone who can authorise system access, and a few practitioners for the shadow-mode review.

Can we skip shadow mode to launch faster?

You can, and it is the most reliable way to damage trust. Shadow mode is where accuracy and confidence are built simultaneously.

What happens after go-live?

Monthly reviews covering accuracy trends, knowledge updates and scope expansion. An AI employee is maintained, not installed.

Free 30-minute session

Ready to hire your first AI employee?

In a free 30-minute session we map your workflows and show exactly where an AI employee saves your team hours every week.