Skip to content

How we work

Five stages from “this is broken” to operated.

No six-month discovery theatre. We map the process, write the blueprint, prove it on your data behind a human review queue, then harden it into production — with an exit plan you own.

Stage by stage

Deliverables, durations,<br />and what we need from you.

Each stage ends with something concrete you keep — a process map, a blueprint, a working pilot, a runbook — whether or not you continue to the next one.

  1. 01

    Discovery

    3–5 days

    We sit with the people doing the work today, map the process as it actually runs — not as the SOP claims — and find where the hours and the errors are.

    • Process map with volumes and cycle times
    • Baseline metrics to beat
    • Risk and data-access review
  2. 02

    Blueprint

    1 week

    A written plan you can argue with: architecture, model choice, integration points, human-review design, evaluation criteria and a cost model per transaction.

    • Solution blueprint
    • Evaluation plan and target accuracy
    • Fixed-scope pilot proposal
  3. 03

    Pilot

    2–4 weeks

    A working system on your real data behind a human review queue. Small blast radius, real numbers, no slideware.

    • Working pilot on production-shaped data
    • Measured accuracy and cost per item
    • Go / no-go recommendation with evidence
  4. 04

    Production

    4–10 weeks

    Hardening: throughput, failure handling, permissions, monitoring, audit trail, runbooks and training for the team that will own it.

    • Production deployment in your cloud
    • Monitoring and alerting
    • Runbook and operator training
  5. 05

    Operate

    Ongoing

    We keep it working: evaluation runs on every model or prompt change, drift review, cost tuning, and a monthly session on what to automate next.

    • Continuous evaluation and regression tests
    • Monthly quality and cost report
    • Quarterly roadmap of new automations

What we need from you

Four things, and the project moves fast.

  • One process owner who can make decisions without a committee
  • Read access to the systems in scope (sandbox is fine to start)
  • A sample of real documents or tickets — a few hundred beats a few thousand
  • A named reviewer from the team that does the work today

The arithmetic

Eleven steps in.
Three steps out.

This is where the value sits: not in a clever model, but in the number of times a person touches an item before it is finished.

Manual steps removed, exceptions kept human. Nothing with a financial consequence posts without either a match or a signature.

before · manual

11 steps
  1. 01 Invoice arrives in a shared inbox
  2. 02 Someone opens the PDF and reads it
  3. 03 Keys 14 fields into the ERP
  4. 04 Checks the PO in a second system
  5. 05 Emails the warehouse about a mismatch
  6. 06 Waits for a reply
  7. 07 Notes the exception in a spreadsheet
  8. 08 Re-keys the correction
  9. 09 Files the PDF in a folder
  10. 10 Month-end: rebuilds the list by hand
  11. 11 Closes late again

after · automated

3 steps
  1. 01 Invoice arrives — parsed, matched and posted
  2. 02 Only exceptions reach a human, with the mismatch shown
  3. 03 Every action logged, reconciled and reportable

Same controls, same audit trail, same month-end — with the keying removed and the exceptions collected in one queue instead of eleven inboxes.

Team mapping a process on paper next to laptops during a working session
Discovery: the process as it actually runs, mapped with the people who run it
Source code on a dark screen during development
Build: durable workflows, typed actions, evaluation harness in the same repository
Close-up of a team working through documents and laptops around a table
Handover: your team owns the console, the runbook and the prompts

How the routing actually works

Humans where it matters,
by arithmetic.

Every step has a threshold, tuned on your own evaluation set. Above it the action executes; below it the item goes to a person with the evidence attached. That is the whole safety story — not a disclaimer, a number.

routing rule

route(x) = HUMAN   if  p̂(x) < θ
           AGENT   otherwise
p̂(x)
calibrated model confidence for step x, not the raw softmax score
θ
threshold tuned on your evaluation set — typically 0.82–0.94
HUMAN
the item lands in the review console with its evidence attached
AGENT
the action executes and is written to the audit log

Thresholds are per step, not per system: a wrong supplier name is cheap to fix, a wrong payment is not.

If the pilot fails

You keep everything: the process map, the blueprint, the evaluation set and the code. There is no proprietary lock-in to walk away from, and no exit fee.

If you want to stop

Monthly engagements run on 30 days’ notice. We hand over repositories, credentials you own, and a written runbook so your team can continue.

If it works

We extend stage by stage — more document types, more steps, more integrations — each one with its own measured target rather than one giant rollout.

Next step

Start with the map, not the model.

Discovery is a fixed fee and takes three to five days. You will know what to automate, what it is worth, and what we would build first.

One business day — and the reply comes from an engineer, not a sales sequence