A New Zealand construction head contractor with residential, commercial and hotel-refurbishment projects receives 500 to 1,000 supplier invoices a month. Before we started, receiving, sorting by site, matching purchase orders, submitting approvals and chasing overdue items all sat with one quantity surveyor (QS). We connected in June 2026, ran real invoices through in July and calibrated at full volume in August. Measured on the client's accounting month (26 July to 25 August): 798 invoice-related documents processed, 730 recognised and routed automatically, manual review down to 68, an automatic routing rate of 91.48%. This article is not about the result. It is about what happened in the three months in between.
How were invoices handled at the start of each month?
The client's stack is typical: supplier invoices arrive in an Outlook mailbox, files live in OneDrive, approvals run through ApprovalMax, the books are in Xero, and Xero locks each month. The 26th to the 3rd is the cross-month window when invoices pile up; the first two or three days of a month could hold 600 to 700 of them.
The QS looked at every one: which site, which project type (residential, commercial, hotel, office, or not this company's bill at all), whether there was a matching purchase order (PO), which project manager should approve it, and then submitted it into ApprovalMax by hand. Hotel projects added consolidated bills that had to be checked across several folders. The mapping from approver to site was not written down anywhere. It lived in a few people's experience.
In mid-July the client's business contact wrote one line in the group chat that turned out to define the project: the last two weeks of bills were being recognised by the AI but had not yet been uploaded to the system, and she was worried that if it dragged on, something would be missed. She was not worried about recognition. She was worried about the flow stopping. We did not fully understand that sentence until a month later.
What did we connect?
The principle was connect, don't replace: not one system was swapped. The agent lives between the tools the client already runs.
- Outlook listener:A dedicated invoice mailbox, with a separate app registration in Microsoft 365 that has read-only mail permission and can be revoked at any time.
- Runs locally:The agent runs on a Mac mini in the client's office, orchestrated by OpenClaw. Original invoices never leave the client's environment.
- Recognition and classification:OCR plus a large model extract supplier, invoice number, date, amount, GST and currency; files are filed into the matching OneDrive folder by site and project type; POs are matched.
- Submission:Qualifying invoices are submitted to ApprovalMax automatically. Anything the model cannot recognise or that is missing information goes to a Manual Review folder to wait for a person.
- Status write-back:ApprovalMax approval status syncs hourly into Lark Base, with two dashboards: an Admin version with amounts and an Approver version without.
- People stay in three places:The QS handles Manual Review; project managers approve in ApprovalMax; finance pays and locks the month in Xero. The AI does not approve, pay or post.

Month one: the problem was rules, not recognition
The first quantified count on 21 July: 407 candidate invoices, 63 into approval, a completion rate of 15.5%. Broken down by blocker: 173 with incomplete rules, 132 missing PO, approver or site information, and only 12 genuine recognition failures, under 3%.
In other words, the model read invoices very well and then did not know who to hand them to. Which site does this supplier's invoice belong to, who approves this project, does a power bill count as this company's expense: none of that is printed on an invoice. It was in the client's heads.
On 24 and 25 July we went back and forth with the client over 377 test invoices: 11 could go straight to ApprovalMax, 226 were correctly excluded (not invoices, or not this company's), 112 were waiting on information or rules, 17 were integration issues. A manual sample of 117 found 7 misjudged, a sample accuracy of about 94%. On 27 July we grouped 111 problem invoices into 7 categories and asked the client to answer in writing how each should be handled. The workflow went back to running automatically.

How many invoices are there, exactly?
In the first week of August the simplest number of all stopped reconciling. July's record count went from 369, to 460 after 88 attachments were recovered, to 815 records, and finally to 517 confirmed standard invoices. Two other models re-counting the same data independently returned 515 and 559. The client asked four questions we could not immediately answer: how many did we receive, did we miss any, how many were handled automatically, and how much manual work went away.
There were two root causes. The first was process: recognition, splitting and de-duplication were mixed into one pipeline, so an error at any step made the final number drift. The second was specific: the agent had earlier been asked to clean up June invoices that had already been processed manually, and it executed on invoice date rather than email-received date, removing some July invoices that should have stayed. There was a backup and 33 were recovered the same day, but without an independent reconciliation channel this might never have been noticed.
The fix was to turn the ledger into an independent audit chain: total emails, emails with attachments, attachment count, unique attachments, recognised invoices, cleaned, misjudged, duplicates, reconciled layer by layer so every layer can answer where the missing ones went. From then on the dashboard had two separate blocks, the previous closed cycle as a formal result and the current cycle in progress, with the data source for each stated.

The client did not want a recognition rate. They wanted the flow to keep moving
On 4 August the client reported that invoice volume had visibly dropped and there was nothing to work on. The engineer's first instinct was to keep tuning recognition rules so more invoices would pass automatically. An internal review found the interpretation was off: the client wanted the business to keep flowing first. Invoices the AI could not decide on, even with incomplete information, had to be routed in an orderly way to the right person, not held back until recognition was perfect, and certainly not piled into one Manual Review black box.
That was the most important turn in the project. Everything after it was designed around one question: how does the remaining part get handed to people?
- Sub-folders by cause:Manual Review is split by blocker: no PO, site undetermined, account code undetermined, supplier mapping missing.
- Mark Done when handled:The QS appends the missing information to the file name and marks it Done; the agent scans on a schedule, back-fills, submits, then marks it Submitted so nothing is uploaded twice.
- Urgent bills classified separately:Power, water, council rates, insurance, vehicle repairs, deposits and telecoms carry a high cost of lateness and go to the front of the manual queue.
- Not this company's invoice:Bills for other group entities or sent to the wrong company are excluded through allow and block lists.
- Duplicate check:Duplicates are caught on the combination of supplier, invoice number and amount.

August: full-volume calibration
Late August was close-out. The manual review queue fell from 300-plus to 181, then 110, then 81. On 28 August, by calendar month: 775 invoice-related documents, 181 needing a person (23%), 594 routed automatically (77%). The figure at formal handover on 30 August: 775 documents processed in August, 694 recognised and routed by AI, Manual Review down to 81, an automatic routing rate of 89.55%.
Recalculated on 1 September for the client's accounting period (26 July to 25 August): 798 documents, 730 routed automatically, 68 in manual review, 91.48%. Both numbers are here because they use different bases. The client looks at accounting months; internally we looked at calendar months. A case study has to state both.
| Metric | July cycle (26 Jun to 25 Jul) | August cycle (26 Jul to 25 Aug) |
|---|---|---|
| Invoice-related documents received | 674 | 798 |
| Recognised and routed by AI | 24 pushed to approval | 730 |
| Manual review | 529 awaiting information | 68 |
| Automatic routing rate | about 4% | 91.48% |
| Data date | dashboard, 10 Aug | project review, 1 Sep |


Who can see the amounts, and who cannot?
There are two dashboards. The Admin version for the owners and finance shows amounts and the distribution by project and approver; the Approver version shows each approver only their own queue, without amounts. Three custom roles were created in Lark Base, Approver, Admin and AI Master, and the agent writes data under the AI Master identity, unable to see what it should not.

The platform's boundary is where the AI stops
On 1 September the client asked for something that looked small: move a batch of July invoice dates into August so they land in the right accounting month. ApprovalMax's public API allows the date to change, but changing it resets the approval workflow, and the API has no approve endpoint, so the reset approvals can only be clicked by an approver with permission. Of 37 qualifying invoices, 19 updated successfully, 9 were skipped on a permission conflict (HTTP 409), and 9 were left alone because their status had already become rejected.
The choice was not to work around it. If the platform does not let a program approve on a person's behalf, a person approves. That became a delivery rule: for anything touching approval, payment or an external commitment, the AI prepares and checks, and a person presses the last button.
Four lessons from three months
Looking back, the expensive part was not the technology. It was these four judgments.
- A recognition rate alone points the wrong way:Recognition failures were under 3% in the first month while completion was 15.5%. The real work is in rules and mappings, which have to be asked out of the client's people, and whatever cannot be asked out has to be designed as a manual path.
- The ledger needs an independent audit chain:The automation ran, but when we could not say how many invoices had arrived, the client's trust went straight back to zero. Recognition, de-duplication and counting have to be separate layers that reconcile.
- AI executes instructions literally:People understood clean up the processed June invoices; the AI executed it by invoice date. Anything that deletes needs a dry run, a backup and a second data source to cross-check.
- Tell the client about the process:The team's July and August workload was heavy, but for two and a half months the client only saw that the result had not arrived yet. The 1 September review set three things: move from feature done to business running; lower the bar for front-line staff with fixed steps and multiple-choice prompts instead of open questions; find and close problems proactively rather than waiting for feedback.
Where is it now?
The technical phase was handed over on 30 August and September is an optimisation period: fixing duplicate approval reminders, supplier misclassification and post-lock date adjustments, plus a new agent on a second device for one of the client's managers. The quarterly effectiveness review is scheduled for after the client has used the system for a while.
The best acceptance test was not a number. In early September the client introduced us to the construction insurance broker they work with, another industry that is heavy on documents and heavy on compliance.
Evidence boundary: every figure in this article comes from the project chat record and the project review document of 2 September, each dated. 91.48% is measured on the client's accounting period and 89.55% on the calendar month; both denominators include non-invoice files. The client is a New Zealand construction head contractor and is anonymised; company and personal names in screenshots have been cropped or masked, and dashboards showing amounts were not used.
In this New Zealand AI implementation case study, what does the AI do and what do people do?
Do we need to replace Xero or ApprovalMax?
What does a 91% automatic routing rate mean, and what about the other 9%?
How long until invoice automation shows results?
AI Companies in New Zealand Compared (2026): Providers, Pricing and How to Choose
A side-by-side comparison of the main AI providers in New Zealand — BEE Sigma, Stride AI, Ez-AI, BestAI, HornTech, Datacom, Soul Machines and more — with public NZD price bands and the local compliance points that matter, so businesses of every size can match themselves to the right kind of partner.
Read articleBusiness Software in New Zealand Compared (2026): Accounting, CRM and ERP, and Why Most Companies Shouldn't Switch
Xero or MYOB? HubSpot or Pipedrive? Do you need an ERP at all? A practical comparison of the business software New Zealand SMEs actually use, in four categories: accounting, CRM, ERP and inventory, and collaboration. It ends with a counter-intuitive conclusion: for most companies the problem isn't the software, it's that nobody connects the pieces — exactly the job AI agents should take.
Read article