Blog · AI Implementation Case Study · NZ Cross-Border Retail

New Zealand AI Implementation Case Study: From WeChat Orders to SF Express Labels at a Cross-Border Health Retailer

Reviewed 14 Sep 20268 minBEE Sigma Delivery

This New Zealand AI implementation case study is about a cross-border health-products retailer: 5 stores and an online shop in New Zealand, and more than ten channels in China including Tmall, JD, Pinduoduo and Douyin, selling infant formula, supplements and honey, with SF Express as the main carrier. Before we started, daily China sales were read off screenshots and typed into Excel; SF orders were copied by one staff member from WeChat into a master sheet, then into an import template. The project started on 4 June 2026; on 9 September SF Express began pushing labels back automatically. Here are the hundred days, mostly in pictures.

01

How many times was one order copied?

On 24 July the client’s SF operator described her routine: ‘WeChat orders go into my master sheet, then I copy the content into the order template and import it into the SF system.’ One formula order passed through human hands five times between the WeChat group and the label.

Before and after: five manual steps (recognise the order, copy to master sheet, copy to SF template and import, archive IDs by hand, send numbers and labels back one by one) versus one sentence and one confirmation (@ the agent, split and map, four checks and ‘review passed’, create and archive and write back, SF pushes the label)
Before and after. On the right, the only human step is checking four items and saying ‘review passed’.
02

What happened in a hundred days

Kickoff on 4 June with three data scenarios; a change of route at the end of July when marketplace APIs ran into copyright rules, toward SF Express ordering; SF API connected on 26 August; label callback live on 9 September; coaching mode from 10 September.

Hundred-day timeline: 4 Jun kickoff, 9 Jul first-month review, 28 Jul change of route, 5 Aug WeChat connected, 26 Aug SF API connected, 2 Sep first production order, 9 Sep label callback live, 10 Sep coaching mode
Eight milestones, each traceable in the project chat.
03

Month one: the smallest thing first

Phase one was three things: group message search, a China daily sales report, a store daily report. One loop in four weeks, with the client in a test group from week two. The first-month review on 9 July recorded an agreement that kept proving true: third-party interfaces are atomic capabilities, and ‘there is no such thing as “the interface exists, so it can be used directly”’.

First-month plan: week one development and mock-data testing, week two testing with the client in a test group, week three iteration, week four go-live
The first-month plan of 9 June. The stated goal: let the client feel what the bots do quickly and fix the data problems in front of them. (Chinese-language original.)
First-month delivery overview: message search done, China daily sales agent deployed pending WeChat access, NZ store daily report in testing, Gmail basics done; below, plan versus actual and gaps
First-month review, 9 July: the state of the four scenarios and a plan-versus-actual table. (Chinese-language original.)
04

Two hurdles: the WeChat account and marketplace APIs

For an agent to read WeChat groups, an account must stay logged in on an always-on machine. WeChat’s risk controls on remote and multi-device logins are strict; it took nearly a month of joint effort before the dedicated account was stable on 5 August. That experience is now item one on our kickoff checklist for every WeChat project.

The client wanted the agent on Taobao and WeChat Channels APIs directly. A week of research showed the platforms require a software copyright held by the store entity, which a locally deployed agent cannot obtain. On 28 July the route changed: daily sales stayed in ‘people report, the bot records’ mode, and development moved to SF Express ordering.

Four preconditions for WeChat access: an account registered 3+ months and logged in inside China, an owner available for verification, never cloud and local logins at once, prove it in a test group first
Four preconditions, earned over a month.
Three routes for marketplace APIs: apply directly (requires store-entity copyright, not feasible), go through the ERP or a service provider (not feasible short term), chosen: screenshot recognition plus SF Express
The decision of 21 to 28 July: a week of research, then a different road.
05

The journey of one order

First test-environment order on 28 August, first production order on 2 September. Only one node in the chain requires a person: check recipient, product, payment and ID, then tell the agent ‘review passed’. Without that sentence the agent does not call SF’s production API.

Order flow: WeChat order text → agent parses and splits → verification card → human review passed → SF ISCO order → ID archiving to NAS → SF label callback → write-back to Lark Base
Eight nodes. The gold one is a person.
Lark group screenshot: the engineer replies ‘verification passed, no payment, no ID needed’; the agent writes back the work order and asks whether 6 bags are one fulfilment unit; after confirmation it creates the SF order
A real round trip on 31 August: a person says ‘verification passed’, the agent has one question left, the person answers, and only then does it create the order. Names and numbers pixelated. (Chinese-language original.)
Lark group screenshot: the agent reports ‘accept order and pull waybill’ succeeded, returns waybill numbers and notes SF has not returned a label address
26 August, the day the SF API connected: waybill number yes, label no. The label took another two weeks. (Chinese-language original.)
06

Why the label would not come

For a cross-border seller the label is the customer’s receipt; without it the order is not finished. SF Express generates the label only after the warehouse picks and packs and the outbound order is approved; before that the API returns nothing, even when the web page shows it. The final answer was a callback from SF: live in production on the evening of 9 September, with label paths filling into work orders automatically from 11 September.

Label mechanism: create outbound order → warehouse picks and packs → outbound approved and label generated; polling the query API did not work; final answer SF pushes a callback
Creating an order and getting the label are two separate actions.
SF integration group screenshot: SF explains the label is based on the actual warehouse pack-out and the API only returns the label URL after outbound
SF’s integration contact on 31 August: the label follows the actual pack-out. Group name, names and avatars pixelated. (Chinese-language original.)
07

A ‘false failure’

On the evening of 2 September an order’s quantity changed from 1 to 6; after cancel-and-rebuild the API reported ‘no stock’ and the work order was marked failed. The agent checked itself, found the rebuild had switched SKU, and queried SF: the order in fact existed with a waybill. Its conclusion went into the workflow rules: ‘Do not retry or create another order, to avoid shipping twice.’

False-failure sequence: 19:35 cancel and rebuild → 19:48 API error, marked failed → 19:57 agent self-check finds the SKU switch → 19:59 query SF, order really exists → conclusion: do not retry
A system error does not mean SF did not create the order. On any error, query the other system by transaction number first.
08

A formula that was wrong for a month

On 30 July the client wanted the cost basis in the profit calculation changed from average to latest cost, and we changed the formula. The report went out at 21:00 every day, complete, with no step reporting an error, while the margin sat between 74% and 96% for a month. The client flagged it on 29 August; we fixed it on 1 September. The lesson is ours: a change to a core formula needs an ‘outside the normal range’ alert in the same commit.

Margin line chart: 18.28% on 30 July, jump to 85% from 31 July, 74% to 80% through August, 74.31% on 29 August, 20.46% after the fix on 1 September; normal range 17% to 24%
The margin shown on the daily report; values from the report cards.
Lark Base store report screenshot: Profit Margin column shows 77.82%, 79.74%, 74.65%, 74.71%, 60.69%, total 74.31%; store names and amounts pixelated
The report of 29 August. Complete and on time, which is exactly why nobody looked twice for a month. Store names and amounts pixelated.
09

Why does the bot ‘forget’?

The longer a conversation, the more likely early rules are dropped in automatic compression. ‘It did not forget. It can no longer see it.’ So rules go into skill files and Base tables, never only into chat. This diagram, sent to the client on 12 June, became our default explanation on every project.

Eight-panel illustration: a user asks for a store code, the bot answers from early dialogue; after many turns the history is compressed and the code is dropped; asked again the bot cannot answer; it did not forget, it can no longer see it
How context compression works. The fix is to write rules into files. (Chinese-language original.)
10

The engagement: build one line in Q1, then coach

The first quarter is fully managed: we build, tune and integrate. Quarters two to four are coaching: we provide plans, answers and reviews, and the client’s team implements. The aim is for the client to grow its own AI capability. SF ordering was added to the first quarter after the client made clear that the step consuming staff hours had not yet been freed; that is what ‘build one line and build it through’ means. From 10 September, we coach.

Service-phase table: Q1 fully managed, provider does all building and integration; Q2 to Q4 coaching, plans and answers only, implementation by the client
The service-boundary table from the 9 July review notes. (Chinese-language original.)
11

Five lessons

Looking back, not one of the hardest parts was model capability.

  • Preconditions before development: A compliant dedicated WeChat account, prepared a month ahead, at the top of the kickoff checklist.
  • Judge compliance blockers early: A week of research and a change of road beats three weeks of drift.
  • Creating an order and getting the document are two things: The label exists only after outbound; rely on the carrier’s callback, not our polling.
  • A system error does not mean nothing happened: On any error, query by transaction number first, then decide whether to retry.
  • Core formula changes need threshold alerts: Set the alert in the same change. That part is on us.
12

Where is it now?

On 11 September the daily work orders for WeChat ordering, verification, SF order creation, ID archiving and label return were running normally, and the client’s operator was asking for the next step: an order summary every hour or two with batch label download, and the dozen customer statements she updates daily. China daily sales and the store report go out at 21:00 every day, with margins back in the normal range.

Evidence boundary: numbers and quotations come from the project group chat and the meeting notes of 9 July and 1 September, each dated in the text. Daily sales figures for stores and channels are not given; SF transaction, order and waybill numbers, ID documents and recipient details were not used. The client is a New Zealand cross-border health-products retailer and is anonymised; company names, personal names, avatars and numbers in screenshots are cropped or pixelated; diagrams were drawn for this article.

FAQ · Quick answers

In this New Zealand cross-border e-commerce AI case study, can a WeChat order automatically produce an SF Express label?

Yes, in two steps. The agent splits and verifies the plain-language order and calls the SF Express ISCO API to create the outbound order within minutes. The label is generated only after the SF warehouse picks and packs and the outbound order is approved; SF then pushes it through a callback, and the agent archives it and writes it back. ‘Label at the moment of ordering’ is not possible under SF’s mechanism.

What does the AI do in this case, and what do people do?

The AI parses orders, maps SKUs, splits by carton rule, issues verification cards, archives IDs, calls the SF API, writes back numbers and labels, and compiles daily sales and store reports. People check recipient, product, payment and ID and say ‘review passed’; decide whether customs documents are needed; and decide whether to retry on an error. Without a human instruction the agent does not call the production API.

Do we need to replace the ERP, or move off WeChat and Lark?

No. The client’s ERP, SF Express ISCO, WeChat and Synology NAS were all kept. Agents run on the client’s own Mac mini, orchestrated by OpenClaw, and data lands in the client’s own Lark Base. The only addition is a dedicated WeChat account kept logged in inside China.

How long does WeChat-to-SF-Express ordering take to get running?

SF scope fixed 31 July, API connected 26 August, full test flow 28 August, first production order 2 September, label callback 9 September: about six weeks. Development was less than half; the rest went on SF production parameters, confirming the label mechanism and WeChat account access.
DT
BEE Sigma Delivery

The front-line team that plugs workflows into the systems businesses already run — methods drawn from delivered projects.

Read Next
AI Implementation Case Study · NZ ConstructionReviewed 14 Sep 202612 min

New Zealand AI Implementation Case Study: Site Timesheets and Job Costing for a Construction Company, Run from a Chat Group

AI implementation case study, New Zealand (Lark Base, OpenClaw, Mac mini): a South Island construction company where hours were recalled by foremen after work and totalled in Excel, and project information lived in WeChat groups and email. Three months later a foreman says one sentence in a chat group, an agent writes the timesheet line, and the owner sees 20 projects, 10,000+ hours and 334 daily records on a cost dashboard. This is the story of the three months in between: a first month stuck on stability, a client saying the AI ‘improved too little’, and data errors caught by hand every day after go-live.

Read article
AI Implementation Case Study · NZ ConstructionReviewed 11 Sep 202611 min

New Zealand AI Implementation Case Study: 91% Invoice Automation at a Construction Company

AI implementation case study, New Zealand (Outlook, OneDrive, ApprovalMax, Xero): five hundred to a thousand supplier invoices a month, one quantity surveyor sorting them, matching purchase orders and submitting approvals by hand. Three months later, in one full accounting month, 730 of 798 invoice documents were recognised and routed by AI, and manual review fell to 68. This is the story of the three months in between: recognition was never the problem, the invoice count would not reconcile, and what the client actually wanted.

Read article

Hand your first workflow to the agents

One free scan shows how visible you are in the AI era; the AIM assessment finds your best angle of adoption.