I gave an agent a folder that any Klang small business would recognise: eleven months of supplier invoices as photographed thermal paper, a QuickBooks file two weeks behind, and a WhatsApp export where the same distributor is owed money on three different order numbers. Then I asked it to reconcile, rank what was overdue, and draft the follow-ups. It finished in under ten minutes. The month-end packet it produced, the kind you forward straight to your accountant, was correct. This is the part the launch posts get right.

The frontier moved in five weeks. Anthropic shipped Claude Sonnet 5 on 30 June, posting 63.2 percent on SWE-Bench Pro and 81.2 percent on OSWorld, the benchmark that measures whether a model can drive a desktop by looking at the screen. Google cleared Gemini 3.5 Pro for general availability in July. For a Kuala Lumpur SME owner the number that matters sits elsewhere: Claude for Small Business, announced 13 May, now ships fifteen ready-made office workflows, an invoice chaser, a month-end prepper, a margin analyser, all wired into QuickBooks and PayPal, all queuing their work for you to approve before anything is sent or paid.

The agent finishes the workflows that never leave the ledger, and stalls on the ones that touch a person.

So the honest question has moved past whether an agent can write code, to which back-office jobs it can run with nobody in the room.

The line runs where the data does. An agent finishes the workflows that never leave the ledger, and stalls on the ones that touch a person. Reconciliation is closed-world: the numbers are in the file, the rules are fixed, and the model has read more messy chart-of-accounts than any bookkeeper you could hire. Invoice ageing, margin analysis, the P&L draft, the export packet, these run clean because the inputs are structured and the failure is visible the moment a total does not tie. Give the agent a database and it behaves.

Standard speech recognition posts roughly a 42 percent word error rate on code-switched talk, and a wrong transcription becomes a wrong reminder sent to a real supplier.

Give it a supplier and it does not. The follow-ups are where the closed world ends. Anthropic’s own design tells you this, quietly: the invoice chaser drafts the message and queues it for your sign-off, because the reply is the part it cannot govern. In Southeast Asia the reply is rarely clean text. It is a voice note, or a line of Manglish, or a photo of a delivery order with the quantity scrawled in pen. Standard speech recognition posts roughly a 42 percent word error rate on code-switched talk, and a wrong transcription becomes a wrong reminder sent to a real supplier who now thinks you cannot count. Meta learned the same lesson at scale: its WhatsApp business agent, aimed at a base of 200 million small businesses, will be judged in Jakarta and Kuala Lumpur on whether it survives Bahasa Melayu and the region’s habit of switching languages inside one sentence.1

China is running the same experiment louder, and it is worth watching for the shape of the answer, not the volume. The country processed about 140 trillion tokens a day by April, Tencent dropped its ClawBot agent inside WeChat, and provincial governments now subsidise one-person companies built on agents. The framing there is the token as a settlement unit, admin work priced by the million. But the messy edge is identical. An agent inside WeChat still meets a supplier who sends a fifty-second voice memo about a short shipment, and the model still has to decide whether it heard the number right before it acts on it.

The useful way to buy an agent for your back office, then, is to buy the ledger half and keep a human on the doorway. Let it close the month; it is faster and more careful than a tired founder at 11pm. Read the drafts before they go out. The reconciliation is the product. The voice note is still yours.

Footnotes

  1. The demo that reconciles eleven months in nine minutes and the reply that breaks on a Manglish voice note are, of course, the same product on two different afternoons.