- trends
- images
- small business
Agents that see photos: the fault, the receipt, the note
Customers send photos instead of typing. What an agent can pull from them, where it gets things confidently wrong, and how to test it in twenty minutes.
A customer sends you a photo of the boiler with an error code on the display. Someone else photographs a garage receipt so you can put it through expenses. The driver sends the signed delivery note from his phone: crooked, half in shadow, with a thumb in one corner. All three did the same thing — sent an image instead of typing — and all three assumed someone at the other end would sort it out. That someone has always been you.
It no longer has to be. This is for small-business owners who get photos every day on WhatsApp or by email and work through them by hand. By the end you'll know which kinds of photo an agent can turn into finished work, where it actually gets things wrong, and what to check before you trust a single extracted number.
From template-based OCR to a model that looks
Text recognition has worked for decades, but it worked in one specific way: you told the system "the invoice number sits in this region of the page and the total in that one", and from then on it read. One template per supplier. If the supplier redesigned the layout, it broke. If the paper was skewed, it broke. For a small business with twenty different suppliers, building and maintaining that was never worth it.
Today's models don't read positions — they interpret the whole image. They understand that the thing in the top right is a date, that the narrow column is amounts, and that the big number at the bottom is probably the total, even on a delivery-note layout they've never seen. No templates, no training. That's why it finally fits a small business: the barrier was always setup, and the setup is gone.
What hasn't gone is the error rate. OCRBench v2, an academic benchmark that tests multimodal models on real documents — invoices, forms, handwriting, tables — evaluated 38 models, and most scored below 50 out of 100 on average. Take that at face value: reading a photo of a real-world document is still hard, and anyone quoting you 99% without saying which documents they measured is selling you a story.
Four photos that are already work
This isn't about digitising the filing cabinet. It's about images that arrive on their own, every day, and currently make you stop whatever you were doing.
- The fault photo. A repair service, a garage, a plumber. The image carries the make, the model, often an error code and almost always a hint of how big the problem is. The agent classifies it, checks whether it's the kind of job you take, asks for what's missing ("I need to see the plate with the serial number") and offers a slot before anyone picks up the phone.
- The expense receipt. Date, supplier, net, VAT, total. The agent pulls the five fields and stages the expense line for you or your accountant to approve. This isn't automated bookkeeping — it's the entry arriving half-written instead of living in your camera roll.
- The signed delivery note. Here the value isn't the text, it's the comparison: match what was delivered against what was ordered and flag only the mismatches. Out of twenty notes a day it shows you the three that don't reconcile and stays quiet about the other seventeen.
- The shelf or label photo. A supervisor photographs what's left, the agent identifies the references and drafts the reorder. It works surprisingly well for products already in your catalogue and fairly badly for anything that isn't.
Notice the pattern: in all four cases the photo already existed. Nobody is being asked to change a habit. The only thing that changes is what happens to the image between arriving and being looked at.
What a photo gives you, and what it doesn't
A photo gives you three things: data, evidence and a timestamp. The data is what's written. The evidence is that it existed — the boiler looked like that, the note was signed. And the timestamp is when, which in a dispute is sometimes worth more than the contents.
What it doesn't give you is intent. The receipt doesn't say which job that cost belongs to. The fault photo doesn't say whether the customer wants a quote or wants you there today. That part is still a conversation, which is why an agent that only "reads the image" is of limited use. The useful one reads the image and then asks the missing question.
The photo settles what's there. What to do about it is still yours.
Where it breaks, no varnish
Four things you'll hit in week one that never show up in a demo.
It gets things wrong confidently. This is the serious one. A person who can't quite read a digit tells you it's illegible. A model returns a plausible digit: a 3 that was an 8, a total that was actually the subtotal. If that value flows straight into a charge, nobody catches it until the complaint arrives.
Image quality decides everything. WhatsApp recompresses photos sent in chat, and Meta's API documentation caps images at 5 MB for JPEG or PNG. In practice: a long receipt shot from a distance in bad light arrives with the amounts turned to mush. "Get closer and make sure the total is readable" fixes more cases than switching models does.
Handwriting is a different sport. Detecting that a signature is present works well. Reading a biro note scrawled in the margin of a delivery note, much less so. If your workflow depends on that, test it on your own paperwork before promising anything.
There is personal data in the frame. A photo of a delivery note comfortably contains an address, a name and sometimes an ID number. You're processing personal data exactly as if it had arrived through a form — same questions about lawful basis, retention and who does the processing. Not a blocker, just a box to tick beforehand rather than afterwards.
When not to build this
If you get three photos a week, build nothing. The setup plus watching it closely for a fortnight won't pay back inside a year. The realistic threshold is closer to ten or fifteen images a day of the same type.
Don't build it either where the extracted value becomes money without anyone looking. If a figure comes off a photo and gets invoiced or paid automatically, you've put a probabilistic reader in the middle of a process that demands exactness. There the agent should fill the field and stop; a human validates. And be wary of "let's read everything" — the flows that survive are narrow and boring: one photo type, one clear destination.
How to test it this week
Take twenty real photos you've already processed by hand, bad ones included. Run them through and compare field by field against what you did. Twenty minutes will tell you two things no demo will: what share comes out right and, more importantly, what the failures look like. Obvious errors are manageable. Quiet, believable ones are what cost money.
If you like what you see, start with a single image type and leave the agent in propose mode: extract, fill in, put it in front of you to approve with one tap. After a month without surprises you can decide which fields have earned the right to go through on their own.
That's how we set up the operations agent at Yaqbot: it takes in what arrives — photos included — drafts the quote or the expense line and leaves it ready for you to check. You approve, it executes. What you get back isn't technology. It's the half hour a day you currently spend opening images one at a time.
