We build a grant reconciliation app for nonprofits. Staff log their expenses, attach receipts, and at the end of each month the app puts together a packet for the funder. To save typing, the app reads each receipt with AI and suggests the subtotal, tax and total.
Most of the time this works well. But receipts in real life are not clean scans. They are phone photos taken in a dim restaurant, folded in a wallet for a week, or shrunk down by a chat app. So we decided to test the reader on the kind of receipts people actually upload.
The problem: confident and wrong
One test receipt had a total of $204.75. We made a low resolution copy of it, the kind of photo you get when an image is shrunk and then blown up again. The reader looked at it and suggested a total of $22.50.
There was no warning and no sign of doubt. It just gave a clean, believable number that was off by about $180.
This is the worst kind of mistake an AI feature can make. If the reader says nothing, the person types the amount in themselves. If it says something wrong and sounds sure, the person may accept it, and the wrong number ends up in a report that goes to a funder.
So the goal became clear. When the reader can read a receipt, it should get it right. When it can't, it should say so.
How we tested it
We made a small set of test receipts from one clear original. There was the clear version, a lightly blurred one, a medium blur, a heavy blur and a low resolution copy. We also made a receipt with the totals cropped off and a file with two receipts in it.
Then we ran each one through the real reader several times. Running the same file more than once matters. An AI model does not always give the same answer twice, and a feature that works one time in three is not a feature people can trust.

The results were mixed. Medium and heavy blur were handled well, since the reader declined to give an amount. The clear receipt was mostly right. But the light blur and the low resolution copies were unpredictable. Sometimes the read failed, sometimes it gave a close answer, and sometimes it gave a confident wrong one.
The hidden cause: the AI ran out of room
When we looked at the failed reads, the cause was not what we expected. It was not mainly the blur.
The model we use thinks through a problem before it answers, and that thinking counts toward a limit on how much it can write back. Our limit was 600 tokens. On a hard photo, the model spent all 600 working out what it was looking at and ran out of room before it could give an answer.
About half of the photo reads were being cut off this way. It also explained why the same photo could pass on one run and fail on the next. Some runs needed a little more thinking than others, and those were the ones that hit the limit.
Three fixes
We made three changes to the reader.
More room to think. We raised the limit from 600 to 2000 tokens. The model now has space to work through a hard photo and still give a full answer.
Full detail on photos. The reader now sends photos at full detail. Before this, a photo could be scaled down before the model ever saw it, which threw away the small print that matters most on a receipt.
A clear rule to decline. We told the reader directly that if it can't read the amounts with confidence, it must say so and return nothing. A guess is not allowed.
After these changes we ran the same tests again. The low resolution copy that had produced $22.50 now declined on every run. The clear receipt gave the correct amounts on every run. Medium and heavy blur still declined, which is what we want.
Cleaning up blurry photos first
That left the lightly blurred photos. These are the most common problem in practice, and they are often readable by a person, so declining every time felt like giving up too early.
We tried cleaning the photo before sending it. We tested a few versions on a lightly blurred photo, ten runs each. Enlarging the photo on its own gave the right amounts once in ten runs. Enlarging it, boosting the contrast and then smoothing out the noise gave the right amounts five times in ten. Most of the other runs declined, which is the safe result.
Some ideas made things worse. Boosting contrast without smoothing afterwards caused every read of the blurry photo to fail. Sharpening the photo did the same. The lesson for us was to test each step on its own and not assume that a clearer looking photo is easier for the AI to read.

Cleanup does have a ceiling. Medium blur and very low resolution photos still could not be read, whatever we did to them. For those the reader declines, and the person types the amount in.
What we accepted, and why it's fine
The reader is not perfect, and we wrote down exactly where it still falls short.
A lightly blurred photo can still be off by a few cents. A receipt with a fancy script logo can still produce the wrong shop name. Heavily blurred and tiny photos are not read at all.
We are comfortable with this because of one rule the whole app follows. The AI suggests, the person confirms, and nothing saves on its own. Every amount the reader finds is shown as a suggestion that the person checks before applying. A few cents off is easy to spot when you are looking at the receipt. A confident $22.50 on a $204.75 receipt is much easier to miss, and that is the mistake we removed.
We also dropped a feature we had planned. We thought we might need a way to pick and add up single items on long receipts. Testing showed that real receipts print their own totals, and the reader already reads those. Not building it kept the app simpler.
If you are adding AI to your own product, our advice is short. Test it on the messy inputs your users really have, run each test more than once, and make sure the AI is allowed to say "I don't know." If you want help doing that for your app, talk to us at Mantaq.



