ProjectsProject Fongoing
Language models in daily operation: no text without a number check
A daily briefing for every user, a short profile per customer and a weekly report per salesperson, written by a local language model. Every number in the text is checked against the data. And when AI agents reviewed the platform, I checked the reviewers: 3,258 statements, 82 of them wrong.
- Role
- Design, implementation and operation, alone
- Period
- since 09/2026
- As of
- 10/2026
- Setting
- Platform from project A, without a cloud service
- texts in the sample passed the number check
- 46 of 46
How measured
Every number in a generated text must appear in the data the text was made from. A text that fails is not shown.
- active customers with a short profile
- 616
How measured
Number of customers for whom the platform generates a short profile from its own data. On top of that, 16 salespeople get a weekly report on Mondays.
- statements from the AI agent review verified
- 3,258
How measured
The agents wrote 13 reports. I checked every statement against code and data before any of it became a task.
- of these statements were wrong, 311 imprecise
- 82
How measured
Result of the verification. The wrong statements did not make it onto the board, the imprecise ones only after correction.
Short version: the first paragraph of each section, plus key figures and chart.
Starting point
The platform from project A shows many numbers on many pages. Anyone who wants to know in the morning what has changed has to collect them. A language model can summarise that. It can also invent a number that sounds plausible. In a company that buys and slaughters by these numbers, an invented number is worse than no text at all.
A second question came on top. AI agents can review a whole platform and write reports. How much of it is true?
Decision
First: the model runs locally, on a workstation with a graphics card. No customer data leaves the house. I measured which model to use beforehand. I describe the comparison on the page How I work.
Second: the model writes, but it does not count. The numbers come from the platform. The model gets them as context and phrases the text. Afterwards a rule checks every number in the text against that context. If one does not match, the text is not shown.
Third: what the model cannot do reliably, it does not do. When a case shows that a sentence can go wrong, that sentence comes from fixed code from then on.
Fourth: agent reports are leads, not findings. Every statement is verified before it turns into work.
Implementation
The start page shows every user each day what is new, what remains open and what has been settled. The list is filtered by their permissions and their customers. Above it sits a summary from the language model whose numbers are checked against the list. At launch 39 of 42 users saw this briefing.
Sales gets two texts. Each of the 616 active customers has a short profile on its customer page. 16 salespeople get a weekly report on Mondays: revenue against their own average and the same week last year, the strongest gains and declines per customer, open alerts. In a sample, 46 of 46 texts passed the number check.
One case still stood out. A text said “no alerts” although two existed. No number was wrong, so the check did not fire. I did not patch that with a second check. The sentence about alerts no longer comes from the model at all, but from the platform itself.
At the end of September I had the whole platform reviewed: features, interface, service, security, performance, data and models, operation. AI agents working in parallel wrote 13 reports. I verified all 3,258 statements in them. 82 were wrong, 311 imprecise. Only then did around 40 tasks go onto the board, 2 of them critical. A load test was part of it: up to 30 simultaneous users the platform ran without errors, and the bottleneck is named.
I use agents for building as well. On one evening in September, 20 agents worked on tasks in parallel, each in its own set of files so they would not overwrite each other. Judgement and acceptance stayed with me: tests, visual checks, comparison with the real data.
Result and measurement
| Measurement | Value |
|---|---|
| texts in the sample passing the number check | 46 of 46 |
| customers with a short profile | 616 |
| salespeople with a weekly report | 16 |
| agent statements verified | 3,258 |
| of which wrong | 82 |
| of which imprecise | 311 |
The number check does what it promises: no text with a number that is not in the data. It does not stop everything. The alerts case shows that a text without a wrong number can still be wrong. That is why every sentence that has shown such an error moves from the model into fixed code.
For the agents the result is clear. Unchecked, 82 wrong and 311 imprecise statements would have gone onto the board as tasks. The verification is not an extra, it is the actual work step.
What I have not measured: whether the texts are read and whether they change decisions. So I claim nothing about it.
What I would do differently
Keep the list of sentences that must never come from the model from the start. Statements about the absence of something belong on it. A model easily says “nothing”, and no number check notices.
Make the sample larger and regular. 46 texts are a good start, but a one-off measurement.
Tell the agents beforehand how to back up their claims. A statement with a location in the code or the data can be checked in minutes. One without has to be searched for first.
Technology
- Local language model
- Python
- FastAPI
- React
- AI agents