AI inside an existing backend
I work with systems that already run in production and cannot be stopped. Not from a blank page.
What people come with
Two states. First: the company decided it is time for AI and nobody knows where to start. Second: there already was a pilot, it was demoed, it never shipped, and nobody can explain why.
The second is more common and costs more, because money and time were already spent before I showed up.
What counts as a result
The metric is defined during the diagnostic and goes into the contract. Here is how it looks:
- Share of requests closed without a human.
- Time from an incoming document to a filled record.
- Cost of processing one request, in money.
- How often the system admitted it did not know instead of inventing an answer.
The last one matters more than it looks. A model that honestly declines to answer is cheaper than a model that lies with confidence.
How this differs from a turnkey bot
A turnkey bot is an interface. I build the pipeline: queue, retries, validation of what the model returned, human fallback, logs that show the reason behind every decision, and a cost meter. The interface is the last step, not the first.
What I do not promise
A savings percentage before the diagnostic
«We will cut costs by 15–30%» is the standard line in this market, and it has nowhere to come from: nobody has looked at your system yet. I quote a number after the diagnostic, and only one I am ready to defend.
That the model will replace a team
Part of the flow goes to the machine, part stays with people. The size of the second part is the project metric, not an awkward remainder, and planning starts from it.
That it works without changing the process
If the operation currently holds together because a person fills in the context, the model will not inherit that context. The process has to be written down, and that work is on your side.
That you need this at all
The diagnostic can end with «do not build this». That is a result, not a failure: it costs 90 000 ₽ and saves the 350 000 ₽ you would have spent on a pipeline you do not need.
Where this practice comes from
I run AI Builders, an open showcase of dissected LLM systems with the textbook underneath it. Every entry is a pipeline, numbers I measured myself, and a list of what the system still lacks. The textbook «AI/LLM Integrations: from the model call to a working feature» is complete: ten parts, 58 chapters. Free, no sign-up, in Russian.
The book carries the same practice I bring into projects: response contracts, a token budget, system behaviour when the model fails, quality checks where two runs give different answers. The difference is that the book is written for everyone, while a project applies it to your system.
Reading a few chapters before we talk is useful: you will see how I think about this work, and the conversation starts further along.
It all starts with the diagnostic
Five working days and 90 000 ₽ instead of a half-million decision made blind.