Maciej Nuzia · AI/LLM integrations

I integrate AI models with applications and handle provider limits, costs and outages

A production model call needs more than an API key and a prompt. The application has to handle timeouts, rate limits, retries and the cost of each request. This layer also includes tests that show whether a model or prompt change affected the responses.

This page is about models called through a provider's API. A model running on hardware of your own is a different job.

Production

A prototype works for one call. Production requires more safeguards

In production there are several calls at once, some of them come back refused by the provider, and for others no answer arrives at all. Then there is the invoice at the end of the month: one figure that somebody in the company tries to map onto features of the application. All of this plays out on the provider's servers.

  • the key sits in the code, because that was the quickest way to start
  • under real traffic the provider refuses some of the calls and nobody planned for it
  • the bill arrives as one figure, with nothing to show what pushed it up
  • the prompt was edited straight on production and nobody knows how it read last week

Who it's for

For teams moving an AI prototype into production

There is a key already, there is a prompt, and there is one feature somebody in the company likes. Then comes real traffic, an outage at the provider, and the invoice at the end of the month. That is where I come in.

  • Teams with a prototype ready for production
  • SaaS products with a model inside
  • Applications where the model answers a user live
  • IT departments that inherited a finished integration
  • Companies running on a single provider
  • Teams a few months into paying for calls

The connection layer

What has to work around every model call

A key you can swap without shipping a release

The key lives in configuration and the application picks it up from there at startup. When it has to be revoked, the swap is an edit in configuration and it does not wait for somebody to ship the application again.

A refusal that shows up only under load

A provider counts requests inside a time window and refuses whatever goes over. As traffic grows, calls line up in a queue and come back spaced out, because an immediate retry earns the same refusal.

Silence on the far side

Providers have outages, and the application needs its own behavior for them: what the user sees meanwhile, and what becomes of the work that was waiting on an answer.

Who changed the prompt, and from when

A prompt is part of the system and lives in the repository next to the code. A change goes through the same review and release as any other fix, so every version carries a name and a date.

Questions you already know the answer to

A list of questions with the reply you expect, walked through on every prompt change and every model change. Without it, editing a prompt is shipping on a hunch.

The call log

What went to the model, what came back, how long it took and what it cost. A week later that is the only place where an argument about one particular answer can be settled.

What a model actually does for a company belongs on other pages. Answering from documents is covered by RAG for business, entries on a customer record by the ChatGPT and CRM integration, and back-office tasks by AI process automation. The same connection layer sits under all of them.

Before real traffic

Before launch, I decide what the application does without a model response

The order follows from one thing: the provider sets the terms here.

01

Where the call will sit

I read the code at the point the request leaves from, and ask whether a user is waiting at the screen or the work runs in the background. That settles how much delay is acceptable and whether anything can be pushed onto a queue.

02

Account, keys and a spending cap

I keep the keys apart, one for tests and one for production, both issued on your account. Before real traffic starts I check what spending limits and alerts your provider offers, and set up what is there.

03

The call, the retry and the queue

A call gets a deadline after which I stop waiting for it, and a retry spaced out from the first attempt. Anything nobody is waiting on goes to a queue and comes back at its own pace. Every retry costs on its own, so only the calls known not to have landed get repeated.

04

The set of cases and the log

Before anything reaches production I put together a set of cases, each with the answer it should produce. From then on no prompt or model change goes out without a pass through that set. After launch I read the call log: what came back with an error, what took longest, and where the cost gathers.

Technologies

I keep the provider connection in one module

Calls live in one place and the rest of the application talks only to it. Changing the model then means rewriting that module, and the prompts get written again for the new one, because every model reads them its own way.

  • OpenAI API
  • Claude / Anthropic
  • REST / GraphQL
  • Node.js
  • Python
  • PostgreSQL
  • Queues and retries
  • Webhooks
  • AWS

The provider's terms

The provider can change how the integration has to work

A model version with an end date

Providers retire older model versions and set the date themselves. After the move to a newer one, some of the answers come out differently, sometimes better. That is what the set of cases is for: so the difference shows up with me before it shows up with your user.

A retry counted twice

Your side stops waiting and sends the request again. On the provider's side the first one went through and was counted. So any call that changes something or costs something carries a marker that lets it be recognized as the same job.

A key in the repository history

A key committed once stays in the history even after it is taken out of the file, so the only way out is revoking it with the provider. I watch the logs the same way, because full call logging leaves them holding whatever went to the model.

FAQ

Questions about how the integration works and what it costs

Whose provider account does this run on?

Yours. The account and the keys stay with the company that pays for the calls, so usage is visible to you in the provider dashboard without asking anyone. A leaked key has to be revoked straight away, which is why I do not keep it in the code. Who gets to read it is decided on your side.

What does the application do when the model does not answer?

That gets settled before launch, because a provider has its own outages and its own rate limits. The application waits for an answer up to a set limit and then reaches for it once more, after a pause. If nothing comes back even then, the user gets a message and the job moves to the queue or over to a person. What exactly the user sees depends on whether they are waiting at the screen or the work runs in the background.

Can this be moved to a different model later?

Calls to the provider live in a single module, so a change means rewriting that module. The prompts get written again for the new model and taken through the set of cases. How many answers come out differently shows on that pass, and that is when the cost of the move becomes known.

What drives the bill for the calls?

The number of calls and how much text goes into each one. The heaviest part is usually the context attached to the question: documents, conversation history, the instruction for the model. All of it travels to the provider with every request and all of it is counted. I am not going to reprint provider rates here, because the price list belongs to them and it changes without me. What the bill is doing shows in the call log as it happens, with a cost against every entry.

What is kept from each call?

That is something to settle together, and it is a decision of its own. Storing whole requests along with the answers helps when somebody asks why an answer read the way it did. It has a price: the text sent to the model then also sits in your logs, under the same rules as the data in your database. The other option is markers alone: prompt version, size, duration, cost. That choice comes back every time somebody asks about one particular answer from last week.

One sentence

Show me where the application calls the model

One sentence by email is enough: where in your code the model call sits today, and what a user sees when no answer comes back. That is already enough to say what I would start with. The estimation tool breaks whole projects into parts, so reach for it when the rest of the application has to be built alongside the model.