All notes
7 min read

Custom AI development: how to add AI features to an existing app

What custom AI development usually means today, which AI features are worth building, how to architect them safely, how to test and cost them, and what to ask an AI development company.

AAAsghar AliFounder & Lead Engineer · Daniotech
A request path from left to right. The user's app sends a request to your server, which checks who the user is and applies limits; the API key stays there. The server fetches only the records this user may see from your data, sends them with the request to the model API, then checks the answer: format, rules and confidence. The checked answer goes back to the app. Below, three notes: keep a test set and score every change; log cost and latency per feature; let a person review anything that cannot be undone.
A request path from left to right. The user's app sends a request to your server, which checks who the user is and applies limits; the API key stays there. The server fetches only the records this user may see from your data, sends them with the request to the model API, then checks the answer: format, rules and confidence. The checked answer goes back to the app. Below, three notes: keep a test set and score every change; log cost and latency per feature; let a person review anything that cannot be undone.

"Custom AI development" sounds like training your own model. For most products it is not. The large language models available through an API are already better at general language tasks than anything a single company would train, so the work is building the product around a model: choosing the job it does, giving it the right data, checking what it returns, and keeping cost, privacy and failure under control.

That work is ordinary software engineering with a few new failure modes. This guide covers which AI features are worth building, how to architect them, how to know whether they work, and what to ask before you hire an AI development company.

Start with the job, not the model

Pick one task that a user or your team does today, that is slow or repetitive, and where a mostly-right answer is useful. The common kinds:

Feature What the model does Works well when
Search and answers over your content Finds the relevant documents, then answers from them (often called retrieval-augmented generation, or RAG) The answer is in your data and you can show the source
Extraction and classification Turns messy input (emails, PDFs, forms, tickets) into fields or categories The output has a fixed shape you can check
Drafting Writes a first version: a reply, a summary, a description A person edits before anything is sent
Speech Transcribes audio, or reads text aloud Accuracy can be spot-checked and corrected
Agents and tool use Calls your own functions (look up an order, book a slot) in several steps Every action is limited, logged and reversible

Be wary of features whose only description is "add AI". If you cannot say what the user does differently afterwards, and how you would measure it, the feature is not ready to build.

Build, buy or fine-tune

  • Use a hosted model through an API. The default. You pay per use, get the strongest models, and change provider without rewriting your product if you keep the model behind one interface in your code.
  • Run an open-weight model yourself. Worth it when data cannot leave your infrastructure, when volume makes per-use pricing expensive, or when you need a fixed model version for years. You take on hosting, scaling and updates.
  • Fine-tune. Useful for a narrow, repeated format or style. It rarely fixes missing knowledge; giving the model the right documents at request time usually does that better and stays current.
  • Buy a finished product. If an existing tool does the job, for example a support chatbot product or a transcription service, buying it is often cheaper than building and maintaining your own.

An architecture that holds up

The diagram above is the shape most AI features should have.

  1. Calls go through your server, never straight from the app. The API key stays on the server. The server knows who the user is, applies rate limits and a spending cap, and logs what was asked.
  2. Retrieval respects permissions. If the model answers from your data, fetch only the records this user is allowed to see, before the model sees them. A model cannot be trusted to keep a secret it has been given.
  3. Treat the model's output as untrusted input. Validate the format (a JSON schema for extraction), check it against your rules, and never pass it straight into a database query, a shell command or an HTML page.
  4. Stream long answers, time out slow ones. Users accept a wait when they can see progress. Set a timeout and a fallback for when the provider is slow or down.
  5. Keep people in the loop for anything that cannot be undone. Sending an email, refunding a payment or changing a record should need a confirmation, at least until you have the evidence that it is reliable.

Prompt injection

Any text the model reads, such as a web page, an uploaded document or an incoming email, can contain instructions aimed at the model. You cannot fully prevent a model from following them, so design so that it does not matter: give the model only the tools and data the current user is allowed, require confirmation for actions, and never give it secrets.

Data and privacy

  • Know what you send to the provider. Read its data-use and retention terms, and pick the settings or plan that match your obligations.
  • Send the minimum. Strip or mask personal data the task does not need.
  • Tell your users. Say which features use an AI provider and what is sent, in your privacy notice and, where it matters, next to the feature.
  • Regulated data needs its own review. Health, financial or children's data brings legal requirements that an AI feature does not remove. Check them before the design is fixed, not after.

How to know it works

AI features fail differently from ordinary code: the same input can give different answers, and a change of prompt or model can quietly make things worse.

  • Build a test set before you build the feature. Fifty to a few hundred real examples (anonymised), each with the answer you would accept.
  • Score every change against it. Prompt edits, model upgrades and retrieval changes all get the same test. Exact checks where you can (fields, categories), a reviewed sample where you cannot.
  • Measure in production too. Track how often users accept, edit or reject the output. That number tells you more than any demo.
  • Decide the bar in advance. "Correct on 95% of the test set, and never wrong about X" is a decision you can make before launch. "It looks good" is not.

Cost and speed

Hosted models charge by the amount of text in and out, so cost grows with use and with how much context each request carries.

  • Log cost and latency per feature from day one, so the bill never surprises you.
  • Send less. Retrieve the five relevant passages, not the whole manual.
  • Use the smallest model that passes your test set, and route only the hard cases to a larger one.
  • Cache answers to repeated questions and reuse unchanged context where the provider supports it.
  • Cap spending per user and per day in your own code, not only in the provider's dashboard.

Questions to ask an AI development company

  1. What will the first release do, for whom, and how will we measure it? You want a narrow feature and a number, not "an AI platform".
  2. How will you test it? Ask to see the test set and the scores, and how a change is checked before release.
  3. Where does our data go? Which providers, under which terms, and what is stored, logged or used for training.
  4. What happens when the model is wrong? The fallback, the review step and how a user corrects it.
  5. What does it cost to run? An estimate per user or per request, with the assumptions, and the cap you will set.
  6. Can we change model or provider later? The model should sit behind one interface in our code.
  7. Who owns the code, prompts and test sets? All of them should be yours, in writing.

Where Daniotech fits

We build AI features into web and mobile products: the server side, the retrieval, the checks and the interface, on top of hosted model APIs. We start with one job and a test set, keep keys and data handling on the server, and ask for a working demo at every milestone, like any other project we take on. You own the code (how and when is set out in the agreement), and for larger work we recommend a short, paid discovery first.

Frequently asked questions

Do I need to train my own AI model? Usually not. Hosted models handle general language tasks well. Most of the value comes from the data you give them at request time and the product you build around them.

What does an AI development company actually build? The parts around the model: the interface, the server that calls the model, the retrieval over your data, the checks on the output, the evaluation and the monitoring. The model itself is usually rented.

How long does it take to add an AI feature? It depends on the feature and the state of your data. A narrow feature on clean data is a small project; a feature that needs data cleaned, permissions mapped and a review workflow is a larger one. A short discovery gives you an estimate you can rely on.

Is it safe to send customer data to an AI provider? It can be, with the right provider terms, the minimum data, and your users told. Some data, such as health records, needs a specific legal review first.

Can the AI feature work in a mobile app? Yes. The app calls your server and the server calls the model, so the same feature serves web and mobile and the key never ships inside the app.

Where to go from here

If you have an app and a task in mind, send us the task, who does it today and how you would know it worked: contact us and we will reply with the questions we would ask first. If the product does not exist yet, start with MVP development; for ongoing work, see how to hire dedicated developers. And whatever you build, estimating the cost works the same way.

  • #AI
  • #Custom software
  • #LLM
  • #Product scoping

Working on this?

Planning a custom software project?

Send us the problem, the users and the deadline. We will reply with the questions we would ask before writing any code.