Skip to content
syncra
All notes

AI features that hold up on real data

A demo needs one good answer. A production feature needs a sensible result for every input real users send it. Here is what I've learned building AI features into products people use.

By , founder6 min read

Over the last few years I've built AI features into products that real people use every day. As a contract senior engineer on FreeLogo.com (opens in a new tab), an AI logo and brand platform in the HostPapa family, I built the AI logo generation and the vector editor that sits next to it. Before that, at GWIN AI, I worked on a cross-platform financial app with an assistant that fills in-app widgets with live data.

Neither is a Syncra project. Both were products owned by the companies I worked for, and their internals stay with them. What I can share is the engineering practice I now bring to every AI feature we build at Syncra. None of it is exotic. Most of it is ordinary software discipline, applied to a component that is slow, probabilistic and occasionally confidently wrong.

The gap between a demo and a feature

A demo runs on inputs you picked. Production runs on whatever arrives: text pasted from a PDF with broken line endings, a request in three languages at once, an empty field, a user who presses the button twice, someone who types "ignore your instructions". A model that looks brilliant on your ten examples will meet all of these in its first week.

So the question is never "can the model do this?" It is "what does the product do when the model doesn't?" Every lesson below is a version of that question.

1. Treat model output as untrusted input

Model output is user input that happens to come from your own server. Validate it at the boundary, exactly as you would validate a form submission.

Ask for structured output wherever the result feeds code rather than a human reader, and check it against a schema before anything else touches it:

const ChartRequest = z.object({
  metric: z.enum(["revenue", "orders", "visitors"]),
  period: z.enum(["week", "month", "year"]),
});

const parsed = ChartRequest.safeParse(modelOutput);
if (!parsed.success) return showFallback();

Decide in advance what happens when validation fails: one retry with the validation error in the prompt, then a defined fallback state the UI knows how to render. "The model returned something odd" should be a handled case with a designed screen, not an exception in your logs.

2. Let the model choose. Let your code supply the facts

When an assistant's answer ends up in a UI, such as a chart, a card or a widget, split the work cleanly. The model decides what to show and with which parameters. Your code fetches the actual numbers from your own data and renders them.

The model should never type a figure that your database already knows. This one rule buys you a lot:

  • Correctness. Figures come from the source of truth, not from a model's guess at it.
  • Freshness. The widget shows the data as it is now, not as it was in a prompt.
  • Permissions. Your data layer already knows what this user may see. A model with no direct data access can't leak what it never received.
  • Testability. You can test "the model picked the right widget" separately from "the widget shows the right data".

Tool calling makes this natural. Give the model a small set of well-described tools with strict parameters, and treat its job as routing and phrasing, not arithmetic.

3. Design the wait

Generation takes seconds, not milliseconds, and the time varies from request to request. Treat it as a job, not a function call:

  • Start the work, return immediately, and show honest progress.
  • Let people cancel, and stop paying for work nobody is waiting for.
  • Persist the result, so a refresh or a dropped connection doesn't throw it away.
  • Make retries idempotent, so a double click doesn't run, and bill, the same generation twice.

Stream text when partial output is useful to read. For images and other all-or-nothing results, give the waiting state real design attention. It is part of the feature, and on a slow day it is most of what the user sees.

4. Put the output somewhere people can fix it

AI output is rarely exactly what someone wanted. If the only recourse is "regenerate", users end up rolling dice until something is close enough.

The stronger pattern is to land the result in a form people can edit: a generated logo opens as editable vector shapes, a drafted reply lands in a text field, suggested values fill a form the user still submits. Generation gets someone to a good starting point quickly. The editing surface is where they finish, and it also covers for the cases where the model missed.

This shapes the output format too. If people will edit the result, ask the model for something structured and editable, not a flattened final artifact.

5. Ground answers in your own content

Many "the AI made something up" problems are really retrieval problems. If the answer should come from your documentation, your catalogue or your records, find the relevant pieces first and give the model only those.

Embeddings and semantic search do most of this work, and they don't need much infrastructure to start: a vector column in the Postgres you already run, such as pgvector, covers a lot of products before you need anything dedicated. Spend the effort on what goes in. Chunk content along its real structure, keep each chunk's source with it so answers can cite it, and re-index when the content changes.

6. Build an evaluation set from real inputs

Before launch, collect a set of real or realistic inputs, and deliberately include the ugly ones: empty, very long, the wrong language, ambiguous, adversarial. For each, write down the properties a good answer must have rather than an exact expected string: valid against the schema, mentions the right entity, refuses when it should.

Run the set on every prompt change and every model change. Prompts are code. Version them, review them, and never edit one in production because a single example looked better. Model upgrades deserve the same treatment: a newer model can be better on average and still worse on the cases your product depends on.

7. Plan for the provider having a bad day

Providers time out, hit rate limits, have outages and retire models. Put the provider behind your own small interface, so the rest of the codebase never calls a vendor SDK directly. Set timeouts. Decide on a fallback: another model, a cached result, or a clear message with a way to try again.

Log enough to debug: the prompt version, the model, the latency, the validation result. Keep personal data out of those logs or redact it, since prompts collect exactly the information people would rather not have stored. And watch cost per request the way you watch latency. It changes with every prompt edit.

What this looks like at Syncra

We build applied AI: assistants, semantic search, retrieval and AI-powered workflows connected to a client's own data and tools. We don't train models. Every AI feature we ship starts from the practice above: validated output, facts from your systems, a designed wait, an editing surface where it fits, an evaluation set and a fallback for the day the provider fails.

None of it makes a demo more impressive. All of it decides whether the feature still works a month after launch.

More notes