All insights

AI Product UX

18 Aug 2026 · AI Product UX · Trust · Human-in-the-loop · AI Native

How to design AI features so they feel usable and safe — trust, prompting flows, human-in-the-loop, and the difference between a chat box and a product.

AI is easy to add and hard to live with.

A model in a chat box is not a product. It is a demo with a text field. People do not know what it can do, what it will invent, when a human is still in the loop, or what happens if they trust the wrong sentence.

That is the UX problem. Not the model card. Not the prompt in the repo. The experience of using something that is sometimes brilliant, sometimes wrong, and always confident.

I design AI-native products so the feature feels usable and safe: the job is obvious, the output is labelled, the human can take over, and the product does not pretend a draft is a fact.

If you are bolting “AI” onto a roadmap because the deck has it, stop. If you are trying to make an agent part of a real workflow — healthcare, operations, a founder tool — this is the work.

Chat is a fallback, not a strategy

Blank “Ask anything” is what you ship when you have not designed the job.

Power users will type. Everyone else will stall. They do not know the verbs. They do not know the limits. They will either under-use it or over-trust it.

Prompting flows, done properly, look like product design:

  • A starting job, not an empty field (“Summarise this visit,” “Draft a reply,” “Find the exception”)
  • Suggested next moves after an output, so they are not back in the void
  • Constraints in the UI — tone, length, source, audience — instead of hoping they prompt like an engineer
  • A way to attach the real object (the record, the file, the thread) so the model is not guessing context
  • Edit, retry, and undo as first-class, because generation is a loop, not a submit

The prompt can live under the hood. The person should be choosing a job and steering a result.

If the only interface is a transcript, you have not designed an AI feature. You have exposed a model.

Trust is a UX surface

People do not read your model policy. They read the sentence in front of them.

Trust breaks when the product looks certain. A paragraph with no provenance. A diagnosis-shaped answer. A number with no date. A “done” state on work a human has not seen.

I treat trust as something you can design:

  • Label the generation. Draft, suggestion, extract, prediction — not “the answer”
  • Show what it used. Sources, the record, the window of data. If it had nothing, say so
  • Separate fact from inference. What was retrieved versus what was written
  • Confidence without theatre. Do not fake a percentage. Do say when this class of task is weak
  • Status that stays honest. Pending review is not complete. Suggested is not sent

Pledge My Tree had to be explicit: pledged is not planted. AI products have the same duty. Generated is not verified. The interface should make that impossible to miss.

Healthcare made this non-negotiable on Raiqa — patients, practitioners, agents. If the ecosystem could not stay intelligible, the AI was not a feature. It was a risk.

A UX audit before launch should include this pass: where does the product imply certainty it has not earned?

Human-in-the-loop is a journey, not a disclaimer

“A human reviews this” in the footer is not a loop.

A loop has a place in the workflow: who sees it, what they can change, what happens if they do nothing, and what the user is told in the meantime.

Design the handoff:

  • When the model may act — low-stakes, reversible, bounded (a draft email, a tag, a summary)
  • When a human must act — irreversible, clinical, financial, legal, anything you would not undo in one click
  • What the reviewer needs — the original input, the suggestion, the diff, the reason it was flagged. Not a dump of tokens
  • Time — if review takes a day, the product has to live in “waiting” without looking broken
  • Escape — the user can ask for a person, or the system can escalate, without shame or a dead end

The point of HITL is not to make the AI look responsible. It is to keep the outcome safe when the model is wrong — which it will be.

If your loop is “we’ll watch the logs,” you do not have a loop. You have hope.

Usable and safe at the same time

Safety that makes the product unusable will get bypassed. Usability that hides risk will get someone hurt, or just quietly lose the account.

The tension is the design job:

  • Defaults that are conservative without being hostile (suggest, don’t send)
  • Destructive actions behind confirm, with the generated content visible in the confirm
  • Rate and scope limits as product, not only as infra — what this agent is for
  • Memory that is inspectable. If it remembers, they should see what it thinks it knows
  • Clear failure. “I don’t know” and “I need X” are better UX than a fluent wrong paragraph

Test it like any other journey. Not “does the model impress.” Can a stranger complete the job, tell when to distrust the output, and recover when it is wrong? Behaviour, not a wow demo.

How I work still holds. If you are testing whether AI can perform the task, a working agent beats a clickable mock. Fidelity follows the question. Then you decide whether it deserves to be in the product — or whether it is a slide.

What I will not call AI-native

A sparkle icon on a text field. Autocomplete with no accountability. A chatbot taped to a settings page. A roadmap item named “GPT” with no job, no loop, and no owner when it hallucinates in front of a customer.

AI-native means the product changes because the model is there: the workflow, the trust contract, the states, the human who still decides. Product design for startups is that standard applied to the whole company. This page is the part that is easy to get theatrically wrong.

Speed is useful. Unexamined speed just ships confusion — and now it talks.

If the feature works in a demo and scares you in production

That gap is the UX.

Bring the flow as it is — the prompt, the output, the place a human is supposed to sit. Start a conversation. We will make the job obvious, the generation honest, and the loop real enough that the product can be used without you in the room.

The goal is not more AI. It is an AI feature someone can trust enough to finish the job — and doubt enough to stay safe.