Designing an AI product is not designing software with a chat box bolted on. It breaks the assumption every other interface quietly depends on: that the same action produces the same result. Once that assumption is gone, most of what a designer knows stops applying, and the interesting problems move somewhere unfamiliar.
This is a list of agencies and independents who understand that, plus what to ask before you hire any of them. I should say up front that I do AI product design and I am on this list. Read the rest knowing that, and judge each entry on whether the description is fair rather than on whether the author is disinterested. This is the AI specific cut. For the broader field, UX agencies ranked by research process covers shops across every category.
Why AI product design is a different job
In normal software, a designer can enumerate the states. Click here, this happens. Every path can be walked, drawn, and reviewed. That is why design systems work and why usability testing gives you clean answers.
An AI product does not behave that way. The model can be right, wrong, or plausibly wrong in a way the user cannot check. It can be slow in a way no loading spinner honestly represents. It can produce something excellent for one user and nonsense for the next, with no visible difference in what they did.
So the design questions change shape entirely:
- Trust. How does the user know whether to believe this output? What evidence does the interface hand them?
- Uncertainty. When the model is unsure, does the product say so, or does it project the same confidence either way?
- Correction. When the model is wrong, how quickly can the user fix it, and does that correction teach the system anything?
- Legibility. Can the user tell why the model said what it said, or is it a black box that they eventually stop trusting?
- The confidently wrong case. This is the one that kills products. Not failure, which people forgive, but fluent, well-formatted, completely wrong output that a busy user accepts.
None of that shows up in a portfolio screenshot. All of it shows up three months after launch, when people quietly stop using the feature.
Pavle Lucic
Best for: teams that want the person designing the AI product to also be the person building it.
I design AI products and then ship them, which means I have spent a lot of time on the specific problem of getting a person to trust a machine that is sometimes wrong. Drafting tools where the output has to sound like the user and not like a robot. Classifiers that are confidently mistaken often enough to matter. Autonomous agents where the design job is showing someone what is about to happen before it happens.
The pattern is always the same: the design work was never the chat box. It was the moment where the user has to decide whether the machine is right.
The honest tradeoff: I am one person. If you need a team of five working in parallel, that is not this. Most engagements start with a fixed price audit, which is a cheap way to find out whether the fit is real.
Clay
Best for: well funded AI companies that want a recognized name.
A San Francisco agency with strong AI and enterprise work. Polished, expensive, staffed like an agency. If you are raising or already raised, and you want a partner your board recognizes, they belong on the shortlist. Most funded AI products are SaaS underneath the model, so it is worth also checking the best SaaS design agencies for a shortlist built around pricing, onboarding, and packaging rather than the model layer.
Lazarev
Best for: complex AI products where nobody can articulate the user's job yet.
Research led, and unusually willing to spend time in discovery before drawing anything. If your AI product sits in a domain your team does not deeply understand, that gap is worth closing before it becomes a design.
Fuselab Creative
Best for: AI on top of dense data.
Strong at data visualization, which turns out to matter enormously for AI products that surface insight from large datasets. If your AI's job is to explain a number, the chart is the product.
The Gradient
Best for: AI native teams who want design partners who speak the same language.
Focused specifically on AI and machine learning products. Fluent in the vocabulary, which shortens the part of every engagement where the designer learns what a model actually does.
Punchcut
Best for: AI across devices and surfaces.
Long track record in emerging interfaces, including voice, ambient, and multi surface products. If your AI does not live in a browser tab, that experience is rare and valuable.
Neuron
Best for: AI in enterprise workflows.
Enterprise focused, with the patience that enterprise AI requires. AI inside a workflow that people are paid to follow is a very different problem from AI in a consumer app, and it needs someone who has sat through the compliance conversation.
The independents
Not every good AI designer works at an agency.
Emil Kowalski is a design engineer at Linear whose open source work is used across an enormous number of AI products, whether or not their teams know his name. Not for hire in the usual sense, but his writing is the closest thing the field has to a reference text on interaction quality.
Rauno Freiberg works at Vercel and writes about interaction detail at a level almost nobody else does. Same caveat: mostly not available, but the standard he sets is the standard worth measuring against.
The relevant point is not that you can hire these people. It is that the design engineers who set the bar for AI interfaces are, overwhelmingly, people who write code. That is not a coincidence. You cannot design for a model's uncertainty if you have never watched one behave.
The four failure modes worth naming
Before the questions, it helps to know what actually goes wrong, because the failures are predictable and they repeat across almost every AI product that quietly dies.
The demo trap. The feature is designed for the moment it is shown, not the moment it is used. In a demo, the operator picks the input, so the model always performs. In production, users pick the input, and they pick badly, and the model does something strange, and nobody designed for that. Any agency whose portfolio is a series of beautiful demos should be asked what the tenth real user did.
The fluent lie. The model is wrong, but it is wrong in complete sentences with confident formatting. This is the most dangerous state in an AI product, because a busy person accepts it. The interface has to make wrongness visible in a way the model itself cannot, which means designing for a failure the system does not know it is having.
The trust cliff. Users extend trust generously at first and withdraw it all at once. One bad output that costs someone real time, and they stop believing any output, including the good ones. Most AI features do not die from being bad. They die from one bad day that nobody designed a recovery from.
Correction that goes nowhere. The user notices the model is wrong, fixes it, and the fix vanishes into the void. Nothing is learned, nothing is remembered, and next week the same mistake arrives again. Users notice this faster than product teams expect, and they conclude, correctly, that their effort is not worth anything. Getting stop, steer, and undo controls right is what actually closes that loop.
Each of these is a design problem, not a model problem. A better model does not fix any of them.
What to ask, before you sign anything
Ask to see an error state.
Not the demo. Not the happy path where the model works and everyone is delighted. Ask what the product does when the model is confidently, fluently wrong.
An agency that has shipped AI will have an immediate answer, probably with a story attached, probably about a specific launch that went badly before it went well. An agency that has only styled AI will show you a nicer version of the demo.
That one question sorts the field faster than any portfolio review.
Then ask what happened afterwards. Not what they designed, but what users actually did with it. AI features have a specific failure mode: strong launch, good demo, quiet abandonment about six weeks in when people realize they cannot tell when to trust it. Anyone who has been through that will tell you about it, because it is the most instructive thing that ever happened to them.
If you are not sure you need an agency yet
Plenty of teams reach for an agency when what they actually have is a narrower problem: an AI feature that people tried once and never came back to, or an output nobody trusts, or a flow where the model is fine but the interface around it is not.
Often it is not an agency you need but a clearer idea of how AI interfaces should behave in the first place. The patterns for designing AI products, from uncertainty and hallucination to when not to use a chat box, cover most of what goes wrong before anyone gets hired.
Those are diagnosable in about a week. That is what a UX audit does: you get a prioritized report on where the product loses people, and a backlog your team can build from. If it turns out the answer is bigger than an audit, you will at least know what you are buying before you buy it.
If you want that run on your AI product, book a call and we can scope it in twenty minutes.