Designing for a model
that is sometimes wrong
Role
Design lead
First designer hired
Timeline
2024–2025
Zero to shipped
Type
iOS sourcing app
CNN + NLP search
Focus
Trust in AI output
And the system under it

Zero to one
They could design it.
They couldn't describe it.
The founders had a thesis — AI could match a founder to a factory from a reference image. They had no research and no shipped product.
I was the first designer hired, and led design across the product. So the first job was finding out whether the thesis was the real problem.
No shared vocabulary
A founder can design something and still not know how to name it in the terms a factory quotes against.
No way to spot a ghost
Listings look alike. Nothing on the page separates a real supplier from one that will take a deposit and vanish.
The tools assume bulk
Existing sourcing platforms are built for buyers ordering thousands of units, not founders ordering two hundred.
TL;DR
The gap
Independent founders can design a product and still not know what to call it to a factory, or how to tell a real supplier from a ghost.
The reframe
Interviewed 12+ founders, reframed the product from search to translation, and designed results that show their reasoning instead of asserting a match.
What shipped
A shipped 7-screen iOS app, a design system that cut sprint time about 40%, and a launch that went from zero to 350+ in three months.
12 founders
12+ founders,
one word that kept coming up.
I ran the interviews the product had never had, with the people it was supposedly for.
The reframe
It was never a search problem.
It was a translation problem.
Founders were not failing to find factories. They were failing to say what they wanted in a language a factory could quote against. That changed what the model was for, and what the interface had to do with its output.
Three rules
A match is a claim.
It has to show its evidence.
Three rules the interface had to follow, because the thing behind it is probabilistic and will sometimes be wrong.
Never assert, always attribute
Any score appears with the attributes it came from. A bare number is something a founder has to take on faith.
Make uncertainty visible
A weak match looks weak. Nothing is rounded up into confidence the model does not have.
Always leave a manual path
When the match is poor the screen hands over filters, so a dead end becomes a starting point.
The product
Two ways in,
and a result that explains itself.
Two ways in
Photo, or plain words
Founders don't know the vocabulary
An image search reads the reference you already have; a plain-language one takes “breathable linen, small runs” and does the translating. Neither asks you to know what a factory calls it.

Confidence
A score that shows its work
Every match carries its reasons
Fabric, minimum order, turnaround — the things the match was made on, attached to the match. A number on its own is a claim; a number with its inputs is something a founder can act on.

Failure
When the model is unsure
The screen has to say so
A low-confidence result says what it matched on, what it could not, and hands over the filters to narrow it by hand. Designing the failure state is most of designing for a model.

System & brand
The system was how
a team of this size shipped.
Tokens, components and dev-mode handoff, so engineers could build without waiting on a screen for every state.
It is the single reason sprint time came down about 40%, and the reason the product still looks like one product across seven screens.

What shipped
Zero to shipped,
with the reasoning attached.
My learning
What the reframe unlocked
The reframe. Once the problem was translation rather than search, every screen had an obvious job: turn what a founder has into what a factory understands.
What I never tested
Test the confidence display with founders before shipping it. I designed the honesty I believed in without checking whether a low score reads as useful or as broken.
The number that would settle it
How often a low-confidence result leads to a manual filter rather than an exit, and whether showing reasoning changes who founders actually contact.
AI has to explain itself or it doesn't earn trust. The model could rank suppliers from a photograph. That was never the hard part. The hard part was designing an interface that lets someone decide whether to believe it.