August 31, 2026

Lead Product Designer @ UserEvidence
UX, AI Design, Product Strategy

AI design and automations


When I design an AI feature I ask myself: does it leave the user with less work, or just different work?

A piece of content that looks convincing but gives me no way to understand why AI recommended it, doesn’t save me time.

An unattributed customer testimonial that isn't specific enough, a search result with unverifiable sources, or five thousand recommended responses from a recent survey that need to be sifted through… I go from reviewing, to auditing, and add on time deciding whether I can trust whatever the LLM spit out.

The solutions we build won’t eliminate judgment, it changes where and how often human judgment needs to happen. The design challenge is figuring out which judgments should stay with the user, and which ones AI can help make. Eventually, we can determine which ones are predictable enough to automate.

Project Specifics

Auto-Publish: Reduce the review queue

The problem

A marketer creates a survey in our platform and sends it to their customers. They want to understand sentiment, general experience, and collect feedback on new features. Someone then has to read through every response, looking for the content worth publishing. Hundreds of responses, sometimes thousands, most of them are unusable, but all of them require a judgment call.


What we built

A skill that evaluates each incoming response against a set of publish criteria: a high satisfaction score, signs of real emotion, evidence of trust, and an attributable named source. Responses that clear the bar publish straight to the library or move into a review queue. Nobody has to spend countless hours reading the entire pile anymore.

What does worth publishing mean and translating judgment into rules

The best testimonials weren't necessarily the ones with the highest NPS. We defined a set of signals that together indicated something worth surfacing: strong satisfaction, positive emotion, enough substance to be useful, and most importantly, a source that could be attributed and verified.

Judgement criteria


A strong NPS score

It's not the most accurate signal, but it's familiar to our customers and they already relate to it. It's always attached to a person, so it instinctively signals legitimacy.

Signs of real emotion

Emotional content reads as human, and human reads as real. Prospects weigh this content more heavily than content full of buzzwords or tech vomit.

Evidence of trust or third-party verification

A survey date, number of participants, or any way to show the paper trail and full data set behind the content makes it easier for a user to validate what they're seeing.

A named, attributed source

An anonymous quote is a claim. An attributed one is evidence of a real human behind the quote.

Model specifics for all us nerds

Attribution and a documented source were the foundation of verification. From there, an NPS of 8+ with supporting signals represented the strongest candidates, while responses with an NPS as low as 5 could still surface when the other credibility signals were present.

The goal was to find the best responses and most likely to get published.

What we learned

We launched the first part of this work and customers were excited about the recommendations, but we missed the mark on a few things. A few rounds of interviews and usage data helped us understand how and why.

More specific criteria, and fewer results

Users wanted more specific recommendations in the queue, and fewer of them. They still felt like they were reviewing too much.

Ability to set their own criteria

We knew this was coming and built the infrastructure to support a custom model, allowing customers to create their own automations rather than relying solely on our presets. The more we learned, the clearer it became that "worth publishing" couldn't be one definition we owned.

Integrate more sources

Customers wanted us to pull in other sources as well — G2 reviews, Gartner data, and internal documents. The source problem was bigger than the survey itself. The underlying need was to identify credible, usable evidence wherever it lived.