Log in or create a free Rosenverse account to watch this video.
Log in Create free account100s of community videos are available to free members. Conference talks are generally available to Gold members.
Building impactful AI products for design and product leaders, Part 2: Evals are your moat
This video is featured in the AI and UX playlist.
Summary
The secret ingredient for impactful AI products is “evals”—an architecture for ongoing evaluation of quality. Without evals, you don’t know if your output is good. You don’t know when you’re done. Because outputs are non-deterministic, it’s very hard to figure out if you are creating real value for your users, and when something goes wrong, it’s really tricky to figure out why. Simply Put’s Peter van Dijck will demystify evals, and share a simple framework for planning for and building useful evals, from qualitative user research to automated evals using LLMs as a judge.
Key Insights
-
•
AI product development involves three layers: model capabilities, context management, and user experience, with evals central to experience quality assurance.
-
•
Automated evals help scale testing of AI with inherently open-ended inputs and outputs, enabling faster iteration cycles with confidence.
-
•
LLMs can serve as judges (evaluators) of other LLM outputs, which works because classification is cognitively easier than generation.
-
•
Defining what 'good' means for an AI system is a detailed, evolving process informed by research, domain expertise, and observed risks.
-
•
A three-option evaluation (e.g., yes/no/maybe) works better than fine-grained scales for consistent automated scoring by LLMs.
-
•
Synthetic data, generated by LLMs based on manually created examples, efficiently expands dataset breadth and usefulness.
-
•
Domain experts are essential for tagging data and establishing quality criteria, especially for high-stakes areas like healthcare or legal.
-
•
Building effective evals requires substantial effort—expect 20-40% of project resources devoted to this work.
-
•
Cultural differences impact subjective evals like politeness, requiring localization and careful domain definition.
-
•
AI product quality management is a strategic ongoing commitment, extending beyond initial development into production monitoring and iteration.
Notable Quotes
"AI products almost always have both open-ended inputs and outputs, which makes testing really hard."
"You have to build a detailed definition of what is good for my system to do meaningful automated evals."
"It’s much easier to classify an answer than to generate an answer, and that’s why LLM as a judge works."
"You don’t want to give too many options like rating from one to ten because consistency gets lost between different LLM calls."
"Synthetic data is useful because it’s easier to generate more examples of something you already have than to create entirely new data."
"If you launch in the US and politeness is an issue, first try to fix it with prompts; only if that fails should you build an eval."
"Evals are really your intellectual property—they define what good looks like in your domain."
"Domain experts are crucial for tagging data because users might say ‘that’s great,’ but experts can tell it’s totally wrong."
"You should plan 20 to 40 percent of your project budget on evals—it’s a lot more work than most people expect."
"This is where UX and product strategy bring huge value—defining what good means rather than leaving it to engineers alone."
Or choose a question:
More Videos
"We decided with the COP to view our budget as a moral document."
Alexia Cohen Adriane AckermanIncreasing Health Equity and Improving the Service Experience for Under-Served Latine Communities in Arizona
December 4, 2024
"The timing is key—you want to interrupt them when they are in that flow state."
Marc Majers Tony TurnerInterrupted UX - Add A Dose of Reality To Usability Testing
March 11, 2022
"Our appliances are attacking us somehow and bringing down major parts of our infrastructure."
Jorge ArangoDesign as an Antidote to VUCA
May 9, 2019
"Growth and learning is your long term change management plan—does it take letting go of a clear outcome? Yes, but it’s worth the leap."
Brenna FallonLearning Over Outcomes
October 24, 2019
"The Shakers’ principled approach to design was a precursor to Bauhaus’s form follows function and today’s user-centered values."
Daniel GloydWarming the User Experience: Lessons from America's first and most radical human-centered designers
May 9, 2024
"The four Cs—consistency, convenience, confidence, and customizability—are not just good for accessibility, they make a great experience for everyone."
Sam ProulxOnline Shopping: Designing an Accessible Experience
October 3, 2023
"Every student has unique strengths that should be recognized and nurtured."
Kristin SkinnerFive Years of DesignOps
September 29, 2021
"Racism is by design, and there's no way to counter it unless we counter it with design."
Denise Jacobs Nancy Douyon Renee Reid Lisa WelchmanInteractive Keynote: Social Change by Design
January 8, 2024
"You can send NDAs and informed consent forms directly through the platform, cutting down admin overhead."
Roberta Dombrowski Lianna Aduana5 Reasons to Bring your Recruiting in House
September 30, 2021
Latest Books All books
Dig deeper with the Rosenbot
How does Rally’s research infrastructure facilitate participant management and integration with enterprise data systems?
Why might a paradigm shift be necessary to sustainably rebuild rural maternal health systems?
What human oversight processes are necessary to ensure the accuracy and relevance of AI-enhanced research repositories?