Rosenverse
Building impactful AI products for design and product leaders, Part 2: Evals are your moat

Log in or create a free Rosenverse account to watch this video.

Log in Create free account

100s of community videos are available to free members. Conference talks are generally available to Gold members.

Building impactful AI products for design and product leaders, Part 2: Evals are your moat

Wednesday, July 23, 2025 • Rosenfeld Community

This video is featured in the AI and UX playlist.

Share the love for this talk
Building impactful AI products for design and product leaders, Part 2: Evals are your moat
Speakers: Peter Van Dijck
Link:

Summary

The secret ingredient for impactful AI products is “evals”—an architecture for ongoing evaluation of quality. Without evals, you don’t know if your output is good. You don’t know when you’re done. Because outputs are non-deterministic, it’s very hard to figure out if you are creating real value for your users, and when something goes wrong, it’s really tricky to figure out why. Simply Put’s Peter van Dijck will demystify evals, and share a simple framework for planning for and building useful evals, from qualitative user research to automated evals using LLMs as a judge.

Key Insights

  • AI product development involves three layers: model capabilities, context management, and user experience, with evals central to experience quality assurance.

  • Automated evals help scale testing of AI with inherently open-ended inputs and outputs, enabling faster iteration cycles with confidence.

  • LLMs can serve as judges (evaluators) of other LLM outputs, which works because classification is cognitively easier than generation.

  • Defining what 'good' means for an AI system is a detailed, evolving process informed by research, domain expertise, and observed risks.

  • A three-option evaluation (e.g., yes/no/maybe) works better than fine-grained scales for consistent automated scoring by LLMs.

  • Synthetic data, generated by LLMs based on manually created examples, efficiently expands dataset breadth and usefulness.

  • Domain experts are essential for tagging data and establishing quality criteria, especially for high-stakes areas like healthcare or legal.

  • Building effective evals requires substantial effort—expect 20-40% of project resources devoted to this work.

  • Cultural differences impact subjective evals like politeness, requiring localization and careful domain definition.

  • AI product quality management is a strategic ongoing commitment, extending beyond initial development into production monitoring and iteration.

Notable Quotes

"AI products almost always have both open-ended inputs and outputs, which makes testing really hard."

"You have to build a detailed definition of what is good for my system to do meaningful automated evals."

"It’s much easier to classify an answer than to generate an answer, and that’s why LLM as a judge works."

"You don’t want to give too many options like rating from one to ten because consistency gets lost between different LLM calls."

"Synthetic data is useful because it’s easier to generate more examples of something you already have than to create entirely new data."

"If you launch in the US and politeness is an issue, first try to fix it with prompts; only if that fails should you build an eval."

"Evals are really your intellectual property—they define what good looks like in your domain."

"Domain experts are crucial for tagging data because users might say ‘that’s great,’ but experts can tell it’s totally wrong."

"You should plan 20 to 40 percent of your project budget on evals—it’s a lot more work than most people expect."

"This is where UX and product strategy bring huge value—defining what good means rather than leaving it to engineers alone."

Ask the Rosenbot
Discussion
2017 • Enterprise Experience 2017
Gold
Adrian Howard
Sturgeon’s Biases
2024 • DesignOps Summit 2024
Gold
Yulya Besplemennova
[Demo] Stress-testing GenAI in user research synthesis
2024 • Designing with AI 2024
Gold
Charlotte Lee
Theme 1 Intro
2021 • Civic Design 2021
Gold
Jorge Arango
[Demo] How to re-categorize content at scale using LLMs
2024 • Designing with AI 2024
Gold
Sara Conklin
A UXer’s 12-Month Journey from Climate Concern to Climate Credibility
2025 • Climate UX Interest Group
World Usability Day Panel Discussion
2022 • DesignOps Community
Dr Chloe Sharp
Using Evidence and Collaboration for Setting and Defending Priorities
2023 • Design in Product 2023
Gold
Jodi Forlizzi
Design and AI innovation
2024 • Designing with AI 2024
Gold
Alicia Mooty
Design Staffing Models
2021 • DesignOps Summit 2021
Gold
Joerg Beringer
Scaling User Research with AI: Continuous Discovery of User Needs in Minutes
2025 • Designing with AI 2025
Gold
Shipra Kayan
Make your research synthesis speedy and more collaborative using a canvas
2025 • Rosenfeld Community
Eduardo Ortiz
Day 3 Theme Panel
2025 • Advancing Research 2025
Gold
Darian Davis
Lessons from a Toxic Work Relationship
2024 • Enterprise Experience 2020
Gold
Séamus Byrne
Aligning Teams with Choreography
2024 • Enterprise Experience 2020
Gold
Shipra Kayan
How Tess Dixon Facilitates Team Engagement and Collaboration at Condé Nast Using Miro 
2021 • DesignOps Summit 2021
Gold

More Videos

Alexia Cohen

"We decided with the COP to view our budget as a moral document."

Alexia Cohen Adriane Ackerman

Increasing Health Equity and Improving the Service Experience for Under-Served Latine Communities in Arizona

December 4, 2024

Marc Majers

"The timing is key—you want to interrupt them when they are in that flow state."

Marc Majers Tony Turner

Interrupted UX - Add A Dose of Reality To Usability Testing

March 11, 2022

Jorge Arango

"Our appliances are attacking us somehow and bringing down major parts of our infrastructure."

Jorge Arango

Design as an Antidote to VUCA

May 9, 2019

Brenna Fallon

"Growth and learning is your long term change management plan—does it take letting go of a clear outcome? Yes, but it’s worth the leap."

Brenna Fallon

Learning Over Outcomes

October 24, 2019

Daniel Gloyd

"The Shakers’ principled approach to design was a precursor to Bauhaus’s form follows function and today’s user-centered values."

Daniel Gloyd

Warming the User Experience: Lessons from America's first and most radical human-centered designers

May 9, 2024

Sam Proulx

"The four Cs—consistency, convenience, confidence, and customizability—are not just good for accessibility, they make a great experience for everyone."

Sam Proulx

Online Shopping: Designing an Accessible Experience

October 3, 2023

Kristin Skinner

"Every student has unique strengths that should be recognized and nurtured."

Kristin Skinner

Five Years of DesignOps

September 29, 2021

Denise Jacobs

"Racism is by design, and there's no way to counter it unless we counter it with design."

Denise Jacobs Nancy Douyon Renee Reid Lisa Welchman

Interactive Keynote: Social Change by Design

January 8, 2024

Roberta Dombrowski

"You can send NDAs and informed consent forms directly through the platform, cutting down admin overhead."

Roberta Dombrowski Lianna Aduana

5 Reasons to Bring your Recruiting in House

September 30, 2021