Rosenverse
Hands-on AI #2: Understanding evals: LLM as a Judge

Log in or create a free Rosenverse account to watch this video.

Log in Create free account

100s of community videos are available to free members. Conference talks are generally available to Gold members.

Hands-on AI #2: Understanding evals: LLM as a Judge

Wednesday, October 15, 2025 • Rosenfeld Community

This video is featured in the Evals + Claude playlist.

Share the love for this talk
Hands-on AI #2: Understanding evals: LLM as a Judge
Speakers: Peter Van Dijck
Link:

Summary

If you’re a product manager, UX researcher, or any kind of designer involved in creating an AI product or feature, you need to understand evals. And a great way to learn is with a hands-on example. In this second talk in the series, Peter Van Dijck of the helpful intelligence company will show you how to create an eval for an AI product using an LLM as a judge (when we use a Large Language Model to evaluate the output of another Large Language Model). We’ll have a look at how that works, but also dig into why this even works. Are we creating problems for ourselves when we let an LLM judge itself? This talk is hands on; and there will be plenty of time for questions. You will go away understanding when and how to use LLM as a judge, and build some product sense around how the best AI products today are built, and how that can help you use them more effectively yourself.

Key Insights

  • Evals are a foundational feedback loop defining what 'good' means for AI products, helping to measure and improve systems continuously.

  • Evaluating fuzzy, subjective AI outputs requires innovative approaches such as using LLMs as judges to score results.

  • Binary (yes/no) scoring is more reliable than rating scales with ranges because LLMs lack internal memory and consistency.

  • Starting evals early (week one of a project) drastically improves AI product outcomes, but many teams delay due to perceived complexity.

  • High-risk or important tasks should be prioritized for evals instead of attempting broad coverage.

  • Assigning a dedicated owner or 'benevolent dictator' for evals who works closely with domain experts accelerates feedback and quality.

  • Creating a written constitution of principles helps concretize AI behavior goals and guides prompt and model training.

  • Most current eval tooling is too technical, slowing iteration cycles and making expert involvement inefficient.

  • Custom feedback interfaces tailored to expert users significantly speed up evaluating AI outputs in domains like healthcare and law.

  • Diverse perspectives from UX, product, strategy, and domain experts are critical in defining and refining what 'good' means in AI systems.

Notable Quotes

"Evals are everywhere, right? Everybody's talking about evals. It is like one of the key things in developing useful AI products."

"You want to ask an LLM to evaluate the fuzzy stuff because there’s no black and white output."

"LLMs don’t have memory, so rating on a scale from one to five is pretty random. Better to have yes or no answers."

"One of the biggest problems in AI building is evolving your prompts and having a fast feedback loop."

"By starting to categorize risk in detail, you naturally lead to better prompts and better evals."

"A constitution is a very good exercise: write down your system’s principles and values to help guide its behavior."

"Use custom systems for experts to quickly review and rate outputs, making feedback cycles much faster."

"Evals define a shared definition of good with tests to measure it, and that is the secret sauce for building great AI products."

"Model companies are students in a classroom wanting good points—they’re happy to run external expert evals to improve."

"The more I work with evals, the more I think UX and product people need to be involved because of the need for diverse perspectives."

Ask the Rosenbot
Ellie Krysl
Planned Right. Managed Right. Designed Right.
2023 • Enterprise UX 2023
Gold
Kate Koch
Flex Your Super Powers: When a Design Ops Team Scales to Power CX
2021 • DesignOps Summit 2021
Gold
Louis Rosenfeld
Founder’s Welcome
2022 • Design in Product 2022
Gold
Marjorie Stainback
Transforming Strategic Research Capacity through Democratization
2019 • DesignOps Summit 2019
Gold
Angelos Arnis
State of DesignOps: Learnings from the 2021 Global Report
2021 • DesignOps Summit 2021
Gold
Jim Kalbach
Jazz Improvisation as a Model for Team Collaboration
2019 • Enterprise Experience 2019
Gold
Sheryl Cababa
Integrating Systems Thinking Into Your Practice as a Designer
2025 • Rosenfeld Community
Laura Weiss
Turn Down the Heat: 3 Ways to Handle Conflict in the Moment
2024 • Rosenfeld Community
Mary-Lynne Williams
Exit Interview #4: From Product Design Leadership to Sound Healing
2026 • Rosenfeld Community
Louis Rosenfeld
GenAI for UXers: A Rosenbot Demo and Discussion
2025 • DesignOps Summit 2025
Gold
Rachael Dietkus, LCSW
AI: Passionate defenses and reasoned critique [Advancing Research Community Workshop Series]
2024 • Advancing Research Community
Frances Yllana
DesignOps Exposed: What do our peers really think of us?
2025 • DesignOps Summit 2025
Gold
Catherine Blizzard
Using Integrated Insight to Drive Growth
2022 • Advancing Research 2022
Gold
Matt Bernius
Learnings from Applying Trauma-Informed Principles to the Research Process
2022 • Advancing Research 2022
Gold
Charlotte Vorbeck
Pipeline to Civic Design
2021 • Civic Design 2021
Gold
Marisa Bernstein
It Takes GRIT: Lessons from the Small, but Mighty World of Civic Usability Testing
2021 • Civic Design 2021
Gold

More Videos

Nathan Shedroff

"Reading books and articles is no longer central for many young people. What will replace books and articles for collective knowledge building?"

Nathan Shedroff Hugh Dubberly Thomas J. McLeish

How Will Design be Taught When the Schools Shut Down?

May 8, 2026

John Paul de Guzman

"Spending more time doesn’t automatically make you productive, it just means you spent more time doing things."

John Paul de Guzman

10k Screens Later: How We Became a Data-Driven Design Organization

September 24, 2024

Aleksandra Korczynska

"The key to combating survey fatigue is short surveys triggered at the right context, making respondents feel listened to and valued."

Aleksandra Korczynska Caroline Jarrett Justyna Parmee

Survey Tools

March 12, 2026

Kristin Skinner

"We need to define success in ways that make our work meaningful and purposeful."

Kristin Skinner

Theme 2: Introduction and Provocation

January 8, 2024

Rachael Dietkus, LCSW

"The conference is a single track, so everyone listens to the same content together, fostering community focus."

Rachael Dietkus, LCSW Victor Udoewa Jennifer Strickland

Everything You Need to Know about the Civic Design 2022 Call for Presentations

May 17, 2022

Uday Gajendar

"Meta design is about designing the conditions for design itself to happen well."

Uday Gajendar

The Rise of Meta-Design: A Starter Playbook

May 19, 2022

Vanessa Varin

"You can't do one without the other. Design the system, set rituals in the quality bar. They reinforce each other."

Vanessa Varin

Feedback: The Other F-Word

September 10, 2025

Hugh Dubberly

"We are moving from direct work to mediated work, from wanting things perfect to good enough for now, and from complete to adaptive and growing systems."

Hugh Dubberly

Problems with Problems: Reconsidering the Frame of Designing as Problem-Solving

June 19, 2019

Sean McKay

"Engineers started questioning old assumptions and product wasn’t just adjusting roadmaps, they were reframing decisions around user needs."

Sean McKay

Coexisting with non-researchers: Practical strategies for a democratized research future

March 11, 2025