Rosenverse
Hands-on AI #2: Understanding evals: LLM as a Judge

Log in or create a free Rosenverse account to watch this video.

Log in Create free account

100s of community videos are available to free members. Conference talks are generally available to Gold members.

Hands-on AI #2: Understanding evals: LLM as a Judge

Wednesday, October 15, 2025 • Rosenfeld Community

This video is featured in the Evals + Claude playlist.

Share the love for this talk
Hands-on AI #2: Understanding evals: LLM as a Judge
Speakers: Peter Van Dijck
Link:

Summary

If you’re a product manager, UX researcher, or any kind of designer involved in creating an AI product or feature, you need to understand evals. And a great way to learn is with a hands-on example. In this second talk in the series, Peter Van Dijck of the helpful intelligence company will show you how to create an eval for an AI product using an LLM as a judge (when we use a Large Language Model to evaluate the output of another Large Language Model). We’ll have a look at how that works, but also dig into why this even works. Are we creating problems for ourselves when we let an LLM judge itself? This talk is hands on; and there will be plenty of time for questions. You will go away understanding when and how to use LLM as a judge, and build some product sense around how the best AI products today are built, and how that can help you use them more effectively yourself.

Key Insights

  • Evals are a foundational feedback loop defining what 'good' means for AI products, helping to measure and improve systems continuously.

  • Evaluating fuzzy, subjective AI outputs requires innovative approaches such as using LLMs as judges to score results.

  • Binary (yes/no) scoring is more reliable than rating scales with ranges because LLMs lack internal memory and consistency.

  • Starting evals early (week one of a project) drastically improves AI product outcomes, but many teams delay due to perceived complexity.

  • High-risk or important tasks should be prioritized for evals instead of attempting broad coverage.

  • Assigning a dedicated owner or 'benevolent dictator' for evals who works closely with domain experts accelerates feedback and quality.

  • Creating a written constitution of principles helps concretize AI behavior goals and guides prompt and model training.

  • Most current eval tooling is too technical, slowing iteration cycles and making expert involvement inefficient.

  • Custom feedback interfaces tailored to expert users significantly speed up evaluating AI outputs in domains like healthcare and law.

  • Diverse perspectives from UX, product, strategy, and domain experts are critical in defining and refining what 'good' means in AI systems.

Notable Quotes

"Evals are everywhere, right? Everybody's talking about evals. It is like one of the key things in developing useful AI products."

"You want to ask an LLM to evaluate the fuzzy stuff because there’s no black and white output."

"LLMs don’t have memory, so rating on a scale from one to five is pretty random. Better to have yes or no answers."

"One of the biggest problems in AI building is evolving your prompts and having a fast feedback loop."

"By starting to categorize risk in detail, you naturally lead to better prompts and better evals."

"A constitution is a very good exercise: write down your system’s principles and values to help guide its behavior."

"Use custom systems for experts to quickly review and rate outputs, making feedback cycles much faster."

"Evals define a shared definition of good with tests to measure it, and that is the secret sauce for building great AI products."

"Model companies are students in a classroom wanting good points—they’re happy to run external expert evals to improve."

"The more I work with evals, the more I think UX and product people need to be involved because of the need for diverse perspectives."

Ask the Rosenbot
Jorge Arango
Scale Smart: AI-Powered Content Organization Strategies
2024 • DesignOps Summit 2024
Gold
Lori Muszynski
Keeping Design Weird
2023 • DesignOps Summit 2023
Gold
Cornelius Rachieru
Handling Complexity: Framing a Scale of Design
2021 • Design at Scale 2021
Gold
Bria Alexander
Theme Two Intro
2023 • DesignOps Summit 2023
Gold
Caroline Jarrett
Garbage in, garbage out? Measuring error rates to get ready for AI
2026 • Rosenfeld Community
Erika Flowers
The Handoff is Dead: Design-Led Engineering with AI Agents
2026 • Rosenfeld Community
Dan Willis
Enterprise Storytelling Sessions
2018 • Enterprise Experience 2018
Gold
Jose Coronado
From Zero to Hero
2022 • DesignOps Summit 2022
Gold
Lija Hogan
Practical Principles of Inclusive Research
2023 • Advancing Research 2023
Gold
Sarah Gallimore
Inspire Progress with Artifacts from the Future
2022 • Civic Design 2022
Gold
Dawn Ressel
Full-Stack User Experiences: A Marriage of Design and Technology
2016 • Enterprise UX 2016
Gold
Feyikemi Akinwolemiwa
Play to innovate: How curiosity and experimentation transform UX
2026 • Advancing Research 2026
Gold
Dave Hora
Research in the Face of Complexity: New Sensibility for New Situations
2025 • Rosenfeld Community
Peter Merholz
The 2025 State of UX/Design Organizational Health
2025 • Rosenfeld Community
Husani Oakley
Bias Towards Action: Building Teams that Build Work
2018 • Enterprise Experience 2018
Gold
Sarah Williams
A Framework for CX Transformation
2021 • Design at Scale 2021
Gold

More Videos

Alexandra Schmidt

"Designers need better training to work with off-the-shelf enterprise software like Sitecore, Salesforce, and SharePoint."

Alexandra Schmidt

Enterprise UX Playbook

December 1, 2022

Jen Crim

"Our offices have stand-up desks, nice collaboration areas, and comfy seating with a fresh, on-brand look."

Jen Crim Jess Quittner Saritha Kattekola Alex Karr Gurbani Pahwa

Culture, DIBS & Recruiting

June 11, 2021

Christian Crumlish

"Product managers are responsible for value – making sure something valuable is being created that meets real needs."

Christian Crumlish

AMA with Christian Crumlish, author of Product Management for UX People

March 24, 2022

Libby Maurer

"Having an employee resource group member on interview panels creates a safe space for candidates to disclose more about themselves."

Libby Maurer

Treating Diversity & Inclusion in Hiring as a Design Problem

December 5, 2019

Ali Jeffery

"Technology is a tool, not a solution in itself; it needs human input."

Ali Jeffery Sheri Chudow

How DesignOps Helped Enable Wall Street to Work Remotely

October 22, 2020

Daniel Gloyd

"Daniel Day-Lewis went so far as to suffer three broken ribs immersing in his role for My Left Foot."

Daniel Gloyd

Designing From the Inside Out: How Method Acting Can Inspire Design Research

February 12, 2026

Kristin Skinner

"Technology can revolutionize how we think about education."

Kristin Skinner

Five Years of DesignOps

September 29, 2021

Johanna Kollmann

"Confirmation bias is when we ignore or explain away data that doesn’t support what we already believe - Nina Belk."

Johanna Kollmann

Insights-Driven Product Strategy: Get your Research to Count

December 6, 2022

David Sternberg

"With QFI, we go upstream: simulate, model, predict user behavior before shipping, not just react after."

David Sternberg

Uncovering the hidden forces shaping user behavior

July 17, 2025