Log in or create a free Rosenverse account to watch this video.
Log in Create free account100s of community videos are available to free members. Conference talks are generally available to Gold members.
Hands-on AI #1: Let’s write your first AI eval
This video is featured in the Evals + Claude playlist.
Summary
If you’re a product manager, UX researcher, or any kind of designer involved in creating an AI product or feature, you need to understand evals. And a great way to learn is with a hands-on example. In this talk, Peter Van Dijck of the helpful intelligence company will walk you through writing your first eval. You will learn the basic concepts and the tools, and write an eval together. This talk is hands on; you can follow along, and there will be plenty of time for questions. You will go away with an understanding of the basic building blocks of AI evals, and with the confidence that you know how to write one. And more importantly, you’ll build some intuition, some product sense, around how the best AI products today are built, and how that can help you use them more effectively yourself.
Key Insights
-
•
Evals consist of a task, a golden dataset with known correct outputs, and an evaluator that measures correctness.
-
•
Manual AI prompt testing is slow and inconsistent; automated evals accelerate and scale evaluation.
-
•
UX and product teams can and should learn evals as a practical, non-technical skill.
-
•
Creating your own golden dataset is essential and cannot be outsourced or fully automated.
-
•
Models are fixed once trained; improvements happen by refining prompts and context design, not retraining the model.
-
•
Evaluations measure task performance, not the underlying model itself, allowing comparison across models.
-
•
Outputting a confidence score from models is unreliable due to lack of internal memory and inconsistent scale interpretation.
-
•
Biases are baked into models during training via evals used in post-training refinement.
-
•
LLMs can be used to judge other LLM outputs to evaluate tasks with non-binary answers.
-
•
Effective eval work requires collaboration across data analysts, engineers, subject matter experts, and UX/product teams.
Notable Quotes
"Evals are like a way to define what good looks like."
"The model was baked and once it’s baked, it does not learn again until they bake a new one."
"You need to be looking at the data. Nobody wants to, but that’s core work."
"Without a golden dataset, you have to build the golden dataset yourself."
"We’re not teaching the model anything; we’re improving our prompts and context."
"Confidence scores from the model are not a good idea because the model has no memory."
"Biases are baked in through the evals used during model training and post-training."
"LLMs judging other LLMs might sound crazy, but if you do it right, it works."
"Evals are a product and UX skill; learning them lets you make these systems do what you want."
"There is a large and growing capability overhang in these models we haven’t discovered yet."
Or choose a question:
More Videos
"Increasing page load time by three seconds can nullify all your design improvements."
Erin WeigelUX Lessons from running more than 1,200 A/B Tests
July 10, 2024
"To innovate smarter, you need to get access to the roadmap as early as possible and start research even when not asked for it."
Mike OrenWhy Pharmaceutical's Research Model Should Replace Design Thinking
March 28, 2023
"Understanding how assistive technology evolves is crucial because it foretells the future of interface design."
Lija Hogan Milan Mijatovic Sam Proulx Louis RosenfeldThree Years Out: Perspectives on the Near-Term Future of User Research
March 15, 2024
"Giving users the ability to reverse AI-driven actions is critical but currently underexplored."
Heidi TrostWhen AI Becomes the User’s Point Person—and Point of Failure
August 7, 2025
"Behavior over time is your culture; it’s how you choose to behave and incentivize others to behave."
Adam CutlerPeople + Places + Practices = Outcomes
June 8, 2016
"It’s good practice to offer one-on-one interactions with older adults due to their learning preferences."
Rittika BasuAge and Interfaces: Equipping Older Adults with Technological Tools
February 23, 2023
"Accessibility ownership should never fall on just one person but be understood and shared across entire product teams."
Saara Kamppari-MillerDesignOps for Inclusive Design and Accessibility
May 26, 2022
"We want you to leave feeling like all those missed Slack messages and emails are worth it."
Bria AlexanderTheme Two Intro
October 3, 2023
"Design Ops teams exist in nearly all industries and for all design functions, growing rapidly year over year."
Laine Riley Prokay Lisa GordonCarving a Path for Early Career DesignOps Practitioners
September 9, 2022
Latest Books All books
Dig deeper with the Rosenbot
How can teams create lightweight visibility systems that improve workload transparency without reducing productivity?
What are examples of play used for executive-level prioritization and trade-offs?
Why is assessing design students challenging, and what are the limitations of traditional grading systems?