Evals
Labelbox · Design & engineering
Customer evaluation product
Overview
Evaluation Studio is a customer-gated Labelbox product for AI teams investigating model performance — compare runs, chase regressions, and drill into failure modes instead of staring at opaque scorecards. I owned design and engineering, staying close to a Staff data scientist/researcher so the product matched how customers actually validate data and decide what to fix next.
Context
Scorecards alone don’t tell AI teams what to fix. Customers needed a product for comparing runs, chasing regressions, and validating examples with researchers in the loop.
How it evolved
Owned design and engineering beside a Staff data scientist/researcher so customer needs and data validation shaped the investigation UX — not afterthought polish on a prototype.
Where it is now
Core Labelbox customer product. Product mechanics and demos live in the Evals case study.
What it delivered
Design and engineering ownership on a customer product — not a handoff between mockups and implementation
Partnered with Staff data scientist/researcher so customer needs and data validation shaped the investigation UX
Grew from a customer POC into a core Labelbox service — full product story in the Evals case study
How it ships
Chose React + GraphQL + Plotly so customers could move from aggregate metrics into example-level detail without leaving the product
Tied UI contracts to the shared design system and auth gateway so Evals could ship as a real multi-tenant customer product, not a one-off demo
Launch narrative: product blog
Surfaces
Evaluation Studio · /case-studies/evals
Stack
Team
Design & engineering ownership; close partnership with Staff data scientist/researcher