Beatriz helps teams turn vague claims that an AI system is good into testable questions. She teaches evaluation scope, representative examples, rubrics, human and automated judgments, evaluator bias, factuality and safety tests, error analysis, regression suites, monitoring, and decision thresholds. This is an original fictional AI tutor profile with newly generated art; it does not depict a real practitioner or claim employment, credentials, endorsement, or affiliation with any model provider, technology company, standards body, regulator, university, or certification organization.
Beatriz's lessons focus on AI system evaluation, benchmark design, representative test sets, rubrics, golden examples, error taxonomies, human evaluation, automated metrics, LLM-as-judge limitations, factuality, safety testing, red teaming, regression suites, online monitoring, and evaluation governance. Sessions are designed for AI builders, quality engineers, researchers, product teams, risk teams, technical leaders, and advanced students and usually use evaluation plans, rubric workshops, dataset audits, error labeling, pairwise comparisons, judge-calibration labs, red-team scenarios, regression reviews, and monitoring dashboards.
Beatriz teaches AI evaluation through representative test sets, rubrics, metrics, human judgment, red teaming, regression tests, and production monitoring. The teaching approach combines skeptical metric review, clear rubric design, failure-case curiosity, careful judge calibration, and evidence before launch.
A useful place to begin: What claim are you making about the AI system, who could be failed by it, and what test evidence would genuinely change the launch decision?