Free lessonsEvaluating LLMs and Agents
Evaluating LLMs and Agents
19 lessons
Lesson 4 · 58:44
A standard evaluation workflow
Experiment:Data, metric, then a conclusion you can rerun.
Open this lesson on Bilibili学掌门Bilibili seriesThe lesson is in Chinese and plays from Bilibili.
Not translated yet — showing the Chinese originals.
About the series
How AI evaluation differs from software testing, how to score and grade defects, then how to run batch evals with EvalScope and Langfuse.
- Who it is for
- Testers and engineers who need to accept a model or an agent, not just demo it.
- What you can do afterwards
- You can name the object, the metric, and the defect grade, then reproduce a batch run.
Training a whole team? It is arranged by scope. Book an evaluation