Free lessonsEvaluating LLMs and Agents
Evaluating LLMs and Agents
19 lessons
Lesson 3 · 24:18
Why accuracy is not enough
Experiment:Look at which class the errors fall into before trusting one score.
Open this lesson on Bilibili学掌门Bilibili seriesThe lesson is in Chinese and plays from Bilibili.
Not translated yet — showing the Chinese originals.
About the series
How AI evaluation differs from software testing, how to score and grade defects, then how to run batch evals with EvalScope and Langfuse.
- Who it is for
- Testers and engineers who need to accept a model or an agent, not just demo it.
- What you can do afterwards
- You can name the object, the metric, and the defect grade, then reproduce a batch run.
Training a whole team? It is arranged by scope. Book an evaluation