Free lessonsEvaluating LLMs and Agents
Evaluating LLMs and Agents
19 lessons
Lesson 17 · 10:56
Langfuse 的 Trace 和 Evals
Experiment:简介写了从 Observation、Trace、Session 讲到 Evals、Datasets 和 Metrics。
Open this lesson on Bilibili前端屎不了Bilibili seriesThe lesson is in Chinese and plays from Bilibili.
Not translated yet — showing the Chinese originals.
About the series
How AI evaluation differs from software testing, how to score and grade defects, then how to run batch evals with EvalScope and Langfuse.
- Who it is for
- Testers and engineers who need to accept a model or an agent, not just demo it.
- What you can do afterwards
- You can name the object, the metric, and the defect grade, then reproduce a batch run.
Training a whole team? It is arranged by scope. Book an evaluation