CodexQA

Not translated yet — showing the Chinese originals.

About the series

How AI evaluation differs from software testing, how to score and grade defects, then how to run batch evals with EvalScope and Langfuse.

Who it is for
Testers and engineers who need to accept a model or an agent, not just demo it.
What you can do afterwards
You can name the object, the metric, and the defect grade, then reproduce a batch run.

Training a whole team? It is arranged by scope. Book an evaluation