CodexQA

Not translated yet — showing the Chinese originals.

About the series

A nine-part path through AI application evaluation: metrics, datasets, evaluators, Agent and RAG evaluation, and a practical LangSmith workflow.

Who it is for
Testers and engineers building AI applications, Agents, or RAG systems who need repeatable evaluations.
What you can do afterwards
Build an evaluation set, choose metrics and evaluators, and run a repeatable evaluation for an LLM, Agent, or RAG application with LangSmith.

Training a whole team? It is arranged by scope. Book an evaluation