Scoring agent traces in Langfuse
August 25, 2026 · Judgeworthy
Traces tell you what your agent did. Scores turn that into a quality signal you can chart, threshold and argue from. Here is how to add them with the Python SDK.
Read moreAugust 25, 2026 · Judgeworthy
Traces tell you what your agent did. Scores turn that into a quality signal you can chart, threshold and argue from. Here is how to add them with the Python SDK.
Read moreAugust 21, 2026 · Judgeworthy
Most agent failures are visible in the tool calls, which are structured data. Here is how to assert on them exactly with Promptfoo and gate a merge on the result.
Read moreAugust 13, 2026 · Judgeworthy
Most agent teams can see every step their agent takes and still cannot say whether it took the right ones.
Read moreAugust 1, 2026 · Judgeworthy
Model-graded evals are unavoidable for agents, and most of them are wrong in ways nobody has checked.
Read more