DEV Community

#evaluation

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Production-Ready Multi-Turn Evaluation

Production-Ready Multi-Turn Evaluation

Comments
7 min read
A Free Server Is Enough to Test a New Model Before You Trust It

A Free Server Is Enough to Test a New Model Before You Trust It

Comments
3 min read
Writing the code is no longer the bottleneck

Writing the code is no longer the bottleneck

Comments
3 min read
Why AI Benchmarks Mean Less Than You Think

Why AI Benchmarks Mean Less Than You Think

Comments
6 min read
One Quality Score Is a Lie: Split Your RAG Judge Into Retrieval, Groundedness, and Relevance

One Quality Score Is a Lie: Split Your RAG Judge Into Retrieval, Groundedness, and Relevance

1
Comments 1
5 min read
We Almost Deployed a Temporal Knowledge Graph. The Eval Said No.

We Almost Deployed a Temporal Knowledge Graph. The Eval Said No.

Comments
7 min read
Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations

Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations

Comments
7 min read
Dividing your RAG score by retrieval recall overstates your generation quality, and here is by how much

Dividing your RAG score by retrieval recall overstates your generation quality, and here is by how much

2
Comments 2
9 min read
How EvalPort's Grader System Works: 11 Types for LLM Evaluation

How EvalPort's Grader System Works: 11 Types for LLM Evaluation

Comments
2 min read
Measure the Judge Before You Trust It: Self-Consistency Comes Before Human Agreement

Measure the Judge Before You Trust It: Self-Consistency Comes Before Human Agreement

1
Comments 2
6 min read
RAG Beyond the Demo: Pipeline, Citations, Evaluation, and When Not to Bother

RAG Beyond the Demo: Pipeline, Citations, Evaluation, and When Not to Bother

Comments 2
7 min read
OpenEval: Why LLM Evaluation Needs a Standard Format

OpenEval: Why LLM Evaluation Needs a Standard Format

Comments
1 min read
Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations

Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations

1
Comments 2
7 min read
What Coding Agents Say When They Talk to Each Other

What Coding Agents Say When They Talk to Each Other

Comments 4
13 min read
Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.

Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.

1
Comments
6 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.