Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
evaluation
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Production-Ready Multi-Turn Evaluation
Humza Tareen
Humza Tareen
Humza Tareen
Follow
Aug 25
Production-Ready Multi-Turn Evaluation
#
multiturn
#
evaluation
#
python
#
docker
Comments
Add Comment
7 min read
A Free Server Is Enough to Test a New Model Before You Trust It
Quinn Li
Quinn Li
Quinn Li
Follow
Aug 17
A Free Server Is Enough to Test a New Model Before You Trust It
#
ai
#
python
#
evaluation
#
agents
Comments
Add Comment
3 min read
Writing the code is no longer the bottleneck
The Engineering Manager’s Desk
The Engineering Manager’s Desk
The Engineering Manager’s Desk
Follow
Aug 15
Writing the code is no longer the bottleneck
#
engineeringculture
#
evaluation
#
ai
#
techdebt
Comments
Add Comment
3 min read
Why AI Benchmarks Mean Less Than You Think
The AI Downside
The AI Downside
The AI Downside
Follow
Aug 15
Why AI Benchmarks Mean Less Than You Think
#
benchmarks
#
llms
#
evaluation
#
hype
Comments
Add Comment
6 min read
One Quality Score Is a Lie: Split Your RAG Judge Into Retrieval, Groundedness, and Relevance
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Aug 19
One Quality Score Is a Lie: Split Your RAG Judge Into Retrieval, Groundedness, and Relevance
#
ai
#
evaluation
#
dotnet
#
testing
1
 reaction
Comments
1
 comment
5 min read
We Almost Deployed a Temporal Knowledge Graph. The Eval Said No.
Guatu
Guatu
Guatu
Follow
Aug 14
We Almost Deployed a Temporal Knowledge Graph. The Eval Said No.
#
aiagents
#
knowledgegraph
#
evaluation
#
rag
Comments
Add Comment
7 min read
Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations
Sanjeev Kumar
Sanjeev Kumar
Sanjeev Kumar
Follow
Aug 12
Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations
#
ai
#
llm
#
evaluation
Comments
Add Comment
7 min read
Dividing your RAG score by retrieval recall overstates your generation quality, and here is by how much
Maya Andersson
Maya Andersson
Maya Andersson
Follow
Aug 25
Dividing your RAG score by retrieval recall overstates your generation quality, and here is by how much
#
rag
#
llm
#
evaluation
#
machinelearning
2
 reactions
Comments
2
 comments
9 min read
How EvalPort's Grader System Works: 11 Types for LLM Evaluation
Adha AK
Adha AK
Adha AK
Follow
Aug 4
How EvalPort's Grader System Works: 11 Types for LLM Evaluation
#
llm
#
evaluation
#
testing
#
opensource
Comments
Add Comment
2 min read
Measure the Judge Before You Trust It: Self-Consistency Comes Before Human Agreement
Saurav Bhattacharya
Saurav Bhattacharya
Saurav Bhattacharya
Follow
Aug 12
Measure the Judge Before You Trust It: Self-Consistency Comes Before Human Agreement
#
ai
#
evaluation
#
dotnet
#
testing
1
 reaction
Comments
2
 comments
6 min read
RAG Beyond the Demo: Pipeline, Citations, Evaluation, and When Not to Bother
Xinyang Wu
Xinyang Wu
Xinyang Wu
Follow
Aug 3
RAG Beyond the Demo: Pipeline, Citations, Evaluation, and When Not to Bother
#
rag
#
llm
#
embeddings
#
evaluation
Comments
2
 comments
7 min read
OpenEval: Why LLM Evaluation Needs a Standard Format
Adha AK
Adha AK
Adha AK
Follow
Jul 30
OpenEval: Why LLM Evaluation Needs a Standard Format
#
llm
#
evaluation
#
ai
#
testing
Comments
Add Comment
1 min read
Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations
Sanjeev Kumar
Sanjeev Kumar
Sanjeev Kumar
Follow
Aug 12
Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations
#
ai
#
llm
#
evaluation
1
 reaction
Comments
2
 comments
7 min read
What Coding Agents Say When They Talk to Each Other
JaviMaligno
JaviMaligno
JaviMaligno
Follow
Aug 22
What Coding Agents Say When They Talk to Each Other
#
ai
#
agents
#
evaluation
Comments
4
 comments
13 min read
Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.
Maya Andersson
Maya Andersson
Maya Andersson
Follow
Jul 21
Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.
#
statistics
#
machinelearning
#
datascience
#
evaluation
1
 reaction
Comments
Add Comment
6 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account