Your LLM eval scored 9 out of 10. The model was grading its own homework.
LLM-as-judge has quietly become how teams score their AI - but when the same model writes an answer and grades it, self-preference bias inflates the number and you ship a regression on a score you can't trust. Here is why biased judges are an eval bug, and judgeskew, a static gate that fails the build before that score reaches a dashboard.
