LLM Judges: Evaluating Agreement as a Reliability Indicator
Amazon Science’s recent blog post, “When LLM Judges Agree, Should We Believe Them?” explores the growing use of large language models (LLMs) as automated decision‑makers in legal and regulatory contexts. The article outlines how LLMs can be prompted to evaluate case facts, apply precedent, and produce verdicts that mirror human judgments. It highlights experiments in which multiple LLMs were asked to judge the same set of legal scenarios, noting that high agreement rates often correlate with higher accuracy, yet also pointing out systematic biases that can arise when training data is unrepresentative or when prompts are ambiguous.
The post references a discussion thread on Hacker News (Y Combinator) where 32 upvotes and 13 comments reflect a mixed reception. Contributors praised the technical feasibility and potential for scaling judicial assistance, while others raised concerns about transparency, accountability, and the risk of embedding existing legal inequities into algorithmic outputs. The article concludes by calling for rigorous external audits, clear interpretability standards, and a cautious approach to deploying LLMs in high‑stakes decision‑making until these safeguards are in place.