{"version":1,"type":"story","url":"https://digestai.news/story/megagon-labs-releases-mawile-workbench-for-auditing-llm-judges","json":"https://digestai.news/story/megagon-labs-releases-mawile-workbench-for-auditing-llm-judges.json","markdown":"https://digestai.news/story/megagon-labs-releases-mawile-workbench-for-auditing-llm-judges.md","slug":"megagon-labs-releases-mawile-workbench-for-auditing-llm-judges","headline":"Megagon Labs releases mawile workbench for auditing LLM judges","summary":"MAWILE is a new developer‑facing workbench that lets researchers and engineers audit the sensitivity of large‑language‑model (LLM) judges.  LLM judges are increasingly used to score model and agent outputs, but their verdicts can shift when irrelevant details of the prompt, rubric, input, or output change.  MAWILE addresses this by taking a user‑provided judge and a set of evaluation items, then generating controlled perturbations across four surfaces – the judge prompt, the scoring rubric, the target‑system input, and the target‑system output.  For each perturbation it re‑runs the judge and reports whether the verdict should stay the same or move in a specified direction, highlighting both robustness to noise and sensitivity to meaningful changes.\n\nThe tool works with binary, ordinal, and pairwise judges and does not require gold‑standard labels, making it practical for a wide range of evaluation pipelines.  MAWILE’s code is released openly on GitHub, allowing the community to adopt and extend the framework for more reliable LLM benchmarking.","keyPoints":["MAWILE audits LLM judges across prompt, rubric, input, and output perturbations","It validates controlled changes and localizes sensitivity without needing gold labels","The open‑source code is available at github.com/megagonlabs/mawile-judge"],"whyItMatters":"Reliable evaluation of LLM outputs is essential for model progress; MAWILE helps detect hidden biases and robustness gaps in judges, strengthening benchmark credibility.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":["Megagon Labs"],"models":[],"people":[]},"firstPublishedAt":"2026-09-22T04:00:00Z","updatedAt":"2026-09-22T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.AI","title":"MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators","url":"https://arxiv.org/abs/2609.22599","publishedAt":"2026-09-22T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Megagon Labs releases mawile workbench for auditing LLM judges\", 22 September 2026, https://digestai.news/story/megagon-labs-releases-mawile-workbench-for-auditing-llm-judges","publisher":"Digest AI","title":"Megagon Labs releases mawile workbench for auditing LLM judges","datePublished":"2026-09-22T04:00:00Z","url":"https://digestai.news/story/megagon-labs-releases-mawile-workbench-for-auditing-llm-judges"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}