DigestAI news desk

Cut through the AI noise.

Research

Megagon Labs releases mawile workbench for auditing LLM judges

MAWILE is a new developer‑facing workbench that lets researchers and engineers audit the sensitivity of large‑language‑model (LLM) judges. LLM judges are increasingly used to score model and agent outputs, but their verdicts can shift when irrelevant details of the prompt, rubric, input, or output change. MAWILE addresses this by taking a user‑provided judge and a set of evaluation items, then…

1 source primary source

Key points

  • MAWILE audits LLM judges across prompt, rubric, input, and output perturbations
  • It validates controlled changes and localizes sensitivity without needing gold labels
  • The open‑source code is available at github.com/megagonlabs/mawile-judge

The tool works with binary, ordinal, and pairwise judges and does not require gold‑standard labels, making it practical for a wide range of evaluation pipelines. MAWILE’s code is released openly on GitHub, allowing the community to adopt and extend the framework for more reliable LLM benchmarking.

Read the original at arXiv cs.AI · by Jackson Hassell, Farima Fatahi Bayat, Pouya Pezeshkpour, Estevam Hruschka primary sourceOpen source ↗
Topics · follow one to build your own front page
Megagon Labs

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories