DigestAI news desk
Research updated

XAI-Arena uses LLMs to benchmark explainable AI quality

Researchers have introduced XAI-Arena, a framework designed to address the subjectivity and scalability issues inherent in evaluating Explainable AI (XAI) methods. By leveraging Large Language Models (LLMs) as judges, the system provides a reproducible mechanism for assessing explanation quality across multiple dimensions, including clarity, faithfulness, and trust calibration.

1 source primary source

Key points

  • XAI-Arena is an LLM-as-a-judge framework for scalable, reproducible evaluation of XAI explanation quality.
  • The framework assesses explanations across eight dimensions, including clarity, faithfulness, and trust calibration.
  • Human validation showed a strong positive correlation (Spearman's rho=.693) between LLM and human ratings.

The study benchmarks various XAI methods across different datasets and machine learning models while accounting for different stakeholder personas. The approach aims to make comparative assessments more consistent and scalable than traditional human-only evaluations.

Validation against human ratings revealed a strong positive association, with a Spearman's rho of .693 (p<.001). This suggests that LLM-based evaluations can effectively capture systematic differences in explanation quality, offering a viable tool for standardizing XAI research and development.

Read the original at arXiv cs.AI · by Yanfei Hu Fleischhauer, Alona Zharova, Nadja Klein, Stefan Feuerriegel primary source Open source ↗

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.

Comments

via GitHub Discussions

More in Research

All →

Related stories