Proceedings of the First Workshop on Comparative Performance Evaluation: From Rules to Language Models

Chapitre

Beyond BLEU: Ethical Risks of Misleading Evaluation in Domain-Specific QA with LLMs

Auteur

Ayoub Nainia, Régine Vignes-Lebbe, Hajar Mousannif, Jihad Zahir

Maison d'édition

Association for Computational Linguistics (ACL)

Année de publication

2025

ISBN

978-954-452-102-8

Type

chapitre

Chercheur

ZAHIR Jihad

Voir le profil du chercheur →