New Project on Reliable Evaluation of Generative AI

The Vector Foundation funds the project ZEGeKI (Zuverlässige Evaluation Generativer KI, English: Reliable Evaluation of Generative AI).
The motivation of this project lies in the observation that, following open science principles, benchmarks for evaluating AI systems (not limited to, but including Generative AI) are published on the Web. At the same time, Generative AI systems are trained using large corpora of Web data, which may include the benchmarks. This makes the evaluation very unreliable: if the AI has “seen” the benchmark data during training, performance measurements on this benchmark data are not very significant, since they can be, at least partly, attributed to data leakage. The problem is further excarbated for AIs systems being able to perform online search on the Web, and for those being continuously trained based on user inputs.
The goal of the project is to come up with a different benchmarking strategy. Instead of using fixed datasets, the proposal is to generate benchmark data on the fly, and to evaluate AI systems based on such dynamic benchmarks intead of static benchmark data, while still maintaining an evaluation setup which allows for fair comparison and evaluation.