Back to Search View Original Cite This Article

Abstract

<p>KI-Benchmark Deutsch is a recurring computational benchmark for evaluating large language models on realistic German-language business and administrative tasks. Each run evaluates a fixed model roster on 24 held-out tasks across six categories using fixed rubrics, reference answers where applicable, and a name-blind cross-vendor judge panel with leave-one-provider-family-out assignment. This protocol defines answer generation, judge eligibility, score parsing, aggregation, coverage safeguards, quality assurance, and versioned publication. Applying it yields model- and category-level aggregate scores on a 0–100 scale plus a machine-readable snapshot; held-out tasks and raw answers remain private to limit contamination and gaming.</p>

Show More

Keywords

tasks fixed model heldout answers

Related Articles

PORE

About

Connect