A German AI consortium has released Soofi S, a 30 billion parameter open-source language model that performs competitively on benchmarks in both English and German. The release comes with a transparency caveat: the team discovered that test questions from the GPQA science benchmark accidentally leaked into the training data.

The community flagged the contamination after examining the publicly available training dataset. Rather than ignore the issue, the consortium removed GPQA from its evaluation suite and recalculated all benchmark results in version 3.0 of its technical report. This move reflects a growing emphasis on methodological rigor in AI benchmarking, where data leakage can artificially inflate performance metrics.

Soofi S positions itself as a competitive open alternative to proprietary models in two major markets: English-speaking regions and German-speaking Europe. The 30B parameter size sits in a practical sweet spot for organizations seeking models that run on modest hardware while maintaining strong reasoning capabilities. The bilingual strength matters for enterprises in Germany and Austria that need native-language performance without relying on US-based closed models.

The data contamination discovery underscores a broader challenge in the AI field: training datasets often contain internet-scraped material that overlaps with public benchmarks. Previous models from larger labs have faced similar criticism. The consortium's response—transparency about the error and recalculation of results—sets a standard that helps the community trust performance claims.

What remains unclear is whether Soofi S still maintains its benchmark advantages after the corrections. The team's willingness to publicly acknowledge and fix the problem demonstrates maturity in the open-source AI space, though it also raises questions about how thoroughly baseline benchmarks were vetted before publication.

For German-speaking organizations, Soofi S offers a locally-trained alternative that sidesteps reliance on English-first models. The open release enables fine