DG-009HighProviderDetective
Statistical Testing and Validation
Providers must apply appropriate statistical methods to validate that training data meets quality thresholds and that the resulting model generalises appropriately to the intended deployment distribution.
Articles:Article 10(3)Article 15(2)
Evidence Examples
- Statistical validation methodology
- Distribution shift test results
- Validation set performance report
Standards
ISO 42001:2023 §8.4ISO/IEC 24029-1
Related controls
- DG-001Training Data Quality RequirementsArticle 10(3) requires that training, validation, and testing data sets are subject to data governance practices that ensure relevance,…
- DG-002Bias Detection and CheckingProviders must examine training, validation, and testing datasets for possible biases that could affect health, safety, or fundamental rights, and…
- DG-003Data Documentation and ProvenanceProviders must document the origin, collection methodology, labelling process, and relevant characteristics of all data sets used in training, validation,…
- DG-004Data Representativeness AssessmentArticle 10(3) requires that training data sets are sufficiently representative of the intended population and use-case context, and providers must…
- DG-005Data Relevance and Completeness CheckProviders must verify that data sets used for high-risk AI systems are relevant and complete for the system's intended purpose, documenting any known gaps…
- DG-006Data Annotation Quality AssuranceWhere data labelling or annotation is performed, providers must implement quality assurance procedures to ensure consistency, accuracy, and…
Related EU AI Act terms
- Data GovernancePractices and policies applicable to training, validation, and testing data sets used for high-risk AI systems, covering the design choices, data collection, and preparation processes, and ensuring datasets are relevant, sufficiently representative, and free of errors and complete.
- AccuracyThe requirement that high-risk AI systems achieve an appropriate level of accuracy in relation to their intended purpose, as specified in the technical documentation. Providers must declare the level of accuracy in the instructions for use.
- RobustnessThe ability of a high-risk AI system to maintain its level of performance under adverse conditions — including technical limitations, adversarial inputs, errors, or unexpected situations — or within foreseeable operating conditions outside the intended purpose.
- CybersecurityThe requirement that high-risk AI systems are resilient against attempts by third parties to alter their use, behaviour, or performance in ways that could result in risks to health, safety, or fundamental rights, including protection against data poisoning, adversarial examples, and model evasion attacks.
Upgrade when it needs to be
The free questionnaire returns preliminary signals. The Full Assessment turns real system material into a governed decision record - extraction, component separation, evidence, national overlays, and a review-ready dossier.
Start free risk preview