Data governance
Data governance covers eight practices applicable to the training, validation and testing data sets of a high-risk AI system, and the substantive standard that data sets be relevant, sufficiently representative, and, to the best extent possible, free of errors and complete. The legislator does not demand flawless data sets, but documented diligence proportionate to the intended purpose.
The requirements are set out in Article 10(1)–(6). Paragraph (2) prescribes eight concrete practices — including documenting relevant design choices, data-collection processes, data-preparation operations (labelling, cleaning, enrichment, aggregation), examination for biases, and the corrective measures taken — while paragraph (3) sets the substantive standard: relevance, representativeness, and error-free and complete data sets to the best extent possible, proportionate to the intended purpose. Paragraph (5) provides a narrow, conditional exception: it permits processing of special-category personal data solely for the purpose of bias detection and correction, subject to six cumulative safeguards — since the Digital Omnibus entered into force on 27 July 2026, this has been extended to all AI systems and general-purpose models, subject to a necessity test. Data governance is closely linked to Article 9 of the GDPR, but stands as an independent AI Act obligation.
Anyone developing or fine-tuning high-risk AI cannot simply say "this is the data we had" — examining for biases and documenting corrective steps is an independent, enforceable obligation.
Need documented AI-literacy training?
Article 4 is a duty of diligence: what counts is not knowledge in the abstract, but demonstrable, documented effort. Our starter package lets you begin free.
Start freeRelated terms
- Risk management systemThe risk management system is a continuous, documented, testing-based process spanning the entire lifecycle of a high-risk AI system, which identifies, evaluates and mitigates the risks the system poses to the health, safety and fundamental rights of third persons. It does not manage organisational risk — it protects a different interest than ISO 31000 or ISO/IEC 42001, and neither creates a presumption of conformity, since neither is a harmonised standard.
- Technical documentationTechnical documentation is the record demonstrating a high-risk AI system's compliance, covering the nine points of Annex IV, which must be drawn up before placing on the market, kept up to date, and retained for ten years. It is the record from which a regulator can later reconstruct how the system and its classification decision came about.
- High-risk AI systemAn AI system is high-risk if it is a safety component of, or is itself, a product covered by the Union product-safety legislation listed in Annex I, or if it falls within one of the eight areas and listed use cases of Annex III. Classification attaches to the intended purpose, not to the underlying technology.
- Data Protection Impact Assessment (DPIA)The Data Protection Impact Assessment (DPIA) is a separate legal instrument under the GDPR — not the AI Act — which the data controller must carry out where a type of processing is likely to result in a high risk to the rights and freedoms of natural persons. Many AI-based HR or credit-assessment solutions are typically subject to a DPIA, even where the AI Act's FRIA does not apply to the actor in question.
Related questions in the knowledge base (Hungarian)
- Milyen adatkormányzási és adatminőségi követelmények élnek a magas kockázatú AI tanítóadatára?A 10. cikk (2) nyolc adatkormányzási gyakorlatot ír elő a tanító, validációs és teszt adathalmazokra, a (3) bekezdés pedig az anyagi mércét: relevánsak, kellően reprezentatívak, a lehető legnagyobb mértékben hibamentesek és teljesek legyenek. Hibátlan adathalmazt a jogalkotó nem követel, hanem a rendeltetéshez mért, dokumentált gondosságot.
- Mikor kezelhető különleges személyes adat AI-torzítás vizsgálatához?A 10. cikk (5) bekezdése szűk, feltételekhez kötött kivétel: kizárólag torzításdetektálás és -korrekció céljából engedi különleges kategóriájú adat kezelését, hat kumulatív garanciával. A Digital Omnibus 2026. július 27-i hatálybalépése óta a felhatalmazás minden AI-rendszerre és általános célú modellre kiterjed, szükségességi teszt mellett.