Request a pilot

Tell us more about your project and data

Email is not valid.

Email is not valid

Company name is not valid

Phone is not valid

Some error text

Thank you for contacting us!

Thank you for contacting us!

We'll get back to you shortly

Geberit Quotes

After running pilots with several annotation providers, Label Your Data delivered the strongest results by a clear margin, standing out on turnaround time, annotation quality, and the responsiveness of their feedback loops.

Quotes
Geberit
Maxime Debarbat

Maxime Debarbat

Senior ML Engineer (GenAI)

Trusted by ML Professionals

Ouster
Searidge Technologies
Zendar
Advanced Farm
ABB
Toptal
UiPath
Respeecher
Yale
Thorvald
Back to blog Back to blog
Published September 30, 2026

EU AI Act: What It Means for Your Training Data

Karyna Naminas
Karyna Naminas Linkedin CEO of Label Your Data
EU AI Act: What It Means for Your Training Data

TL;DR

  1. Your obligations depend on the AI system’s intended use and your company’s role. Being based outside the EU doesn’t automatically put your team outside the Act’s scope.
  2. Transparency rules began applying in August 2026. The main high-risk requirements follow in December 2027, with relevant systems embedded in regulated products following in August 2028.
  3. For high-risk systems, Article 10 covers how teams source and prepare training, validation, and testing data, including annotation practices, bias assessment, and gaps in coverage.
  4. Record labeling decisions, QA findings, and known data limitations as you develop the model. These records help your team explain its data choices and identify work still needed before deployment.

Data Annotation Services

First annotation is FREE

LEARN MORE

The EU AI Act makes training data quality a compliance issue for teams building high-risk AI systems. That brings familiar development questions into focus: where did the data come from, how did annotators label it, and does it cover the conditions the model will face?

For machine learning teams, answering those questions starts during dataset preparation. Keeping labeling instructions and QA results tied to each dataset version gives you a record of how the data changed and which limitations remain.

Here’s how EU AI Act implementation in 2026 affects that work, which dates matter, and what your team should prepare.

EU AI Act Updates 2026: Key Changes for AI Teams

The August milestone brought new transparency obligations and enforcement powers. Separately, the AI Omnibus entered into force on July 27, 2026, extending the main high-risk deadlines.

The EU AI Act transparency requirements include informing people about certain AI interactions and marking or disclosing AI-generated content. These requirements concern how AI systems and their outputs are identified. Data annotation serves a different purpose: labeling the examples used to train, validate, and test models.

The Commission’s timeline distinguishes the EU AI Act entry into force date from the dates its requirements start applying. There isn’t one EU AI Act effective date for every obligation:

DateMilestone
August 1, 2024The Act entered into force.
February 2, 2025Prohibited-practice rules began applying.
August 2, 2025General-purpose AI (GPAI) model obligations began applying, subject to transitional rules.
August 2, 2026Transparency obligations and new enforcement powers began applying.
August 2, 2027Providers of GPAI models placed on the market before August 2, 2025 must comply with the applicable GPAI obligations.
December 2, 2027High-risk requirements apply to Annex III systems, such as certain recruitment and creditworthiness systems.
August 2, 2028High-risk requirements apply to relevant AI systems embedded in regulated products.

The dates above follow the Commission’s current implementation timeline.

Some transitions are narrower. For example, providers of relevant systems already on the market before August 2, 2026 have until December 2, 2026 to meet Article 50(2)’s marking and detection obligation, according to the Commission’s enforcement FAQ.

For AI teams preparing high-risk systems, the extension provides time to fix data gaps and establish review processes. Document dataset decisions now, while the people who made them can explain the reasoning.

Who Must Comply With the EU AI Act?

EU AI Act risk levels and related obligations

The EU Artificial Intelligence Act applies to providers placing AI systems or GPAI models on the EU market, EU-based deployers, and certain non-EU providers and deployers whose AI system outputs are used in the EU.

A provider develops a system, or has it developed, and markets or puts it into service under its name. A deployer uses a system in a professional context. GPAI model providers have a separate set of obligations.

The commonly used EU AI Act risk categories help explain the different requirements:

CategoryWhat it means for an AI team
Prohibited practicesCertain uses are banned, subject to the law’s specific conditions and exceptions.
High-risk systemsRequirements cover areas including risk management, data governance, technical documentation, and oversight.
Systems with transparency dutiesCertain AI interactions or generated outputs require notices, marking, or disclosure.
Minimal-risk usesThe Act generally imposes no specific system requirements, though other laws can still apply.

These categories overlap: transparency duties can also apply to a high-risk system. GPAI obligations sit alongside the system rules.

For a CV team, an image recognition application doesn’t become high-risk simply because it processes images. Assess its intended purpose against Article 6 and the relevant annexes. Choosing a different machine learning algorithm won’t, by itself, change that purpose or the applicable classification.

EU AI Act Requirements for Training Data

High-risk system providers and general-purpose AI model providers have different data-related obligations. A company can hold both roles, so assess each separately.

EU AI Act Article 10: Data quality and annotation

EU AI Act: Training data quality requirements

Article 10 addresses data governance for training, validation, and testing datasets used to develop relevant high-risk systems. It explicitly includes data annotation, alongside other preparation operations.

The datasets must be relevant and sufficiently representative for the system’s intended purpose, and as complete and free of errors as possible. The law doesn’t prescribe one annotation-accuracy percentage for every project.

For your machine learning datasets, that means examining suitability alongside label quality. The following records are practical examples of evidence your team can prepare, rather than mandatory document formats:

Area to assessExample record
Data origins and collectionSource register with collection context and known restrictions
Annotation and preparationGuideline versions, transformation history, and review results
Representation of intended conditionsCoverage breakdown against the deployment setting
Bias and identified gapsAssessment findings, mitigation work, and unresolved limitations

Consider a pedestrian-detection dataset. Reviewers might agree that its bounding boxes follow the labeling instructions, yet almost every frame could show clear daylight. If the system must operate at night, that agreement says little about whether the dataset covers the intended environment.

The annotation guidelines also need to resolve cases such as partially visible pedestrians. Otherwise, different batches may encode different definitions of the same class. A useful review records both the disagreement and the rule the team adopted, then identifies earlier annotations that need correction.

At Label Your Data, annotation specialists work with your team to clarify complex edge cases and refine guidelines as the project develops. QA checks follow your project’s requirements, helping keep labels consistent as annotation volume grows. 

EU AI Act Article 53: GPAI training records

EU AI Act: GPAI training records

Article 53 sets separate obligations for GPAI model providers. These include technical documentation, information for downstream providers, a copyright policy, and a public summary of training content.

Providers of qualifying free and open-source GPAI models without systemic risk can be exempt from the technical and downstream documentation duties. The copyright policy and public training-content summary remain required.

The Commission supplies an official training-content summary template. This is a summary obligation, not a general instruction to publish the full training dataset.

Providers of GPAI models with systemic risk face additional evaluation, risk-management, and cybersecurity duties. Using a third-party foundation model doesn’t automatically make your company its provider; assess your role and any modifications you make.

How to Prepare Your Data for EU AI Act Compliance

EU AI Act: Annotation and labeling requirements

EU AI Act compliance requires work across the system’s lifecycle. For high-risk AI systems, that includes conformity assessment where required and post-market monitoring. 

Data teams can prepare by taking the following steps:

  1. Confirm the system’s purpose and your role. Have product, engineering, and legal agree on the intended use, applicable requirements, and deadlines. Revisit that assessment when the use case changes.
  2. Connect each model release to its data. Record which dataset versions supported training and evaluation. Keep the preparation records accessible so a later reviewer can trace a finding to the affected samples.
  3. Turn gaps into assigned work. Give each unresolved issue an owner and an acceptance criterion. For missing night scenes, this could mean collecting suitable data and evaluating performance on that subset before release.
  4. Agree on supplier deliverables. When evaluating data annotation services, specify what accompanies the labels, such as review findings or a record of guideline changes. Confirm what the supplier can provide before production starts.

When comparing data annotation pricing, include the review effort and delivery records your project needs. A quote based only on annotation volume can leave that work unaccounted for.

Security and oversight also need owners beyond the annotation team. Under EU AI Act Article 14, human oversight concerns how people oversee a high-risk system during use, including understanding its limitations and intervening when needed. Reviewing labels supports data quality, but doesn’t establish those operational controls.

Article 15 addresses accuracy, robustness, and cybersecurity, including relevant protections against data poisoning. As a practical measure, restrict who can alter approved datasets and retain a record of changes. Security teams should assess how to detect malicious changes and test the system against relevant attacks.

Keep data preparation and change records accessible to support technical documentation and investigations after deployment. When monitoring reveals a failure associated with a particular condition, the team should be able to inspect the corresponding data and decide whether collection, labeling, or another part of the system needs work.

Label Your Data is an ISO 27001-certified data annotation company with secure workflows for data transfer, storage, and annotation. Our teams can work within your preferred annotation platform, so you can retain your existing tools and review workflows.

About Label Your Data

Label Your Data is an AI training data provider offering managed data annotation services for complex AI environments. Our annotation experts help resolve difficult edge cases and maintain consistent labels. Share your labeling requirements and delivery needs, then assess our workflow and annotation quality through a no-cost pilot before scaling.

Data Annotation for Complex Environments Data Annotation for Complex Environments

Rely on consistent, high-quality output for complex datasets, detailed taxonomies, and edge cases.

Structured Quality From Pilot to Production Structured Quality From Pilot to Production

Get quality engineered into every step through onboarding, evolving guidelines, QA, and continuous feedback.

Flexible and Scalable Operations Flexible and Scalable Operations

Adjust team capacity, project size, and delivery model as you scale, with no setup fees or long-term lock-ins.

An Integrated Delivery Partner An Integrated Delivery Partner

Align on goals, workflows, and expectations with a team that integrates into your process from day one.

Projects Led by Annotation Experts Projects Led by Annotation Experts

Work with former annotators who understand annotation complexity, quality standards, and high-volume delivery.

See how much your annotation project will cost

First annotation is FREE

LEARN MORE

FAQ

What is the EU AI Act?

arrow

The EU AI Act is the European Union’s regulation for artificial intelligence. It bans certain uses, sets requirements for high-risk systems, and introduces transparency duties for specific AI interactions and outputs. General-purpose AI (GPAI) models have a separate set of obligations.

Has the EU AI Act passed?

arrow

Yes, the EU AI Act is law and entered into force on August 1, 2024. Its requirements apply in stages, so your compliance deadline depends on the system and the obligations that apply to your company.

Does the EU AI Act apply to the US?

arrow

It can apply to US companies that place AI systems or general-purpose AI models on the EU market. It also covers certain US providers and deployers whose AI system outputs are used in the EU. Your role and use case determine the applicable obligations.

What are the key changes in the EU AI regulations for 2026?

arrow

Under the EU AI Act, transparency obligations and new enforcement powers began applying on August 2, 2026. The July 2026 AI Omnibus also extended the main high-risk deadlines to December 2, 2027 for Annex III systems and August 2, 2028 for relevant systems embedded in regulated products.

What are the main requirements of the EU AI Act?

arrow

For high-risk systems, requirements cover data quality and risk management throughout development and use. Teams must also address technical documentation, logging, human oversight, accuracy, robustness, and cybersecurity, alongside applicable conformity-assessment and monitoring duties. Other systems and general-purpose AI models have different obligations.

How to comply with the EU AI Act?

arrow
  • Confirm your company’s role and assess each system’s intended use and risk classification.
  • Identify applicable requirements and deadlines, then assign responsibility for each.
  • For high-risk systems, document dataset preparation, quality checks, and known limitations.
  • Implement the oversight, security, and transparency measures that apply.
  • Complete required assessments and maintain monitoring and incident-reporting processes.

Written by

Karyna Naminas
Karyna Naminas Linkedin CEO of Label Your Data

Karyna is the CEO of Label Your Data, a company specializing in data labeling solutions for machine learning projects. With a strong background in machine learning, she frequently collaborates with editors to share her expertise through articles, whitepapers, and presentations.