The following scene is an illustration, not a customer case. In July 2026, the credit risk team at a European lender hears that the EU has pushed back the AI Act’s high-risk rules. Their credit scoring model falls under Annex III, and the obligations that were due on 2 August 2026 now apply from 2 December 2027. Someone says what everyone is thinking: so we have until the end of next year. The documentation project moves into next year’s plan.
Fast-forward to December 2027. A reviewer asks the team to show the data the model in production was trained on: where it came from, how missing income values were handled, and which applicant groups the team checked for bias. The model has been retrained twice since the summer of 2026. The training set behind the current version was overwritten when the pipeline was rebuilt. The engineer who wrote the filter for incomplete applications has moved to another team, and the notebook used for it has changed several times since.
The date that matters is not 2 December 2027. It is the day you freeze the training data for the model you will be running on that date. For a team that retrains on a schedule, that day may come in the next twelve months. The record Article 10 asks for can only be captured at that moment. Afterwards, you can describe the pipeline you have, but not the one that produced the data.
Key takeaways
- The Digital Omnibus moved EU AI Act obligations for stand-alone high-risk AI from 2 August 2026 to 2 December 2027, a delay of sixteen months.
- The effective deadline for the data record is earlier: the day you freeze the training data for the model that will be live on that date.
- A usable record states what was decided about the data, why, what it affected, and who owns what is left, for one specific data version.
What the Digital Omnibus changed, and what it did not
The Digital Omnibus changed when the high-risk rules apply. It did not remove what they ask for. The regulation was published in the Official Journal on 24 July 2026 and entered into force on 27 July (as reported by Law & Technology and summarized by Gibson Dunn).
| Obligation | Original date | Date after the Digital Omnibus |
|---|---|---|
| Stand-alone high-risk AI systems (Annex III) | 2 August 2026 | 2 December 2027 |
| High-risk AI built into products covered by sectoral law (Annex I) | 2 August 2027 | 2 August 2028 |
Two details belong with your legal team rather than in a blog post. The Omnibus also amended some provisions related to Article 10, so read the consolidated text. And systems already on the market before the application date are treated differently from new ones, so whether a retrained model counts as a significant change is a question for counsel. This post is not legal advice. It is about the data work underneath either answer.
Why the real deadline comes before December 2027
The real deadline comes earlier because Article 10 is about the training, validation, and testing data behind a model, not about a document that sits next to it. It asks about the origin of the data, the preparation steps, the assumptions about what the data represents, whether it fits the intended purpose and setting, which biases were examined, and which gaps remain. Each of those answers is true of one data version and false of the next.
That makes the record a by-product of preparing the data, not a report you can write at the end. When an engineer drops incomplete applications, the reason, the scope, and the effect on approval rates are known that afternoon. A year later, the people may have moved on, the code may have changed, and the training set may no longer exist in the form the model saw. The team then has to rebuild evidence from memory, and a reviewer can tell.
What a reviewer-ready data record looks like
Most teams already write something down. The difference is whether the note would answer a reviewer’s next question. Here is the same credit scoring dataset, recorded two ways:
| Article 10 asks about | A weak record | A record a reviewer can use (illustration) |
|---|---|---|
| Origin and collection purpose | Loan application data from the CRM. | Applications from the retail lending system for the stated period, collected to assess loan eligibility. Records from a partner channel were excluded because applicants gave consent for marketing only. |
| Preparation steps | Removed incomplete records. | Removed applications with missing income. Most came from self-employed applicants, so approval rates by employment type were compared before and after, and the result is stored with this data version. |
| Assumptions about what the data represents | Data is representative of customers. | Assumes past approvals reflect current lending policy. That assumption no longer holds after the 2025 policy change, so earlier decisions are labeled and weighted separately. |
| Bias examination and remaining gaps | Checked for bias. No issues found. | Compared outcomes by age band and employment type for the intended use. Applicants with thin credit files remain underrepresented. This is logged as a known limitation, owned by the credit risk lead, with a review date. |
The right-hand column is not longer for its own sake. Each entry names a decision, the reason for it, what it affected, and who owns what remains. That is what lets a reviewer, or your own team a year later, check the work instead of taking it on trust.
What it takes to build this yourself
Teams that build this in-house usually find that writing the record is the easy part. The hard part is the machinery around it. Someone has to capture the record on every preparation run, not only the ones that feel important. Each model version has to stay tied to the exact data version it used, through retraining and pipeline rebuilds. When a source or a field changes, someone has to know which checks to rerun, and for which models. All of it has to keep working after the people who built it move on.
Most teams start with a spreadsheet or a wiki page. Within a few months it describes the pipeline the team meant to run, not the one that ran. Closing that gap between the data work and the record of it is the part CUBIG builds into Syntitan, so data teams do not have to build and maintain it themselves.
Start with the model that will be live in December 2027
You do not need to document every model at once. Start with the one most likely to be in production on 2 December 2027, and work back from its next retraining date.
- Find the freeze point. With your legal team, confirm which systems may fall under Annex III. For each one, find when the training data for the version you will run in December 2027 gets prepared.
- Record decisions as they happen. At that preparation run, write the record in the right-hand column above, next to the code and the data version, not in a separate document weeks later.
- Tie the model to that data version, and recheck on change. Keep the exact data version each model version was trained and tested on. When a source, a field, or the intended use changes, rerun the checks and mark thin evidence as inconclusive rather than passed.
The third step is where most teams have the least in place. Linking a model to the exact data state it ran on is what lets you answer a question about last year’s model with last year’s data. The same habit pays off outside regulation, as we saw in what a second AI deployment should inherit, and it starts with a clear data provenance record.
Where Syntitan fits
Whether a system is high-risk, and whether its documentation meets the Act, are decisions for your legal and risk teams. Syntitan, CUBIG’s AI-Ready Data Platform, handles the data work underneath. Teams prepare data for a defined use case and validate whether it meets the criteria for that task, using the model or agent they selected under fixed evaluation conditions. Each release leaves a data version with before-and-after evidence for its preparation steps, the evaluation conditions, and a validation result of qualified, not qualified, or inconclusive. When the underlying data changes, the team can see what changed against the previous version and revalidate. That is the record a reviewer asks for in December 2027, produced as part of the data work instead of written up afterwards.
If one of your models will still be running in December 2027, find out when its next training set gets built. That is the date to plan for. To see how the data side can work, explore Syntitan.
