AI-Ready Data, Syntitan

SS1/23 Principle 3.2: What Validators Will Ask About Your Model Data

Principle 3.2 of the PRA’s SS1/23 sets out what a UK bank has to show about the data used to develop a model. The data should suit the intended use, match the chosen methodology, and represent the customers, products, and portfolios the model will serve. It should carry no inappropriate bias and respect privacy rules. Where it falls short, any adjustment or proxy used to fill the gap has to be documented and independently validated.

Most of Principle 3.2 reads like good modelling practice. The point teams tend to underestimate is (d): filling a representativeness gap does not end validation. The method used to fill it becomes something the validator has to check, with its own assumptions, rationale, and evidence.

Key takeaways

  • SS1/23 Principle 3.2 asks for development data that is suitable, consistent with the methodology, representative of the target portfolio or customers, free of inappropriate bias, and used in line with privacy rules.
  • Filling a gap with an adjustment or proxy is allowed. Under point (d), the adjustment’s assumptions, factors, and rationale must themselves be documented and independently validated.
  • A useful test for any adjustment: did it improve results for the segment it was meant to fix, or only the overall score?

SS1/23 at a glance

PRA SS1/23 model risk management principles for banks at a glance
ItemSS1/23
Issued byPrudential Regulation Authority (Bank of England)
Applies toUK banks, building societies, and PRA-designated investment firms with internal model approval for regulatory capital
In effectFrom 17 May 2024, first published 17 May 2023, latest edition April 2026
Models in scopeIncludes AI and machine learning models
Self-assessmentUpdated at least annually
StructureFive principles: model identification and tiering, governance, development and use, independent validation, risk mitigants

This post is not regulatory advice, and your model risk and validation teams decide how the statement applies to your models. It covers the data work underneath Principle 3.2.

Principle 3.2, point by point: what a validator will ask

Each point in Principle 3.2 turns into a question during independent validation. Under Principle 4.1, the validation function gives an objective opinion on the accuracy, relevance, and completeness of the development data. These are the questions the data has to answer, according to the April 2026 text.

PRA SS1/23 Principle 3.2 (a) to (e), read as the questions an independent validator asks about model development data (first three columns summarize SS1/23; the last column is our practical example, not regulatory text)
PointWhat it expectsThe validator’s questionPractical evidence (our example)
(a) Suitability and represent­ativenessData suits the intended use, matches the methodology, and represents the target portfolio, products, assets, or customersDoes this data look like the population the model will actually score?Segment coverage compared with the current or expected portfolio, with the date of the comparison
(b) Bias and privacyNo inappropriate bias in development data, and data use in line with privacy rulesWhich biases were checked for this use, and is the data allowed to be used this way?Bias checks tied to the intended use, and the basis for using each personal data source
(c) Non-representative dataImpact assessed, limitation reflected in the model’s tier, owners and users informedWhere is the data not representative, and what does that do to the model?Impact on segment-level error, the tier decision, and the note shared with owners and users
(d) Adjustments and proxiesDocumented and validated, with assumptions, adjustment factors, and rationale independently validated and recorded in the inventoryWhy this adjustment, what does it assume, and what did it change?Reason, method and settings, before-and-after data comparison, and results with and without the adjustment
(e) Interconnected and alternative dataIdentified and recorded in the inventory, with added complexity and uncertainty reflected in the tierWhich inputs come from other models or from alternative and unstructured sources?A list of such inputs per model, with their source and how they were treated in tiering

Principle 3.5 adds that model documentation should describe the data sources, any proxies, and the results of data quality, accuracy, and relevance tests. The first three columns above summarize SS1/23. The last column is our practical example of evidence that can answer each question; SS1/23 does not prescribe these specific items.

Where teams get stuck: when an adjustment helps the wrong group

Take a bank extending a credit model to borrowers whose income comes from several platform jobs. The development data has few of them, so the team fills the gap three ways. It oversamples the cases it has, borrows a proxy (a stand-in variable) for income stability from a similar product, and adds reconstructed rows to balance the classes. On the holdout set, the data kept aside for testing, overall error falls.

Then the team splits the result by segment. Error for the platform-income borrowers, the group the adjustment was meant for, has barely moved. The added data improved the model for customers it already handled well. Is the adjustment a success? Under point (d), the answer has to come from evidence about the adjustment, not from the overall score. Point (c) adds that any remaining limitation should be reflected in the model’s tier, the risk rating that sets how much validation scrutiny the model gets.

A review note for that adjustment could look like this (illustration):

  • Adjustment. Oversampling, an income stability proxy from a similar product, and reconstructed rows for the platform-income segment.
  • Assumption. The proxy moves with platform income. Checked on customers who appear in both products, the relationship holds for salaried income and only weakly for platform income.
  • Effect overall. Error on the holdout set fell with the adjustment.
  • Effect on the target segment. Error for platform-income borrowers was unchanged with and without the adjustment, under the same model and evaluation conditions.
  • Decision. Keep the adjustment for the general population, record the segment as a known limitation owned by the model owner, keep the higher tier, and collect more platform-income cases before the next review.

The note does not prove the model is right. It lets a validator see what the adjustment did and challenge the decision, which is what point (d) is for.

A practical checklist for your next SS1/23 self-assessment

Firms are expected to update their SS1/23 self-assessment at least once a year. The checks below are our practical recommendation for Principle 3.2, not a list that SS1/23 prescribes:

  1. Find the representativeness gaps. List models whose portfolio, products, or customer base has moved away from the development data, including new products and segments.
  2. Inventory every adjustment and proxy. Record each one in the model inventory with its owner, assumptions, and rationale, as point (d) expects.
  3. Measure each adjustment on the segment it targets. Rerun with and without the adjustment under the same model and evaluation conditions, and report the target segment separately from the overall result.
  4. Flag interconnected and alternative data sources. Point (e) asks for these to be identified, recorded, and reflected in the model’s tier.

Most of this stays with your model risk process: the inventory, the bias and privacy decisions, and tiering. The repetitive data work is step 3. The same comparison has to be rerun whenever the data, the model, or the annual review comes around, with the before-and-after data kept for each version. That part is what CUBIG builds into Syntitan and DTS, so teams do not have to rebuild the notebooks each time.

Where Syntitan and DTS fit

Whether a model meets SS1/23, and how it is tiered, are decisions for the bank’s model risk and validation functions. Syntitan, CUBIG’s AI-Ready Data Platform, handles the data work behind step 3. Teams diagnose how classes and segments are skewed, then prepare the data with steps such as cleaning, row augmentation, and class balancing. Where those are not enough, DTS reconstructs the data the task needs.

For each adjustment, the team gets the data state before and after, saved together. For reconstructed data, DTS adds a comparison of distributions, variable relationships, and duplicates against the original. With a supported model and fixed evaluation conditions, the team gets results with and without the adjustment, plus a validation result of qualified, not qualified, or inconclusive for the task. The data, the training and scoring code, and the results can be downloaded, so the bank’s independent validation team has a traceable record to review instead of a notebook to reconstruct.

Start with one model whose development data under-represents a segment, and check whether its adjustment helped that segment or only the overall score. To see how the data side can work, explore Syntitan. For the EU side of the same question, see how the EU AI Act delay moves the real data deadline earlier.

Explore Syntitan, CUBIG’s AI-Ready Data Platform

References

  1. Bank of England, PRA, SS1/23 Model risk management principles for banks (2023)
  2. Bank of England, PRA, SS1/23 Model risk management principles for banks, April 2026 edition (2026)
  3. CUBIG, The EU AI Act Deadline Moved. Your Real Deadline Is Earlier.
  4. CUBIG, What Is DTS?
  5. CUBIG, Syntitan

FAQ

What is PRA SS1/23?

SS1/23 is the Prudential Regulation Authority's supervisory statement setting out model risk management principles for UK banks. It applies to UK-incorporated banks, building societies, and PRA-designated investment firms with internal model approval to calculate regulatory capital, and it covers AI and machine learning models.

When did SS1/23 take effect?

The expectations took effect on 17 May 2024. Firms were expected to complete an initial self-assessment before then and to update it at least annually. The PRA published an updated edition in April 2026.

What does SS1/23 Principle 3.2 require for model data?

It asks firms to show that development data are suitable for the intended use, consistent with the chosen methodology, representative of the target portfolio, products, assets, or customers, free of inappropriate bias, and used in line with data privacy rules. It also covers non-representative data, adjustments and proxies, and interconnected or alternative data.

What happens if model development data are not representative under SS1/23?

The firm should assess the potential impact, reflect the limitation in the model's tier classification, and make model owners and users aware of it. Any adjustment or proxy used to compensate should be documented, independently validated, and recorded in the model inventory.

Can banks use synthetic or augmented data under SS1/23?

SS1/23 does not name synthetic data. In our reading, when augmented, reconstructed, or proxy data are used to compensate for unrepresentative data, they fall under Principle 3.2(d), so the adjustment, its assumptions, and its rationale need to be documented and independently validated. Confirm the approach with your model risk function.

What should model documentation say about data under SS1/23?

Principle 3.5 expects documentation to describe the data sources, any data proxies, and the results of data quality, accuracy, and relevance tests, alongside the methodology, performance tests, limitations, and any model adjustments.