What is Reference Data?

Reference data is a controlled set of values that other data uses for classification, such as country codes, currency codes, units of measure, product categories, or status codes. It changes less often than transactional data, but many systems depend on it at once. Some reference data comes from external standards, like ISO country codes, while organizations define the rest internally, like their own list of claim types.

Because so much depends on it, a small change in reference data can shift results across reports and models. For example, if a company splits one product category into two, a demand model trained on the old category may start grouping sales differently without any change to the model itself. Teams usually manage reference data with versioning and effective dates, so they can tell which set of values a past report or AI run used.

Frequently asked questions

What is the difference between reference data and master data?

Reference data classifies other data with a set of allowed values, such as country codes. Master data describes core business entities, such as customers, products, or suppliers.

Why does reference data need version control?

Many systems depend on the same values. Recording versions and effective dates lets teams trace which values a past report or model run used and explain differences in results.

Who owns reference data?

Often a data governance or data management team, with business owners who approve changes to each code set.