A golden dataset is a fixed, reviewed set of inputs and expected outputs that a team uses as the reference for evaluating a model, an AI application, or a data pipeline. Domain experts usually write or approve the expected answers, and the set covers both common cases and known difficult ones. In business intelligence, the term can also mean a certified dataset that reports should draw from.
Because every run is scored against the same cases, a golden dataset lets a team compare versions fairly. For example, a team testing a new prompt for a contract assistant might rerun the same 200 reviewed questions and compare accuracy with the previous version. The set needs its own maintenance. When the task, the policy, or the source data changes, some expected answers become wrong, and the team has to review and version the dataset so results stay comparable.