AI-Ready Data

Which Version of the Data Passed the Test?

Abstract illustration of three translucent data version cards stacked in sequence, with the middle validated version in deep violet and the newest version in periwinkle

Consider a team that approves a support-ticket routing model in June after it passes validation. The model uses a reference table to map product codes to categories. In August, the product team splits one category into two and renames another. Both table versions are preserved, so the change is easy to inspect. But does the June validation still support using the model with the updated table?

Data versioning records what changed. The team still needs to determine whether earlier validation results apply to the updated data. The scope of revalidation depends on the change and its effects across the system.

Key takeaways

  • Classify the data change and check its dependencies before deciding which earlier validation results you can reuse.
  • Google engineers gave this advice for production machine learning years ago: when an input can change, keep a frozen version and switch only after the new one has been validated.
  • Rerun the cases that use changed records, plus a sample of the rest. As a default, rerun everything when a field’s meaning or the pass criteria change.
  • Record each validation result with the data version, evaluation set version, pass criteria, and model and prompt used for the run. This gives the team a reference for assessing the next change.

New versioning features leave revalidation scope to the team

Neon, Tonic, and Snowflake have each shipped versioning features recently. On September 25, 2026, Neon extended branch reset so that it rolls back a branch’s object storage along with its Postgres data. On September 21, Tonic Fabricate 4.30 added project versions for Enterprise users, with a published version’s data locked by default. Snowflake’s Semantic Studio entered public preview in August with Git-backed version control for change tracking.

These features make it easy to restore a state or compare two states. After a change, the team still has to decide which earlier validation results it can rely on. To assess whether an approval still applies, the team needs to link each data version to the task, the validation run, and its results.

What the research on production ML already says

In Hidden Technical Debt in Machine Learning Systems (NeurIPS 2015), engineers at Google described unstable data dependencies: an input that changes over time can harm the system that consumes it, even when the change is an improvement. Their suggested mitigation is to create a frozen version of the input and use it until an updated version has been fully vetted. The paper also describes the costs of versioning: a frozen version can go stale, and someone has to maintain several versions of the same input. In a separate section, the authors warn that in a learned system, changing one input can change how the model uses the remaining inputs.

Google built change checks into its own ML pipeline. Data Validation for Machine Learning (SysML 2019), also from Google, describes a validation step that asks whether there are significant changes between successive batches of training data, or between training and serving data. Spotting a change still leaves the team to decide how much of the earlier validation it invalidates.

Retrieval systems show a similar effect. In HoH (ACL 2025), outdated information in a retrieval context distracted language models from correct information and lowered answer accuracy, and it could mislead them even when current information was available. Adding current information may not be enough if the model still retrieves outdated records without distinguishing them.

How data changes affect earlier validation results

The routing team can sort the August change before deciding what to rerun. The framework below is CUBIG’s practical guide, not a rule prescribed by the cited papers. Adjust the scope to the system’s dependencies and risk.

When earlier AI validation results may be reused after each kind of data change (CUBIG practical guide, illustration)
Kind of changeExample (illustration)When earlier results may be reusedWhat to recheck
Values in existing records editedOne mapping row now points to a different categoryPotentially reusable for cases that never use the edited rows, if the change has no indirect effects on shared rules, retrieval, or downstream behavior; confirm with regression checksCases that use the edited rows, and any rule built on them
Records added or removedA category is split into two; an old one is retiredPossibly reusable for existing cases if nothing they use was removed; a new category can change how existing tickets are classified, so spot-check a sampleNew cases for the added records; confirm retired records are excluded from inputs and retrieval, or clearly marked
A field’s meaning changed while values look normalA region field now holds sales territory instead of shipping regionDo not reuse results for cases that depend on the field without revalidation, even if the values still pass format checksEvery case that uses the field; rerun the full evaluation
Evaluation set or pass criteria changedStricter criteria for urgent ticketsReassess earlier results against the new evaluation set or criteria before using them to support approvalRerun against the new set or criteria and record a new baseline; keep the old result labeled with its criteria

The routing team’s August change falls in the first two rows: one category was renamed and one was split into two. The cases that use those categories, and any rule built on them, need a new run on the new version of the table. The June results for other tickets may be reusable, but a split can change how existing tickets are classified, so rerun a sample of them as well.

How to decide what to revalidate after a data change

  1. Record what each validation result rests on. Store the data version, the evaluation set version, the pass criteria, and the model and prompt with every validation result.
  2. Classify the change. Use the four kinds in the table above.
  3. Recheck the affected cases first. Rerun the cases that use the changed records and a sample of the rest. As a default, rerun the full evaluation when a field’s meaning or the pass criteria change.
  4. Exclude superseded records from active inputs and retrieval. Keep the archive. If old rows or documents must stay searchable, mark them so the AI can tell them apart from current ones.
  5. Record the new result against the new version. Keep the June result labeled with its version and criteria, so the team can interpret each result under the conditions used at the time.

When possible, evaluate model or prompt changes separately from data changes. This makes it easier to identify which change affected the results. Teams doing this manually must keep each result linked to the data version and evaluation conditions used for that run. Syntitan links data versions to evaluation conditions and validation results, helping teams keep the evidence for each run together.

Where Syntitan fits

The team that owns the task decides which changes matter for the business and whether a result is good enough. Syntitan, CUBIG’s AI-Ready Data Platform, releases each data version with its evaluation conditions and validation result attached: Qualified, Not Qualified, or Inconclusive for a named task with a selected model or agent. Approval for production stays a separate decision the team makes. When the data changes, the team compares the updated version with the version used in the last successful validation, then revalidates it before use.

Pick one AI task already in production. Identify the data version used in its last successful validation and list the changes made since. For each change, check whether the team has evidence that the updated data still meets the task’s criteria. Explore Syntitan to see how data versions, evaluation conditions, and validation results can be kept together.

Related reading

Explore Syntitan, CUBIG’s AI-Ready Data Platform

FAQ

What is data versioning?

Data versioning keeps identifiable versions of a dataset over time, so a team can see what changed, restore an earlier state, and tie a model run or evaluation to the exact data it used.

Is data versioning the same as revalidation?

No. Versioning records and restores states of the data. Revalidation decides whether an earlier validation result still holds after a change, and reruns the parts it no longer covers. Versioning makes revalidation possible but does not decide its scope.

When does an AI validation need to be rerun after a data change?

Rerun the affected cases when values in records they use are edited. When records are added or removed, rerun the cases that use them and a sample of the rest, since a new category or record can change other outputs. As a practical default, rerun the full evaluation when a field's meaning changes or when the evaluation set or pass criteria change, because earlier results alone may not establish whether the system meets the updated requirements.

What should a validation record include?

The data version, the evaluation set version, the pass criteria, the model and prompt or agent configuration, the result, and the date. Without these, the team cannot tell which approval a later change affects.

What does Hidden Technical Debt in Machine Learning Systems recommend for changing data?

The 2015 NeurIPS paper by Sculley and colleagues suggests using a frozen version of an unstable input until an updated version has been fully vetted. It also notes that versioning has costs, such as a frozen version going stale.