AI-Ready Data

Same Customer Data, Two AI Tasks, Two Different Preparations

Two translucent pattern shapes overlapping on one shared base sheet, in purple and blue gradients

The following scene is an illustration, not a customer case. A B2B software company spends a quarter tidying its customer data. Account records from different systems are linked, broken account IDs are repaired, and every timestamp is moved to one time zone. Then two teams ask to use it. The data science team wants a model that predicts which accounts will cancel. The support team wants an AI agent that drafts replies to customer tickets. Both expect the data to be ready, and both find that it is not, for different reasons.

AI data preparation starts from the task, not the dataset. The shared work was worth doing, and much of it carries over. But what counts as one record, which period of data the AI may see, what each field means, what to leave out, and how to check the result are decisions each task has to confirm for itself. Some earlier answers can be reused. Whether they still fit the new task has to be checked rather than assumed.

Key takeaways

  • AI data preparation combines a shared base that many tasks can reuse with task conditions that each AI task has to confirm.
  • The same customer data needs a different record unit, time cutoff, field meaning, exclusions, and evaluation set for a churn model than for a support agent.
  • Identifiers, formats, record links, and some preparation rules can carry over to a new task. Its time cutoff, inputs, and evaluation conditions need to be confirmed again.

What AI data preparation covers: a shared base and task conditions

AI data preparation is the work of getting enterprise data into a state a specific AI task can use. Part of it builds a shared base: identifiers that link records across systems, consistent formats and time zones, and a record of where each source came from. Many tasks can use that base, and the work can start before anyone has chosen a use case. It still needs care. Records that look like duplicates can be repeat events, and old plan codes can carry different contract terms, so the base should keep original values and a history of changes rather than overwrite them. The other part is set by the task: its record unit, time cutoff, field meanings, exclusions, relationships, and evaluation. That part cannot be settled until the task and its success criteria exist.

Same data, two AI tasks: where data preparation splits

Data preparation for the two tasks in the scene splits on six decisions. The table puts them side by side, starting from the same shared customer data.

Six data preparation decisions for the same customer data used by two AI tasks (illustration)
DecisionChurn prediction modelSupport reply agent
What one record isOne account at one point in time, with its usage and tickets rolled upOne ticket thread, linked to the account it belongs to
Which time the data may showOnly what was known before the prediction date; later usage, tickets, and status changes are cutCurrent terms for current questions; past plan terms and expired exceptions stay, marked with their dates and scope
What a field meansPlan code as a category, with old and new codes merged where they mean the same planPlan code translated into the plan name and terms a customer would recognize
What to leave outAccounts still in a free trial, if the task is about paying customersInternal notes and fields the agent should not repeat to a customer
Which relationships to keepTicket counts and severity per account per monthEach ticket’s link to the contract, earlier replies, and any exception approved for this customer
How to check the resultHold out a later time period and compare predictions with what happenedReference replies reviewed for current policy, or past cases replayed with the information available then, scored the same way before and after a data change

For the agent, the team has to say which kind of evaluation it is running. Replaying past tickets with only the information available at the time tests whether the agent handles a past situation correctly. Testing current answers needs reference replies that someone has reviewed and updated against today’s policy. Scoring a current answer against an old reply can mark a correct answer as wrong.

Neither column is the better preparation. They answer different questions. Problems start when one team reuses the other’s table without checking. If the data science team trains on the support team’s current-state view, the model sees each account’s current plan status, and an account that has already cancelled shows up as cancelled. That is data leakage, which the scikit-learn documentation describes as information that would not be available at prediction time. An evaluation built from the same view shares the leak and makes the model look better than it will be in use. Separating the evaluation period helps only if the team also checks that each record’s inputs were actually known at prediction time. If the support team uses the churn table, the agent gets monthly account summaries with no ticket text, contract terms, or approved exceptions to work from.

The support column also shows why exceptions need their own handling. A one-time concession approved for one customer reads like a standing rule once it lands in a summary without its dates and scope, which we showed in an expired exception that an AI summary kept alive.

What carries over to the next AI task, and what to check again

A new AI task can reuse the shared base and parts of earlier task work, but it has to confirm that they fit its own time cutoff, inputs, and evaluation conditions. Suppose the company adds a third task next quarter, a model that suggests which accounts to offer an upgrade. The links between tickets and contracts carry over, and so can the rule for merging old and new codes for the same plan. Account features built for churn may carry over too, but only after they are rebuilt as of the upgrade model’s own prediction date, and the evaluation set has to match what the upgrade model will be judged on.

Change creates the same kind of work. When the shared base gets a new source or a corrected field, the shared rules are applied again and each task that depends on the base needs its checks repeated. Without that, teams cannot tell whether a change in results came from the model or the data. Keeping track of what each task reused, what it changed, and what was checked is the part that grows with every new AI task. It is also the part CUBIG builds into Syntitan, so teams do not have to rebuild that record in each project.

The same split between what carries over and what to recheck applies when one AI product goes to a new customer, as we covered in what a second AI deployment should inherit.

How to prepare data for an AI task

  1. Define the task and its output first. Write what the AI will produce, for whom, and how a good output will be judged, before deciding anything about the data.
  2. Build the shared base without losing the originals. Align identifiers, formats, and sources as shared rules, keep original values and change history, and apply the rules again when the data changes.
  3. Decide what to reuse and what to confirm. For record unit, time cutoff, field meaning, exclusions, relationships, and evaluation, check whether an earlier task’s choice still fits, and record the choice and the reason next to the prepared data.
  4. Build the evaluation set with the task. For a prediction model, hold out a later time period and check that every input was known at prediction time. For an agent, decide whether you are replaying past cases with the information available then or testing current answers against replies reviewed for current policy.
  5. Compare before and after under fixed conditions. A prediction model is usually retrained when its data changes, so fix the model type, training settings, evaluation data, and metrics. For an agent, fix the model version, prompt, and retrieval settings, and change only the data it draws on.
  6. Recheck dependent tasks when the shared base changes. List which tasks use the base, and repeat their checks after a change instead of assuming they still hold.

Whether data is AI-ready is a question about a task, and these steps are how a team answers it for one. For a structured look at where your data stands before choosing tasks, see the AI readiness assessment.

Where Syntitan fits in AI data preparation

The teams that own a task decide what the task is, what counts as a good result, and the six decisions above. Syntitan, CUBIG’s AI-Ready Data Platform, runs and records the data work around those decisions. Assess Data Readiness and Improve Data Readiness work on the shared base and can run before a use case exists. Define Use Case and Prepare Data for Use Case hold the task conditions: the customer’s structure, field meanings and units, business terms, relationships, and protection conditions for that task. Teams then validate data fit with a supported model or agent under fixed evaluation conditions, and review a result of qualified, not qualified, or inconclusive for that task. Which transformations, models, agents, and integrations apply depends on the setup, and teams use the parts they need rather than every step in order.

What the team gets is data versions linked to their task conditions and validation evidence: each carries its preparation steps, before-and-after evidence, the evaluation conditions, and the validation result. Because those steps and that evidence stay with each version, the next task can start from work that already exists and spend its time on the conditions that need checking. When the shared base changes, the team can see what changed against the previous version and validate again each task version that uses it.

Pick one AI task your team has already planned and write down the six decisions for it. Mark whether each one came from earlier work or was decided new, and keep the evidence that it fits this task. A decision with no such evidence has not really been made yet. To see how the data side can work, explore Syntitan.

Explore Syntitan, CUBIG’s AI-Ready Data Platform

FAQ

What is AI data preparation?

AI data preparation is the work of getting enterprise data into a state a specific AI task can use. It combines a shared base, such as linked identifiers, consistent formats, and source records, with task conditions, such as the record unit, time cutoff, field meanings, exclusions, and evaluation set for that task.

How is AI data preparation different from data cleaning?

Fixing duplicates and broken values belongs to the shared base and needs care, since apparent duplicates can be real repeat events. AI data preparation also covers decisions that depend on the task, like which period of data a model may see and how its output will be checked. Cleaning does not settle those decisions.

What are the steps to prepare data for an AI task?

Define the task and its output, build a shared base that keeps original values, decide which earlier choices to reuse and which to confirm, build an evaluation set that matches the task, compare before and after under fixed conditions, and recheck dependent tasks when the shared base changes.

Can the same prepared dataset be reused for multiple AI tasks?

The shared base and some preparation rules can be reused. Whether they fit a new task's time cutoff, inputs, and evaluation conditions has to be checked again. Reusing one task's prepared data for another without that check can cause problems such as data leakage or missing context.

What is data leakage in AI data preparation?

Data leakage happens when information that would not be available at prediction time is used to build a model. Training a churn model on current account status is a common example, because the status already shows who cancelled. An evaluation built from the same data can hide the problem. Separate the evaluation period and confirm that each record's inputs were actually known at prediction time.

How do you know whether data is prepared well enough for an AI task?

Compare results before and after the data change under fixed conditions and check them against the criteria set for the task. For a retrained model, fix the model type, training settings, evaluation data, and metrics. For an agent, fix the model version, prompt, and retrieval settings. If the evidence is too thin, treat the result as inconclusive.