Usable data for AI starts with a specific task, not a purchase order. An enterprise team can buy a dataset, subscribe to a data product, accept a data grant, or approve an internal data request. Once access is approved, a harder decision begins: can this exact data state support the AI task the team intends to run? That decision is also a necessary part of data readiness for AI.
Contracts, permissions, and delivery records can establish who may reach the data and which revision is available. AWS Data Exchange, for example, documents how subscribers can access provider-maintained data in Amazon S3 and analyze it with AWS services. Those records are valuable, but they do not show whether the dataset represents the right population, carries the right labels, or fits the selected model and evaluation method.
In August 2026, Google submitted the successful $10 million bid for a deidentified enterprise dataset from Spirit Airlines, according to a bankruptcy-court filing. The proposed sale remained subject to court approval after objections raised privacy and intellectual-property concerns. The bid shows that access to operational data can carry substantial economic value. It still does not establish whether that data fits a particular AI task.
The UK Government Data Quality Framework expresses the general principle as fitness for purpose: the required quality varies by purpose, data is unlikely to be equally fit for every use, and data quality is more than cleaning. For an AI team, “purpose” must be specific enough to test.
Access and usability are different boundaries
A transaction, subscription, or internal approval can provide valuable evidence. It may identify the provider, define the licensed scope, grant read permissions, expose revisions, and establish how the data is delivered. Those facts belong in the review.
They still do not answer whether the dataset represents the population the AI system will encounter, whether its labels express the right target, whether its collection process introduces relevant gaps, or whether it works with the selected model, prompt, retrieval setup, and evaluation method.
Google Research's Data Cards work makes the separation concrete. Its minimum documentation includes upstream sources, collection and annotation methods, training and evaluation methods, intended use, and decisions that may affect model performance. Its reader-centric guidance goes further: dataset benefits are bounded, and a dataset created for one purpose can have clear shortcomings when used for another.
Documentation therefore changes what a buyer or owner can inspect. It does not turn access into a universal approval.
Access evidence
What the transaction or grant establishes.
Target-fit evidence
What must be established for one AI task.
Data readiness for AI starts with a defined target
The question “Is this good data?” is too broad for an AI release decision. A dataset can be complete enough for reporting and still omit the features a prediction task needs. It can be appropriate for evaluation and unsuitable for fine-tuning. It can be legally accessible while requiring additional policy, privacy, or usage review for the intended workflow.
Consider a demand-forecasting team evaluating a well-documented transaction dataset. The records may be complete, licensed, and easy to query. If they cover online purchases while the target includes store demand, however, the dataset may not represent the population the model will encounter. The issue is not whether the data is valuable. The issue is whether this data state fits this forecasting target under the planned conditions.
The Datasheets for Datasets paper proposed documenting a dataset's motivation, composition, collection process, recommended uses, restrictions, assumptions, risks, and maintenance. The aim is to help dataset consumers make informed decisions and select more appropriate datasets for their chosen tasks. The important word is chosen: appropriateness is tied to what the team is trying to do.
NIST's AI Risk Management Framework applies the same discipline at the system level. It calls for teams to define the tasks an AI system will support, document data-selection considerations such as availability, representativeness, and suitability, specify the application scope, and connect measurement to the deployment context. NIST does not use CUBIG's terminology; its guidance independently supports the narrower principle that data decisions require a defined use and context.
What the transaction proves and what it does not
| Evidence area | Access or ownership can establish | AI use still requires |
|---|---|---|
| Identity | Provider, asset, revision, delivery method. | The exact data state bound to the intended run. |
| Rights | Entitlement, license or grant scope, access duration. | Applicability of those rights and policies to the defined workflow. |
| Meaning | Available descriptions, schema, provenance, collection notes. | Whether definitions, labels, population, and limitations match the task. |
| Evidence | Provider tests or documentation, when supplied. | A target-specific evaluation with recorded conditions and limitations. |
| Change | Provider revisions and availability updates. | Requalification triggers for changes to data, model, prompt, tools, policy, or environment. |
The right-hand column is not a reason to reject third-party or shared data. It is the work needed to turn access into a defensible decision. A strong data product may provide much of the required documentation. The consuming team still has to test whether that evidence applies to its target.
From Core Readiness to Target Fit
In CUBIG's current AI-ready data model, this distinction separates two related questions.
Core Readiness asks whether the data can be understood, governed, inspected, and reproduced as a reliable foundation. Provider identity, provenance, definitions, integrity checks, permissions, and known limitations belong here. A well-documented data product can strengthen this foundation.
Target Fit asks whether one data state and AI setup satisfy the requirements of a defined task. The team needs a Target Profile: the use case, success criterion, model or agent version, prompt, retrieval and tool configuration, and relevant deployment or policy conditions. A Proof Run then tests that combination under recorded conditions.
Passing a Proof Run is still not permanent certification. A provider may release a new revision. The target population may shift. A policy, model, prompt, tool, or environment may change. Continuous Operating Evidence preserves those changes and identifies when requalification is required.
How to assess usable data for AI before use
Before moving an entitled dataset into an AI workflow, ask:
- What exactly was acquired? Record the provider, revision, delivery method, license or grant scope, and the data state the team can access.
- What does the dataset represent? Check collection context, population, time range, annotation method, provenance, transformations, and known gaps.
- Which AI task is being evaluated? Lock one Target Profile rather than relying on a generic statement that the data is valuable or high quality.
- What evidence establishes fit? Define the method, baseline, success criterion, limitations, and owner for a target-specific Proof Run.
- What would invalidate the result? Record provider revisions and internal changes that should trigger review or another test.
This review should be proportional to the decision and its risk. It is not a claim that every dataset needs the same paperwork or evaluation method.
Treat access as the start of the decision
The value of a data marketplace, grant, or internal catalog is real: it can make data discoverable and reachable. The mistake is treating that achievement as the final evidence of AI usability.
Start with one acquired or shared data state and one Target Profile. Connect the access record to provenance, permissions, intended use, evaluation conditions, and change history. Then use a Proof Run to decide whether to scale, fix, retest, or stop.
This is where usable data connects to CUBIG's broader view of AI-ready data. Access and organization support Core Readiness. A defined target and test are still required to establish Target Fit, and operating evidence keeps the decision current.
Test one entitled data state against one defined AI task. See how Syntitan supports AI-ready data decisions.
