AI-Ready Data, Syntitan

Data Cleaning Doesn’t Vanish From Your AI Budget. It Turns Into Headcount.

Hello, this is CUBIG the company behind Syntitan, the AI-ready data platform for enterprise AI. 💎

In its Enterprise Data Infrastructure Benchmark Report 2026, Fivetran found that 53% of data engineering time now goes to maintaining pipelines that already exist, not building new ones. A separate Fivetran survey of 401 data leaders found 67% of centralized enterprises put more than 80 percent of their engineering resources into pipeline maintenance alone. That time has a name on the org chart. It’s headcount, and unlike a software cost, it doesn’t get cheaper the tenth time you pay it.

Minimal geometric illustration of a wide row of data blocks narrowing into a tall stack of smaller data blocks

The maintenance tax nobody puts in the pitch deck

Every AI vendor pitch describes what a model will do once it’s running. Almost none describe what it takes to keep it fed. Fivetran’s 2026 benchmark, based on a survey of 500 senior data and technology leaders, put a number on the gap: 53% of data engineering time goes to maintaining pipelines that already exist. Its earlier 2025 survey of 401 data leaders found the picture was worse at scale — 67% of centralized enterprises allocate more than 80 percent of engineering resources to pipeline maintenance, leaving a sliver for anything new.

That maintenance work is mostly the same task repeated: reconciling a format that drifted, chasing a schema that changed upstream, rebuilding a mapping a client’s system quietly broke last quarter. It is billable, necessary, and structurally incapable of getting faster on its own. A data engineer hired to build AI pipelines this year will spend most of next year re-cleaning the same categories of mess a different client produces. One practitioner on Hacker News put the ratio bluntly: “an hour to build and about 3-5 weeks to handle bureaucracy, data cleaning and detective work to account for the lack of docs.” A second Hacker News commenter described the underlying labor as “soulcrushing” — “huge amounts of soulcrushing human labor for data acquisition, cleaning, labeling.”

An hour to build, three to five weeks to clean

A composite scene, drawn from the pattern above, makes the ratio concrete. An insurance client hands a delivery team data from five regional branches for a claims-automation pilot. Each branch exported it differently: two date formats, three currency conventions, diagnosis codes that mean different things depending on which system entered them. Three junior engineers spend two weeks just aligning the date fields, and they still haven’t finished when the sprint ends. Thirty-five percent of the engagement’s budget is gone before a model has touched real data.

The client’s question in the steering committee is reasonable and unanswerable in the moment: it’s an AI project, so why isn’t there a model yet? They paid for AI. What they are getting, for the first third of the engagement, is data cleaning performed by people whose hourly rate assumes they’d be doing something else. The mismatch isn’t a staffing failure. It’s what happens when the only lever available to absorb inconsistent data is more people looking at it by hand.

Minimal geometric illustration of five mismatched data-block cards converging into a funnel shape

Warehouses store what arrives. They don’t make it agree with itself.

This keeps happening for a structural reason, not a discipline problem. A data warehouse enforces schema: it will insist a field exists and has a type. Schema is not the same thing as consistency. Ten clients sending data in ten formats will all land successfully in the warehouse — inconsistencies and all — because storing a value faithfully was never the same job as making ten versions of that value agree with each other.

Infrastructure’s job is to keep what comes in. Making it consistent has never been anyone’s job by default, which is exactly why it keeps landing on whichever engineer is available that week. Between storage and the consistent signal an AI system can actually trust, there’s a layer that most data stacks simply don’t have. Without it, cleaning cost has nowhere to go but headcount, and headcount has nowhere to go but the project’s margin.

Hiring doesn’t fix a problem that resets at every client

The standard response is to staff around it: add data engineers until the backlog clears. It works for exactly one engagement. The eleventh client’s data is inconsistent in its own way, so the eleventh onboarding costs close to what the first one did, and the team that got faster on client one starts over on client eleven. Headcount added to absorb inconsistency doesn’t compound the way a reusable pipeline does — it resets.

IBM’s 2025 CDO Study, run with Oxford Economics across 1,700 senior data leaders, found more than a quarter of organizations estimate they lose upward of $5 million a year to poor data quality, and 7% put the figure above $25 million. Those losses accumulate quietly, mostly as labor cost nobody labeled “data cleaning” on the invoice. Meanwhile Gartner told its Data & Analytics Summit audience in 2026 that the share of AI spending going to data readiness will grow sevenfold between 2025 and 2029. The market isn’t spending less on this problem as it scales. It’s spending dramatically more, on a fix that mostly means more people doing the same manual reconciliation, one client at a time.

Minimal geometric illustration of an upward-trending bar chart where each bar is built from stacked data blocks

What actually stops compounding as headcount

The alternative isn’t automating the cleaning step faster. It’s reducing how much cleaning a given dataset needs in the first place, by enforcing consistency structurally instead of relearning it by hand at every client. That’s the question underneath CUBIG’s Consistency axis: once the noise is filtered out, is what’s left a signal an AI system can actually trust — not “was this stored,” but “does this agree with itself”?

Most data quality tools on the market today — Great Expectations, Soda, and Monte Carlo among them — are built to detect when something’s wrong after the fact. That’s valuable, and it’s a different job: flagging inconsistency isn’t the same as removing the need for a person to fix it every time it recurs. Syntitan is the layer that sits between data management and AI execution and does the second job — bringing a dataset into a consistent, AI-ready state structurally, so the same category of drift doesn’t have to be rediscovered and hand-fixed at every new engagement. Consistency captured once becomes something the data keeps, rather than a cost the next client resets to zero.

Comparison of approaches to reducing AI data cleaning cost, from manual fixes to structural consistency
Approach Reduces need for repeat cleaning Catches inconsistency before it reaches AI Gets cheaper by the 11th client
Manual data cleaning by hand No Partial No
Hire more data engineers No Partial No
Syntitan (structural consistency) Yes Yes Yes

Before your next data-heavy AI engagement

Five questions for anyone scoping AI delivery work across multiple clients or business units. Each “no” is cost that will land on headcount instead of the statement of work.

  • Does the proposal separate “data preparation” from “data cleaning,” or is cleaning quietly absorbed into a generic engineering line?
  • If this engagement is your eleventh client, will onboarding cost less than the first one did — and can you point to why?
  • When a client’s schema drifts three months in, does anything remember what the last drift looked like, or does someone start from zero?
  • Is the team measuring engineering hours spent building versus hours spent reconciling formats that already existed?
  • If headcount on this account doubled, would the backlog actually shrink, or would it just process the same ratio of new mess faster?

See how much of your data is already consistent enough for AI to trust. Try it on your data, free, and see what a structural consistency check finds before you staff up to find it by hand.

Syntitan, the AI-ready data platform. Try it on your data, free.

References

  1. Fivetran, "The Enterprise Data Infrastructure Benchmark Report 2026" (2026)
  2. Fivetran, "AI & Data Readiness" Report (2025)
  3. IBM Institute for Business Value with Oxford Economics, "2025 CDO Study: The AI Multiplier Effect" (2025)
  4. Gartner, "Data & Analytics Summit 2026 Sydney: Day 2 Highlights" (2026)
  5. Hacker News, comment by jsheard (2025)
  6. Hacker News, comment by pydry (2026)

FAQ

What does it mean that data cleaning cost "turns into headcount"?

Data cleaning work doesn't disappear when it isn't budgeted for explicitly — it gets absorbed by engineers who were hired for other work. Fivetran's 2026 benchmark found 53% of data engineering time already goes to maintaining existing pipelines rather than building new ones, and that time shows up as headcount cost rather than a line item labeled "cleaning."

Why doesn't hiring more data engineers fix the problem?

Each new client's data is inconsistent in its own specific way, so onboarding the eleventh client costs close to what the first one did. Headcount added to handle inconsistency resets with every new engagement instead of compounding the way a reusable pipeline or codebase does.

What's the difference between data consistency and data cleaning?

Cleaning is the manual, repeated work of fixing a specific dataset's problems by hand. Consistency is a structural property: once enforced, a dataset doesn't reintroduce the same category of noise at the next client or the next run, which reduces how much cleaning is needed in the first place rather than doing the cleaning faster.

How is this different from data quality or observability tools?

Tools like Great Expectations, Soda, and Monte Carlo are built to detect when data has gone wrong, which is valuable but leaves the fix to a person each time. Syntitan brings data into a consistent, AI-ready state structurally, so the same category of drift doesn't have to be rediscovered and hand-fixed at every new engagement.