# CUBIG: Full Documentation > Full text of CUBIG pages, concatenated for LLM ingestion. > Index (curated summary): https://cubig.ai/llms.txt @generated: products/company static (2026-07-09); articles + pillar glossaries appended live @note: The product/company/proof pages below are static and updated on major redesigns. Articles and pillar-flagged glossaries are appended dynamically from live content at request time. Excludes blogs (general education), news, /ko translations, and retired pages. --- title: "CUBIG | AI-ready execution for enterprise AI" url: "https://cubig.ai/" source: live --- Skip to content Enterprise AI fails on data, not models. AI-ready isn't a project. It's a product. Syntitan fills the missing layer between enterprise data management and real AI execution. Sign in and run it on your own data. Explore Syntitan Run a sample proof Free trial on sign-up. No sales call, no PoC. Diagnose Refine Release State Proof Run Agents See how ready your data is, scored on six checks. AI-ready data is usable, reliable, and stable in production. CUBIG gets enterprise data there. Past blocked approvals, past messy records, past results that change every run. Gartner® “Emerging Tech: AI Vendor Race: Tech Innovators in Agentic AI — Solution Accelerators” (2026) Gartner® “Emerging Tech: AI Vendor Race: Most Prominent Use Cases in Agentic AI by Industry” (2026) Gartner® “Emerging Tech: Provider Differentiation Strategy—Trends for Hyper-Synthetic Data” (2025) Named a Representative Vendor Problem AI is stuck between data management and AI execution. Storage works. Models work. It's the data in between that stalls projects. no access no access no access Restricted Data your teams can't touch. Approvals take months while workflows stay blocked. null null null null Unusable Records too messy or thin to train on. Missing values and imbalance stop the model before it starts. PoC run · 0.86 Production run · Unstable States that shift between PoC and production. Last month's working run can't be reproduced today. 30% of GenAI projects are abandoned after PoC. Only 4% of IT leaders call their data AI-ready. — Gartner (2024, 2025) Missing Layer Your systems manage data. Syntitan makes it run for AI. It sits on top of Snowflake, Databricks, and Fabric, and fills the layer they leave open. Nothing gets replaced. WHERE AI RUNS AI Workloads Models, LLMs, RAG, and agentic workflows that need data they can trust and trace. Fraud detection Customer analytics Enterprise copilots AI agents Risk simulation Reproducible runs Powered by Syntitan · DTS · LLM Capsule CUBIG AI-Ready Execution The operating layer between data management and AI execution. One platform, Syntitan , holding two core capabilities. Syntitan Platform AI-ready execution Release State Run Binding Diff Reproduce DTS Capability AI-Ready data transformation engine Diagnose Transform Rebuild LLM Capsule Capability Context-preserving data layer for AI Substitute Execute Reconstruct WHERE DATA LIVES Enterprise Data Storage, query, pipelines, BI, and data lakehouse, built for management, not AI execution. Databases SQL · NoSQL Documents Contracts · Internal CRM & ERP Salesforce · SAP Object Storage S3 · Data Lake Logs & IoT Sensors · Streams APIs & Legacy REST · SOAP Every source, none of it AI-Ready Two Entry Paths Start with data, or start with a workflow. Whether your blocker is data or workflow, start with the path that fits your team. One platform. Two entry paths. Path A · AI-ready data For data leaders, ML teams, and AI platform owners Make your enterprise data ready for AI, and keep it that way For data that's locked, scarce, or unstable on AI. Rebuilt so your models can run on it. Path B · Sensitive AI workflow For AI adoption leads and workflow owners Run LLM, RAG, and agent workflows on sensitive data For workflows where sensitive data blocks LLM execution. Enabled without exposing raw data. Platform One platform. Five capabilities. From diagnosis to proof, everything runs in one place: Syntitan. Platform Syntitan The AI-Ready Data Platform. It makes your data ready for AI, keeps it that way, and proves the difference on your own workflow. Explore Syntitan The same six steps you'll see in the product 1 Diagnose AI readiness diagnosis on six axes 2 Refine Fix data values and context 3 Optimize Tune for your metric and model 4 Release Freeze a versioned data state 5 Proof Run Before/after on a baseline model 6 Verify Re-run in your environment Six axes Usability Integrity Context Consistency Reproducibility Traceability Capability DTS AI-ready data transformation engine. Rebuilds restricted, thin, and imbalanced data into datasets your models can learn from. Explore DTS Capability LLM Capsule Context-Preserving Data Layer for AI. Run LLM and agent workflows while original values stay inside your environment. Explore LLM Capsule Validation Quality gates before a data state is released Operating Control Release State · Run Binding · Diff · Reproduce Agent Connection Connects agents and workflows to enterprise systems Inside Syntitan The blocker you'd tackle first. Three walls teams hit between data and AI. Chances are, yours is one of them. Gaps and errors in your data blocking AI? From imbalance correction to synthetic augmentation, rebuild a complete AI-ready dataset. Need to use data without exposing the originals? Original values stay inside. Usable results come back. Want to test customer and market responses first? Simulate synthetic personas instead of recruiting panels. Research in hours, not weeks. All of it happens in one place Syntitan . Use cases Same blocker. Different industries. One platform. Financial services, healthcare, public sector, telecom/NOC, manufacturing/OT. The data state blocking production AI looks the same everywhere. CUBIG removes it. FINANCIAL SERVICES Fraud Detection & AML Analytics HEALTHCARE Clinical Decision Support & Research PUBLIC SECTOR Policy Sentiment & Citizen Services TELECOM / NOC Network Anomaly Detection & NOC Automation MANUFACTURING / OT Predictive Maintenance & Quality Inspection DEFENSE Defense AI Operations & Threat Analytics “Improved anomaly detection reliability and audit-traceable model runs across rare fraud and AML patterns.” OUTCOME “DTS expands rare fraud scenarios with synthetic data · Syntitan binds runs via Release State for audit.” CUBIG REMOVES IT THE BLOCKER Rare fraud and AML patterns are underrepresented in training data. Compliance audits cannot trace which data version produced which decision. EXAMPLE DATASET transaction_id account_id amount merchant_id mcc_code timestamp location is_fraud “Clinical insights and research models generated without exposing PHI, even for rare disease cohorts.” OUTCOME “DTS rebuilds rare clinical cohorts with structure-preserving synthetic data · LLM Capsule runs clinical LLM workflows without raw PHI exposure.” CUBIG REMOVES IT THE BLOCKER PHI restrictions prevent patient data from reaching modern LLM and ML pipelines. Rare disease cohorts are too small for reliable model training. EXAMPLE DATASET patient_id encounter_id diagnosis_code lab_result medication timestamp age_group region “Early detection of policy sentiment shifts and faster citizen-service responses, under Korea's AI-ready public data guidelines.” OUTCOME “DTS transforms regulated public data into AI-ready states · LLM Capsule enables citizen-service LLMs without raw record exposure · Syntitan Release State for audit.” CUBIG REMOVES IT THE BLOCKER Citizen records and policy data are siloed across agencies and regulated by privacy law. LLM-based services cannot consume raw policy data directly. EXAMPLE DATASET case_id agency topic sentiment_score region citizen_age_band timestamp resolution_status “Stable anomaly detection through pipeline updates, with subscriber data never leaving the operator's environment.” OUTCOME “DTS expands rare anomaly scenarios · LLM Capsule keeps NOC LLM workflows on-prem · Syntitan Run Binding stabilizes monitoring through pipeline changes.” CUBIG REMOVES IT THE BLOCKER Subscriber PII and network topology can't be moved to external AI environments. Rare network anomalies are sparse in training data and drift after pipeline updates. EXAMPLE DATASET subscriber_id cell_id traffic_volume packet_loss latency_ms timestamp region alert_level “Higher predictive maintenance accuracy and shorter downtime, without exposing process IP or breaking OT isolation.” OUTCOME “DTS synthesizes rare defect scenarios with structure preservation · LLM Capsule runs OT-side inference without raw telemetry exposure.” CUBIG REMOVES IT THE BLOCKER Process IP and OT telemetry cannot leave the plant for cloud AI training. Defect cases are rare, making quality-inspection models unreliable. EXAMPLE DATASET machine_id sensor_type vibration_rms temp_c pressure_bar timestamp defect_label line_id “AI-assisted operational analysis under air-gapped constraints, without weakening classification or network isolation.” OUTCOME “DTS synthesizes operational scenarios with classification-preserving transforms · LLM Capsule runs analyst LLM workflows entirely inside the air-gapped network.” CUBIG REMOVES IT THE BLOCKER Operational data is classified and cannot leave air-gapped environments. Threat scenarios are rare and AI models can't be trained on enough variation. EXAMPLE DATASET mission_id asset_type region_code threat_level sensor_feed timestamp classification_tier response_action Proof Built for enterprise. Proven in production. Audit trail, data lineage, cloud or on-prem deployment. And you can try all of it without a sales call. More → Key Numbers Customers & partners 15+ Across finance, healthcare, public sector, legal, marketing, and cloud Awards & certifications 10+ 4 Ministerial Prizes · GS · KISA Patents 12 8 domestic (4 registered) · 4 overseas (1 allowed) Founded 2021 Seongnam-si, Korea · UK entity established Certifications & Recognition Data Safety Controls Access control, audit logging, and separation of duties built into the operational workflow. Audit & Traceability Run Binding, Release State, and Diff give full traceability of data lineage, transformations, and AI execution states. Compliance-Ready Designed to operate within regulated industries. Enterprise-grade controls applied throughout. Enterprise Procurement Available via enterprise marketplace channels with procurement support from first contact. Deployment Options On-premises, cloud, or marketplace deployment. Flexible to fit your existing infrastructure and security posture. Policy-based data boundary control Policy-based handling of raw data boundaries and data minimization across all workflows. Learn Understand AI-ready data before you buy anything. Definitions, comparisons, and assessments, written to be useful even if you never sign up. Start here What Is AI-Ready Data? (And Why Clean Data Isn’t Enough) Clean data isn't AI-ready data. The pillar guide to what "ready" actually means, and how to tell where your data stands. Read the guide → Go deeper GLOSSARY AI-Ready Data → ARTICLE AI Readiness Assessment: The Six Readiness Axes → ARTICLE AI-Ready Data vs Clean Data: Why Clean Isn’t Enough → LATEST Ten Clients, Ten Different Failures: When AI Doesn’t Understand Domain Context → LATEST 80% of AI Is the Dirty Work of Data Engineering. Nobody Budgets for It. → LATEST The Real Bottleneck in Enterprise AI Adoption Is Not What the 95% Report Says → Explore the Learn Hub → See the difference on your own data. Explore Syntitan Run a sample proof Free trial on sign-up. No sales call, no PoC. Prefer to talk it through? Book architecture review → --- title: "Syntitan: AI-Ready Data Platform for Enterprise AI | CUBIG" url: "https://cubig.ai/syntitan" source: live --- Skip to content Syntitan · AI-Ready Data Platform Stop hoping it works in production. Prove it on your own data. The model isn’t the problem. Your data is. Syntitan gets your data AI-ready, then proves the difference before production. See how it works Run a sample proof Free trial on sign-up. No sales call, no PoC. What one proof run shows Raw data AI-ready 89% Recall @ fixed precision Flag for review Caught at decision time, before the loss reaches production. Same agent. Same model. Same test. Only the data changed. Representative; reproduce it on your own data and model. Old vs new Stop getting AI-ready. Just be AI-ready. No consultants, no readiness workshops, no guesswork. Sign up and run it. The old way Consultants define “AI-ready” A six-month readiness project PoC after PoC Agents stall on the data “We hope it works” With Syntitan Sign up, no sales call A six-axis readiness score in 30 seconds A working product, not a PoC Agents get sharper on the same data Reproduce the result yourself, exactly The problem Most AI works in the PoC, then breaks in production. 46% of AI PoCs scrapped before production 60% of AI projects stall on data S&P Global, Gartner (2025) Verifiable Data State Your AI doesn’t break on the model. It breaks on the data state. A run that worked yesterday breaks today, and no one can point to what changed. Syntitan controls the data state behind every AI run: the exact version of the data each result came from. Then it ties that state to model performance you verify yourself. Release State Freeze the exact data state an AI used. “What data did we train on in March?” gets a one-click answer for audit. Run Binding Every AI run is bound to a data state, so “this result came from this exact data” is traceable automatically. Diff Compare yesterday’s and today’s data. When output drifts, see in seconds whether the data changed. Reproduce Restore any past state and re-run it. “Worked in March, broke in April” becomes a reproducible investigation, not a guess. AI-ready AI-ready isn’t one thing. It’s six measurable axes. Six checks that tell you whether AI can actually run on your data, and what to fix first. Usability Can sensitive data be used with AI safely? representativeness · enrichment · quantification Integrity Are missing values, duplicates, and skew visible? data quality · bias · accuracy Context Does AI know what each field means? semantics · inference & derivation Consistency Which fields help or hurt the AI task? consistency assessment · uniform records Reproducibility Can this data state be reused in a workflow? versioning · validation · regression testing Traceability Can changes, versions, and authors be tracked? lineage · stewardship · compliance Gartner, AI-Ready Data Essentials Roadmap (2024) · Map Your AI Use Cases by Opportunity (2025) Product From data diagnosis to verified proof. One workspace. Walk through the real Syntitan product, from data diagnosis to model verification. Pick any step to open its workspace. AI Performance Workbench Select a step to open its workspace. 1 Diagnose AI readiness diagnosis 2 Refine Fix data values and context 3 Optimize Tune for your metric and model 4 Release Freeze a versioned data state 5 Proof Run Before/after on a baseline model 6 Verify Re-run in your environment AI Readiness Diagnosis Diagnose whether your data is ready for AI across six axes, and pinpoint what to fix first. Overall AI Readiness 61% Caution Usability 80% Integrity 34% Context 30% Consistency 68% Reproducibility 75% Traceability 80% AI analysis results AI readiness is partial. Integrity (34%) and Context (30%) are critically low. Fix these before training. The other four axes pass. Proof Run Run the before/after data through a baseline model to preview the expected impact of AI-ready prep. Preview Setup Validation Model XGBoost baseline Task Type Binary classification Target Metric Recall @ fixed precision Data State Snapshot vs AI-Ready Release Simulation Result Before · Raw data Approve transaction Recall 61% F1 0.64 False positive 18% After · AI-ready Flag for review Recall 89% F1 0.78 False positive 11% Expected uplift Recall +28 pts F1 +0.14 False positive -31% No model yet? Get an instant read with a standard baseline model. Figures shown are baseline estimates. Verify the real impact with your own model. Optimize for My Model Adjust the AI-ready prep priorities to your model’s target metric and evaluation data. Upload eval dataset Connect an evaluation dataset that includes ground-truth labels. Connected Upload prediction result Upload the prediction output from your current production model. Upload Connect model via MCP Connect model execution from Claude Code or your own environment. Connect MCP ✓ Auto-detected Fraud detection · Target metric: Recall @ 95% precision · Current score: 73% Optimization Plan Handle PII columns Substitute sensitive columns to stabilize model inputs. +6 pts Remove duplicate records Remove duplicates that bias training. +3 pts Standardize schema context Normalize column semantics and link them to task intent. +4 pts Drop noisy columns Drop unused columns to cut token cost and noise. +2 pts Not generic cleanup. We prioritize the prep that moves your model’s performance the most, in order of impact. Uplift figures are baseline estimates; your validated lift is reproduced on your own model. AI-Ready Refinement Fix data values and distribution, and add the context AI needs to use the data. Feature derivation & augmentation Creates new columns through binning and categorization, and adds composite signals via cross-column operations. Sensitive data detection & substitution Replaces sensitive values with non-identifiable ones. Missing value treatment Preserves meaningful missingness patterns as signal and fills the remaining gaps with statistical methods. Outlier, distribution & category refinement Detects outliers and corrects distribution and category skew so models train reliably. Data augmentation & class balancing Generates synthetic samples for minority classes and rebalances class ratios to a normal range. Low-signal column removal Selectively removes low-importance columns and those that could contaminate predictions. Release & Run Binding Freeze the AI-ready result as a versioned Release State. Every AI run binds to it. New Release This will be published as v4 . Release notes (Optional) Publish release Version History Jul 9, 1:50 AM Current v3 Sensitive fields handled. Column meaning added for product_category. Distribution stable vs v2. data-ops · e00566f Jul 6, 11:18 AM v2 Class balancing applied · low-signal columns removed data-ops · afb0844 Jul 6, 10:16 AM v1 Initial upload data-ops · 2e706ec v3 released successfully. Verify via Claude MCP Don’t just take Syntitan’s word for it. Re-run it yourself under identical model, seed, and split, in your own environment. Portable Proof Kit before_snapshot.csv ready after_release.csv ready change_manifest.json ready comparison_harness.py ready eval_config.json exportable README.md exportable Claude MCP Verification ready Run the before/after data in Claude Code and reproduce the result. Open Claude Code Not just a data download. Export a reproducible experiment kit bundling the Change Manifest, Harness, and Eval Config. Proof The proof is in the product. Three screens that show why the AI that passed your test keeps working in production. Readiness diagnosis The data gaps your PoC never showed, scored on six axes and caught before they reach production. Release & Run Binding The exact data state your model ran on, reproducible in one click. Proof that production won’t break without warning. Proof run Same model, same test, only the data changed. Run it yourself and see the lift before production. Three ways to run a proof. Pick whichever fits how you work: bring your own model, tune the prep to your metric, or re-run it yourself. Preview with a baseline No model of your own? Syntitan’s baseline model compares before and after performance instantly. Optimize for your model Syntitan tailors the data-prep direction to your target metric, eval set, and model. Verify via Claude MCP Export the before/after, change manifest, and harness to re-run in your own environment. Platform DTS and LLM Capsule live inside Syntitan. Syntitan is the platform; DTS and LLM Capsule are its core capabilities. You adopt one platform for AI-ready execution, not three tools. Platform Syntitan Diagnose data readiness, release a fixed AI-ready state, and bind every run to it, so what worked in the demo holds in production. Capability DTS Rebuild restricted, imbalanced, or blocked data into AI-ready datasets, so the data you couldn’t use becomes data your model can. Learn more → Capability LLM Capsule Run AI on sensitive data with its context preserved, then restore the results inside your own workflow. Learn more → How the platform puts them to work Validation Score whether the data DTS and Capsule produce is AI-ready on six axes, and preview the lift on your model before production. Verifiable Data State Lock that data state to a version so every run reproduces exactly. Release · Run Binding · Diff · Reproduce. Agent connection Hand the verified data state to every agent via Claude MCP, so outputs stay bound to data you can trust. Agent Analysis Modules Agents that run on ready data, not guesswork. Generic agents guess. Syntitan agents run on verified data that carries its own context. Persona Survey 100K synthetic personas that reproduce real response distributions Generate survey responses from synthetic personas and analyze behavior patterns. Result report Strategy Which customers to focus on, simulated down to ROI Segment customers by behavior and demographics, and simulate ROI-based strategy. Result report Churn Prediction Catch churn signals before customers leave Identify at-risk customers from behavioral signals, predict churn, and suggest retention strategy. Result report Launch Pricing Launch-price simulation that pre-validates revenue and churn impact Analyze price sensitivity and recommend a launch price that accounts for revenue and churn impact. Result report Model Integration Already have a model or agent? Through API connection, we compare performance before and after AI-Ready. Book architecture review Custom Build Can't find the agent you need? Tell us your analysis scenario and we'll design a custom agent. Book architecture review Fit Where Syntitan fits in your AI data stack Syntitan does not replace the tools your team already runs. It fills the missing step between enterprise data and AI execution. Syntitan in one line The path from raw enterprise data to a fixed, traceable, AI-ready state your production AI can run on. Data platforms Store enterprise data: warehouses, lakehouses, pipelines. Syntitan adds The AI-ready layer on top: score the data, fix it, freeze it. Data quality tools Detect issues: null rates, type errors, schema drift. Syntitan adds The verdict that matters: is this data ready for AI, and what is blocking it. Observability tools Detect that something changed in pipelines, models, or systems. Syntitan adds Which data changed: compare two states and bring back the one that worked. Agent tools Run agents and agentic workflows on top of data. Syntitan adds Ready data for agents to run on, so answers are grounded, not guessed. Sensitive-data transformation tools Synthetic transformation and sensitive-data preparation. Syntitan adds One step on a longer path: diagnose, refine, release, then trace every run. Use cases Where your team starts Wherever your team begins, the path lands on the same Release State. IT · Data · MLOps When AI breaks, how do you know if the cause is data or execution? Run Binding, Release State, and Diff narrow the cause from evidence, not memory. Finance · Risk Can you prove which data state produced this result? Every risk analysis is bound to a Release State with version history that internal review can inspect. Marketing · Growth Can you recreate the exact segment used in the last campaign? Release campaign data states and compare before-and-after changes across versions. Research · Strategy How long does it take to reproduce last month's analysis? Recurring analysis stays attached to the same Release State, so reproduction is a click, not a rebuild. HR Do prediction results change every quarter without a clear reason? Prepare sensitive workforce data through a protected path, release the analysis state, then compare quarter to quarter. FAQ Frequently asked questions What is Syntitan? Syntitan is an AI-Ready Data Platform. It scores whether your data is ready for AI, fixes what blocks it, freezes the result as a versioned Release State, and ties every AI or agent run back to the exact data behind it. What is an AI-Ready Data Platform? A platform that prepares enterprise data for AI and keeps it that way: readiness scoring, refinement, restricted-data preparation, and fixed data states that production runs are traced back to. Is your data ready for AI? Most enterprise data is not ready yet. Syntitan scores your data readiness for AI across usability, integrity, context, and traceability, then shows the specific gaps blocking model or agent use before you spend weeks cleaning data. How do you make data AI-ready? In four steps: diagnose readiness on six axes, refine values and context, release a fixed AI-ready state, and bind every model or agent run to it. What is AI Readiness Qualification? AI Readiness Qualification checks whether data can be reliably and traceably used by AI models or agents. It surfaces gaps across Usability, Integrity, Context, Consistency, Reproducibility, and Traceability. What is AI-Ready Refinement? AI-Ready Refinement fixes data values, distribution, and class balance, and adds the context AI systems need to understand what the data means. What is a Release State? A Release State is a fixed AI-ready data state. Once released, it becomes the reference point for analysis, agent runs, and operational review. What is Run Binding? Run Binding connects every AI or agent execution to the Release State used for that run. What is Diff and Reproduce? Diff compares two Release States to narrow down what changed between them. Reproduce restores a previous data state so teams can investigate from evidence. How does Syntitan help with data drift? Syntitan can compare live production data against a released AI-ready baseline and surface which fields and distributions have moved. How does Syntitan help agents use enterprise data? Syntitan agents and agentic workflows run on qualified, data-grounded states with semantic context attached, not on raw files. Their outputs share the same Release State every team uses. How does Syntitan fit with data platforms, data quality tools, and observability tools? Data platforms store and process data. Data quality tools detect issues. Observability tools detect that something changed. Syntitan sits between them and AI execution: it qualifies whether the data is ready, refines data and semantic context, fixes the state as a Release State, and binds every AI or agent run to that state. Other tools describe or detect. Syntitan prepares. Can Syntitan help evaluate a model API? For selected design partners, Syntitan can preview whether internal data has the signal and context required for a target model, before full evals are run. Which data state is your AI running on? Most teams can’t answer. In one upload, Syntitan can. Book architecture review Run a sample proof --- title: "DTS: AI-ready Data Transformation Engine | CUBIG" url: "https://cubig.ai/dts" source: live --- Skip to content DTS · AI-ready data transformation engine Rebuild unusable data into AI-ready datasets. Most enterprise data isn't AI-ready. DTS rebuilds restricted, imbalanced, or incomplete data into an AI-ready dataset you can actually use. It replaces restricted data with privacy-safe substitutes, rebalances skewed datasets by generating additional data, and fills gaps in what the data covers with new AI-ready data. Book architecture review Explore Syntitan Available on AWS Marketplace NCP Marketplace Data problems Three data problems. One engine. Data that can't be shared, can't be used, or can't be accessed. DTS resolves all three. Restricted Data Privacy-safe replacement. Swap compliance-blocked data for a synthetic set with no real personal data. Replace regulated data (GDPR, HIPAA…) with privacy-safe synthetic data Formal ε bound on every output Safe for cross-team, cross-border, external use Unusable Data Coverage & balance expansion. Fix rare classes, imbalance, and thin volume by generating additional data. Augment underrepresented classes at scale Fix class imbalance (too few examples of rare cases) without overfitting Scale small datasets to production volume Non-Accessible Data Safe dataset generation. Generate safe substitutes for data locked in separate systems that can't reach pipelines. Safe replacements from inaccessible sources Unblock stalled validation & testing Keep statistical properties, no data transfer Capability Privacy-safe synthetic data, as a capability. Synthetic data is one capability inside DTS, not its identity. DTS uses it, with differential privacy underneath, to expand coverage and repair imbalance when real data can't be used. Differential Privacy A formal privacy bound, by design. What differential privacy means Differential privacy (DP) is a mathematical framework that bounds how much any single individual's data can influence the synthetic output. Individuals cannot be re-identified, no matter what outside information someone combines it with. DTS applies DP during the generation process itself, not as a post-processing anonymization step. The privacy property is structural, not dependent on masking or field removal, a provable bound, not best-effort masking. The bound The chance of identifying any individual from the synthetic dataset is capped by a defined value, epsilon (ε), regardless of outside knowledge. How DTS generates synthetic data 01 Statistical profiling DTS analyzes the real dataset's statistical properties (distributions, correlations, and other statistical patterns) without storing raw records. 02 DP noise injection Calibrated noise is injected into the statistical model according to DP bounds, so individual data points become mathematically unidentifiable. 03 Synthetic generation New records are sampled from the DP-protected model. Output is statistically representative but contains no real personal information. 04 Fidelity validation Generated data is validated against the original distribution. Quality and utility metrics confirm suitability for training and validation use. Deployment Start with DTS, grow into Syntitan. Mode A · Direct DTS on its own DTS is a core capability of Syntitan you can start with directly, against your own data sources. It fixes AI training-data quality, generating what's missing at scale without touching real data. Fix class imbalance: generate more examples of rare classes with distribution fidelity Augment sparse datasets to production-grade volume Generate edge cases and rare-event samples Mode B · Integrated DTS + Syntitan When compliance blocks data from reaching models, DTS runs inside Syntitan to generate privacy-safe replacements. DTS makes the data. Syntitan versions and tracks it. Replace GDPR, PIPA, HIPAA-restricted data: the original data never leaves your environment Syntitan versions the synthetic dataset and binds it to a Release State Syntitan's change log tracks it from data generation through the AI run In production Finance · IBK Industrial Bank 97.6% fraud-detection accuracy (AI model) · 79 patterns → 1,000 records Fraud and transaction patterns expanded into DP-safe synthetic records. PIPA-compliant, with zero real customer data exported. Finance · Kyobo Life Insurance F1 0.92 churn model · 277,249 synthetic records A 6-month data-retention policy had blocked Kyobo's churn AI. DTS rebuilt DP-safe records from historical data, legally usable after deletion. Marketing / Sales 90% time reduction · 70% cost saving on trend research Annual consumer-trend surveys replaced with AI persona agents trained on synthetic behavioral data. Insights in 1 to 2 days instead of a month. Defense · Ministry of National Defense Zero data exports · classified imagery → AI-ready Deployed on-premise in an air-gapped classified environment. Classified data became AI-ready synthetic datasets within clearance. Comparison DTS vs. other approaches to restricted data. Capability DTS Masking / Anonymization Data Sampling Manual Labeling Privacy bound ✓ Formal DP bound (ε) △ Re-identification risk remains ✗ None ✗ Coverage expansion ✓ Generate at any scale ✗ Can't create new data △ Bounded by real data volume △ Expensive & slow Rare-class augmentation ✓ Targeted generation ✗ ✗ Can't create rare events △ Very high cost Distribution fidelity ✓ Validated against real stats △ Distorted by masking △ Sampling-bias risk △ Annotator variance Cross-border / external use ✓ No real data transferred ✗ Residual risk ✗ ✗ Syntitan integration ✓ Native versioning & binding ✗ ✗ ✗ When to use Five signals your data is blocking AI. Enterprise AI projects stall when data conditions prevent training, validation, or safe deployment. DTS was built for these situations. Restricted Data Data exists but compliance blocks AI access. GDPR, PIPA, HIPAA, or internal retention policies prevent the data from reaching models. Unusable Data Imbalanced datasets or coverage gaps distort model behavior. Rare classes underrepresented, fraud patterns too sparse, edge cases absent from training. Unusable Data Retention policies delete what AI needs. Historical data was deleted per retention policy, so the patterns that trained the previous model no longer exist. Restricted Data Sensitive records can't leave your environment. Classified, patient, or customer data cannot be exported for AI training, even internally. Unusable Data Training-data volume is too low for reliable AI. The original dataset is too small to train a robust model, and collecting more takes months. Outcome In each case, DTS turns data that is restricted or unusable into an AI-ready dataset, without exposing real records. See if DTS fits your data → Proof Proven in production. +30pp F1-Score Lift 58.55% → 88.55% −90% Time to Deploy 4 weeks → 1 day 97.6% Fraud-Detection Accuracy (AI model) IBK Industrial Bank 277K+ Synthetic Records Kyobo Life Insurance Gartner® Representative Vendor AWS Marketplace NCP Marketplace Listed as a Representative Vendor in Gartner®, Emerging Tech: Provider Differentiation Strategy–Trends for Hyper-Synthetic Data (2025). Gartner does not endorse any vendor, product or service depicted in its research publications. GARTNER is a registered trademark of Gartner, Inc. and/or its affiliates. FAQ Frequently asked questions What is DTS? DTS is CUBIG's AI-ready data transformation engine. It generates DP-protected datasets using differential privacy to fix class imbalance, fill coverage gaps, expand training data, and replace restricted or non-accessible data. DTS can be deployed on its own for data transformation work, and operates as a core capability of the Syntitan platform. What is differential privacy in DTS? Differential privacy (DP) is a mathematical framework that puts a hard bound on how much any single person's data can influence the output. This keeps re-identification risk low, no matter what outside information someone combines it with. DTS applies DP during generation, so datasets stay statistically representative while containing no real personal records. Can DTS run without Syntitan? Yes. DTS can be deployed on its own for transformation workloads. As part of Syntitan, its datasets are versioned and bound to Release States. What data problems does DTS solve? Three categories. First, restricted data that privacy or compliance rules keep from being shared. Second, data with coverage gaps or class imbalance that makes models unreliable. Third, data that exists but cannot reach training pipelines. What is zero-access architecture? Original data stays inside the client environment. DTS analyzes statistical properties in place, and only the DP-protected synthetic output moves on. No raw records are transferred outside. This makes the architecture suitable for environments where data cannot move: classified, regulated, or isolated networks. How is DTS different from Syntitan? DTS is the transformation engine; Syntitan is the platform it powers. Syntitan performs data-quality refinement as part of execution stability and can use a subset of DTS capabilities when DP-protected synthetic data is needed, while DTS is the platform's full AI-ready data transformation engine, which can also be deployed on its own. Restricted data. Usable AI. DTS rebuilds the data your AI can't use today into datasets it can train on tomorrow. GS Certified. KISA approved. Book architecture review Explore Syntitan Available on AWS Marketplace NCP Marketplace --- title: "LLM Capsule: Context-Preserving Data Layer for AI | CUBIG" url: "https://cubig.ai/llmcapsule" source: live --- Skip to content LLM Capsule · Context-Preserving Data Layer for AI Your AI stops at the data it can't touch. Capsule gets it through. LLM Capsule lets AI work on operational data that can't move raw. Those values become context-preserving stand-in values (substitutes). The AI runs on those. Usable results come back inside your environment, through a protected mapping layer that never leaves. Book architecture review Explore LLM Capsule Available on AWS Marketplace The problem Why enterprise AI projects stall before production When you put AI to work on sensitive enterprise data, the hardest part isn't the model. It's the data. Each approach breaks at a different layer, and the risk that remains stalls the project right before production. 01 External LLMs raise enterprise ROI ChatGPT, Claude, Gemini, and in-region EU models are already good enough to reshape ops workflows, from root-cause analysis to clinical drafting. BUT 02 PII guardrails alone aren't enough Free-text fields like a CS ticket's Details column mix customer names, contact info, and claim narrative with no structure. Simple PII guardrails miss it. Blanket masking and redaction strip the context AI needs to give a useful answer. AND 03 Legacy and operational data is complex Network operations tickets, plant sensor archives, patient records, and mission briefs aren't clean structured data. They cross-reference each other, stay unstructured, and live in production systems that won't migrate. AND THE RESIDUAL RISK 04 Filtering alone leaves regulated risk standing Filters that only check each request run at the API edge. They don't process documents end-to-end, don't restore outputs, and don't cover cases where data must stay in-country and sending it outside is not allowed at all. SO The fix LLM Capsule This is the context-preserving data layer for AI. Operational values that can't move raw become DP-based, context-preserving substitutes with structure intact. You run them on any approved model path, external or on your own servers (on-prem), and reconstruct inside the workflow it came from. Capabilities Six reasons Capsule works inside real enterprise workflows Each capability fixes a specific way conventional approaches break. CAPABILITY 01 Get real results back AI outputs auto-restore your original names, figures, and references, ready for reports, legal reviews, and client deliverables. No manual reconstruction. CAPABILITY 02 Tables, tickets, logs, and runbooks stay readable to AI Tables, tickets, and document structure stay intact. AI reads the full operational structure instead of broken fragments that produce useless outputs. CAPABILITY 03 Runs inside the systems you already operate Air-gapped networks, on-premise servers, custom data systems, ServiceNow / SharePoint / Jira / OT historians. Capsule runs inside your existing environment with one added API call and no architectural change. Run an external LLM or on-prem local under a single governance. CAPABILITY 04 You define what's sensitive Customer-defined confidentiality markers beyond standard PII: device IDs, circuit IDs, deal terms, M&A code names, mission references, OT identifiers. Enterprise context, not generic privacy. CAPABILITY 05 Your workflow runs where your data already lives Your data stays inside. The model only ever sees protected stand-ins; the originals and the restore map never leave. CAPABILITY 06 You can change the policy tomorrow Time-shifting policy: yesterday's policy archived, today's enforced. When new regulations land, you update the markers without rebuilding pipelines. Versioned Scoped Access-controlled Change-logged Architecture The four-zone architecture Raw operational data stays inside the corporate environment. Only the protected capsule crosses zones, and restored output is reconstructed locally, inside the workflow it came from. ZONE 01 Corporate Internal Network The operational systems already live here. Capsule reads your existing systems in place with one API call. ZONE 02 Data boundary: structure-preserving substitution Detection finds the values you defined as sensitive. Structure-preserving, DP-based substitution replaces them while keeping structure and context intact, and only the protected working version leaves your environment. ZONE 03 In-House Team Governance and routing choose between an approved external LLM (ChatGPT / Claude / Gemini / in-region EU models) and an on-prem local model. Organizational policy and domain context stay intact. ZONE 04 Local: Auto Reconstruction Inside the organization, the AI response is auto-restored to original values. Data that left the boundary cannot be reconstructed outside, and business-ready output goes back to the workflow it came from. See the full diagram → Positioning Built to enable AI work, not to police it. AI gateway Manages model traffic: routing, auth, fallback, caching, rate limits, cost, observability. DLP Detects, classifies, or blocks sensitive content. Employee AI Helps workers search, chat, and automate tasks across company apps. LLM Capsule It changes what your workflow actually sends to the model. Data becomes a restorable capsule before it goes out, and comes back restored, inside your environment. Gateways route the call. Capsule changes what crosses the model boundary. Proof Validated in production environments Telecom · Industrial cybersecurity · Healthcare · Finance · Public sector · Legal · Cloud sovereignty 0.12s Per-page processing 2,200-character document Exact Exact, repeatable reconstruction inside your environment 98% Output similarity vs. processing on raw original AWS Marketplace FAQ Frequently asked questions What is a context-preserving data layer for AI? A context-preserving data layer for AI lets AI work on operational data that cannot move raw. Sensitive values become DP-based substitutes. Document structure, relationships, and meaning stay intact, so models reason over real business context. Usable outputs reconstruct inside your environment, and the original values and reconstruction mapping never leave your boundary. How do you run an LLM on operational data that cannot move raw? LLM Capsule turns sensitive values into DP-based, context-preserving substitutes before the request reaches the model. The model path can be external, but it sees only the protected working version. The original values and the reconstruction mapping stay inside your environment, where the response reconstructs. Can we run RAG or agents on documents that contain sensitive fields? Yes. LLM Capsule substitutes the sensitive elements in your RAG sources and agent context while keeping the structure that retrieval and reasoning depend on. So RAG and agent workflows run on operational data you could not send to an external model before. How is this different from masking or redaction? Masking and redaction destroy the meaning a model needs, and a record full of blanks is unusable. LLM Capsule keeps the format, relationships, and document structure intact with context-preserving substitutes, so the model still understands the task and teams get answers they can act on, with real values restored internally. How is LLM Capsule different from an AI gateway or DLP? An AI gateway routes and manages model traffic. DLP detects and blocks sensitive content. LLM Capsule changes what crosses the model boundary: it swaps operational values for context-preserving substitutes before model execution, then reconstructs the output inside your environment. Where does reconstruction happen, and what leaves our environment? Reconstruction happens only inside your organization. The model path can be external and sees only context-preserving substitutes. The original values and reconstruction mapping never leave your boundary. Reconstruction is a deterministic internal mapping, not a statistical recovery of differential-privacy values. See LLM Capsule run on your own enterprise documents. Bring your documents, deployment constraints, and one real workflow. We demonstrate it on your documents, in your environment, in 30 minutes. Book architecture review Explore LLM Capsule Available on AWS Marketplace --- title: "Company: CUBIG | AI-ready execution for enterprise AI" url: "https://cubig.ai/company" source: live --- Skip to content Company About CUBIG CUBIG builds the AI-ready data operating layer, turning restricted, unusable, and unstable enterprise data into states AI can actually run on. Book architecture review Explore Syntitan Category What is AI-ready execution? AI-ready execution is the operating layer that makes enterprise data usable , reliable , and stable for production AI. Most enterprises have data, but most of it is not ready for AI. Some is restricted and cannot reach AI safely. Some exists but is unusable because of missing values, bias, or coverage gaps. AI that works in a pilot (PoC) often breaks down in production. Once schemas, pipelines, or conditions change, results can no longer be reproduced. CUBIG builds the operating layer that closes these gaps. Two market entry paths, one long-term platform destination: Syntitan. Path A serves AI-ready data demand through Syntitan and DTS. Path B serves sensitive AI workflow enablement through LLM Capsule. Both converge into Syntitan. Mission Why we exist. Enterprise AI stops because data is restricted , unusable , or because execution becomes unstable in production. What we believe Most teams can make AI work in a PoC. Production is a different problem, and the cause is rarely the model or the compute. It is the state of the data underneath every run. That is where most projects stop before reaching production. We believe these three problems, not models or compute, are what keep enterprise AI out of production. CUBIG builds the operating layer that resolves all three. What we do Make data usable , reliable , and stable for production AI. Restricted data Sensitive or regulated data cannot reach AI safely. Compliance constraints keep it out of training, validation, and inference. Unusable data Data exists but is not usable: missing values, bias, coverage gaps, restricted access. The PoC works. Production does not. Unstable execution After deployment, data and execution conditions change, so results cannot be reproduced. Traceability disappears and root cause becomes impossible. We rebuild restricted and scarce data into AI-ready states, so it becomes usable. We fix the data state behind every run, so results stay reproducible. That is what turns a PoC into production. Story How we got here. CUBIG was founded in 2021 by a team that had spent years building enterprise AI in regulated industries: finance, healthcare, defense. We kept hitting the same three walls. Data we could not use because of compliance. Data too damaged for training. AI that worked in a PoC but degraded after deployment. We looked at existing tools. Data governance managed access but did not make data usable. MLOps tracked models but not the data state behind each run. None were designed to work together as one layer. The problem was not any single tool. It was the absence of a layer that handled all three blockers at once. So we built what was missing: Syntitan, the AI-Ready Data Platform that fills the missing layer between enterprise data management and real AI execution. Two core capabilities carry the work. DTS rebuilds locked, scarce, or regulated data into AI-ready states. LLM Capsule runs LLM and agent workflows while original values stay inside your environment. At a glance Where we are today. 2021 Founded Seongnam, Korea · Belfast, UK AWS Marketplace partner LLM Capsule on AWS Marketplace 2 Global entities CUBIG Corp (KR) · CUBIG Ltd (UK) Team The people building it. Our team comes from enterprise AI, data engineering, and privacy technology. We have built and stress-tested AI systems at scale. That is why we know exactly where production AI fails. CUBIG Team AI infrastructure engineers Practitioners who have operated AI in regulated enterprise environments: finance, healthcare, manufacturing. Every product decision comes from something we had to fix ourselves. Research & Privacy Data reconstruction specialists The research team behind the DTS engine and the field-handling layer inside LLM Capsule. Measured guarantees, not policy promises. Enterprise Engineering Platform & integration Responsible for Syntitan: Release State, Run Binding, and the integration layer that connects to existing ML pipelines, data platforms, and runtime environments. Platform & capabilities The structure that makes data AI-ready. Syntitan is the long-term platform. DTS and LLM Capsule are core capabilities that converge into it, not standalone products beside it. Platform Syntitan The AI-Ready Data Platform. Diagnose data readiness, prepare it for use, and fix the data state as a Release State. Every AI or agent run binds to that state, so you can reproduce it, diff what changed, and prove the result with a real run whenever you need to. Release State · Run Binding · Reproduce More → Capability DTS AI-ready data transformation engine. Rebuilds locked, scarce, or regulated data into AI-ready states, expands coverage, and restores data utility. Works within Syntitan and on its own. Rebuild · Transform · AI-ready states More → Capability LLM Capsule Context-preserving data layer for AI. Runs LLM, RAG, and agent workflows on data that cannot leave in its original form. Usable results are reconstructed inside your environment. Structure-preserving substitution with business-ready reconstruction. Structure-preserving · Enablement · Business-ready More → Partners Trusted by enterprise and government. From global cloud and research partners to major Korean financial institutions and national defense, CUBIG operates where the data stakes are highest. Technology & Cloud Marketplace · LLM Capsule Technology partner Cloud partner · DTS Research partner Finance & Enterprise Enterprise partner Industrial Bank of Korea Financial partner Financial partner Defense & Public sector Defense partner Defense partner Medical institution Public sector Values How we work. 01 A layer, not features We build the layer everything else runs on. Features solve single problems. A layer solves a whole class of problems and supports every AI system built on top of it. Every decision starts with which problem it solves and what it makes possible next. 02 Production is the only test A PoC is not proof. We build for production: restricted data, compliance constraints, schema changes, multi-team pipelines. Every decision is tested against one question: does it hold when conditions change after deployment? 03 Evidence over assertion Every claim is backed by operational evidence: before and after outcomes, state comparisons, reproducible runs. We do not say "improves accuracy" without showing what changed and how it can be verified. If we cannot prove it, we do not say it. Contact Get in touch. Enterprise & Architecture Book architecture review Map your production constraints (data that is locked, damaged, or drifting) to the right path across Syntitan, DTS, and LLM Capsule. Book architecture review → General inquiries Press, partnerships, research Research collaboration, press, partnership discussions, or anything not covered above. [email protected] Korea headquarters CUBIG Corp 4F, NAVER 1784, 95 Jeongjail-ro, Bundang-gu, Seongnam-si, Gyeonggi-do, Republic of Korea. United Kingdom CUBIG Ltd 21 Arthur Street, Belfast, Antrim, BT1 4GA, United Kingdom. Make your AI runs reproducible in production. Start with Syntitan, and bring in DTS and LLM Capsule where you need them. Book architecture review Run a sample proof --- title: "Proof: Trust Evidence for Enterprise AI | CUBIG" url: "https://cubig.ai/proof" source: live --- Skip to content Proof · Trust Evidence Proof, not promises. Data that is usable for AI execution, privacy-safe in production, and stable across runs. The case records, certifications, patents, awards, and partnerships that prove it, in one place. < 4 hrs Root cause identification 21 days → 4 hrs · 99% faster 88.55% F1-score (DTS augmentation) 58.55% → 88.55% · +30pp 1 day Model time-to-deploy 4 weeks → 1 day · 90% faster Customers & Partners Across banking, insurance, legal, the public sector, and telecom. Certifications & Awards Backed by international certifications and industry awards. Gartner® “Emerging Tech: AI Vendor Race: Tech Innovators in Agentic AI — Solution Accelerators” (2026) Gartner® “Emerging Tech: AI Vendor Race: Most Prominent Use Cases in Agentic AI by Industry” (2026) Gartner® “Emerging Tech: Provider Differentiation Strategy—Trends for Hyper-Synthetic Data” (2025) Named a Representative Vendor Operational Evidence From PoC to production. Financial Services Model retraining pipeline: schema drift (data format change) detection Execution Stability 21 d Root cause time (before) < 4 hr Detection time (after) 2 Feature columns removed 1 Schema type change Before Schema change in upstream data caused silent model degradation. Root cause took 21 days to identify. By then, downstream decisions had already been affected. After Syntitan Release State detected the schema diff as the data came in. Issue flagged before the next training run triggered. No degraded model reached production. What Changed 2 feature columns removed from upstream feed. 1 schema type change introduced silently. Reproduce Re-run verified Root cause: 21 days → under 4 hours State Card Change Log Re-run Record Schema Diff Telecom Real-time inference service: pipeline version rollback Execution Stability Unknown Drift source (before) < 2 hr Rollback time (after) Matched Score distribution vs baseline Before Preprocessing pipeline update produced inconsistent scores in production. No way to trace which version caused the score drift. After Run Binding linked every score to its exact Release State. Rollback to stable state completed in under 2 hours. What Changed Normalization logic updated across the preprocessing step. Feature scaling range shifted by 12%. Release State diff identified both changes with exact pipeline version reference. Reproduce Re-run verified Matched baseline State Card Change Log Re-run Record Manufacturing Quality inspection model: rare defect class coverage Data Usability 3 Underrepresented classes +30pp F1-score improvement DP-safe Synthesis method Before Rare defect class underrepresented in training data. F1-score capped at 58.55% . Model missed edge cases in production. After DTS generated privacy-preserving synthetic samples (differential privacy) for 3 underrepresented classes. F1-score rose to 88.55% (+30pp) . Coverage gap closed before next training cycle. What Changed 3 underrepresented defect classes augmented with DP-safe synthetic data. Class distribution rebalanced. Augmented dataset versioned within Syntitan Release State. Reproduce Bound to Release State Recall verified on unseen data State Card Dataset Version Re-run Record Class Dist. Log Healthcare Clinical AI validation: restricted patient data replacement Data Usability Blocked Validation status (before) Unblocked Validation status (after) DP-safe Synthesis method Before Real patient records required for model validation could not be accessed due to regulatory constraints. Validation pipeline stalled. After DTS generated DP-safe synthetic patient records matching real distribution characteristics without containing real identifiable information. Validation unblocked. What Changed Non-accessible real records replaced with DP-safe synthetic equivalents. Data distribution preserved. Compliance review passed. Validation pipeline resumed without modification. Reproduce Versioned in Syntitan Re-runnable on demand State Card DP Audit Log Dataset Version Insurance LLM-assisted claims processing: sensitive data substitution Sensitive Workflow Enablement Exposed Sensitive data in prompts (before) Substituted Sensitive fields (after) Preserved Output usability Before Claims documents containing policyholder names, ID numbers, and medical details were sent directly to an external LLM API. Compliance team blocked the workflow. After LLM Capsule substituted sensitive fields with restorable stand-ins before submission. Outputs returned and reconstructed locally for downstream system use. What Changed LLM Capsule layer inserted into the workflow. Substitution covered names, IDs, dates, and medical field patterns. Sensitive raw values stayed in the protected mapping layer inside the environment. Reproduce Every run logged Bound to Release State State Card Substitution Log Mapping Record Re-run Record Retail / E-commerce Recommendation engine: runtime environment drift Execution Stability Days Debugging time (before) < 3 hr Root cause identified (after) Exact Environment reproduced Before Recommendation scores degraded after a routine infrastructure upgrade. Engineers could not reproduce the pre-upgrade behavior. After Syntitan captured every runtime parameter in the Release State at execution time, with Run Binding tying the run to that state. Pre-upgrade Release State re-run in under 3 hours. What Changed Library version bump changed default float precision handling. Embedding normalization behavior altered. Release State diff identified the exact library version delta. Reproduce State restored exactly Score delta measured State Card Runtime Snapshot Change Log Re-run Record Public Sector Aggregate-data release: automated screening & audit trail Execution Stability Manual Release screening (before) Automated Screening (after) 0.94 PII detection F1 Multi-agent Detect · trace · transform Before Data-center users exporting sensitive aggregate statistics required manual, per-desk screening and release review. The process was inconsistent and hard to audit. After A per-desk module and a multi-agent pipeline detect, trace, and transform personal information. The release review is now automated and standardized. What Changed Release State fingerprints the data before and after transformation, so which records were changed, and how, stays traceable for audit. Reproduce Release replayable Ready for regulatory inspection Screening Report Release Audit Log Detection Trace State Card Defense LLM adoption in isolated networks: classified context preserved Sensitive Workflow Enablement Blocked AI adoption (before) Enabled AI adoption (after) 0% Raw context egress N2SF Guideline aligned Before In an air-gapped environment, classified context could not be sent outside, so adopting an external LLM for the work stalled before it began. After LLM Capsule substitutes the sensitive context with restorable stand-ins locally. Only the substituted capsule reaches the external LLM, and the result is reconstructed locally inside the boundary. The original context stays local. That keeps the workflow aligned with N2SF guidelines. What Changed Sensitive context is substituted with local stand-ins before processing and reconstructed locally afterward. The original stays within the local boundary. Reproduce Logged locally Inspectable per request Local Mapping Layer Audit Log N2SF Alignment Industrial · OT/ICS OT network data: AI-ready transformation for threat analysis Data Usability Restricted Raw OT data (before) Enabled AI threat analysis (after) Structure-preserving Transformation Before OT/ICS network data carried sensitive operational details, so it could not be sent to an external AI for automated threat analysis. After Structure-preserving transformation lets an AI agent analyze the network data and answer threat questions. Sensitive values are replaced with stand-ins while relationships stay intact. (Integrated with a global OT security platform's detection solution.) What Changed Network-data sensitive fields are substituted while topology and relationships are kept intact, so the agent can reason over realistic context. Reproduce Fixed data state Re-run verified Transformed Dataset Agent Analysis Log Structure Map Telecom Network operations model: topology change impact tracing Execution Stability Unknown Drift source (before) Traced Root cause (after) Diff Topology change surfaced Before After a change in the network layout, a NOC AI model's outputs drifted, but engineers could not tell which change caused it, since the model itself was unchanged. After Release State captured the network state of each run. The Diff surfaced the exact layout change. The pre-change state was re-run to confirm it. What Changed Specific network-layout and configuration elements changed between runs. The Release State diff identified them against the bound prior run. Reproduce Pre-change run replayed Fix validated State Card Topology Diff Change Log Re-run Record Certifications Standards, third-party verified. Third-party validated certifications across information security, privacy, and operational standards. Information Security Management ISO/IEC 27001:2022 · 2026 International standard for information security management. Demonstrates a systematic approach to managing sensitive information. AI Management System ISO/IEC 42001:2023 · 2026 International standard for AI management systems. Demonstrates responsible AI governance and risk management. GS Grade 1 · DTS GS Certification Grade 1 · 2025 Korean SW Quality Certification, Grade 1 (2025). Verified quality, eligible for public procurement. GS Grade 1 · LLM Capsule GS Certification Grade 1 · 2024 Korean SW Quality Certification, Grade 1 (2024). Listed on the public Innovation Marketplace for procurement. KISA Fast Track KISA · 2024 Selected for the KISA information-security industry Fast Track program. Patents The patents behind the tech. Registered patents and pending applications behind Syntitan, DTS, and LLM Capsule. The technical foundation of the operating layer. ▸ Patent · KR US Registered AI-Based Service Providing Method Without Leaking Private Information and Client Apparatus KR Reg. No. 10-2757651 (App. 10-2023-0133086, Registered 2025-01-16) · US App. No. 18/908,054 (Filed 2024-10-07, allowed for registration 2026-07) Core LLM Capsule patent. Method and client apparatus for AI services without exposing private information. View patent → ▸ Patent · KR Registered Method and Data Processing Apparatus for De-identifying Data While Preserving Target Characteristics KR Reg. No. 10-2926046 · App. No. 10-2023-0167085 · Registered 2026-02-06 Core DTS patent. Method for de-identifying source data while preserving target characteristics such as statistical distributions and label structure. View patent → ▸ Patent · KR Registered · US Pending Synthetic Data Generation Method Without Leaking Target Information and Client Apparatus KR Reg. No. 10-2818137 (App. 10-2024-0017564, Registered 2025-06-04) · US App. No. 19/039,319 (under examination) DTS synthesis patent. Client-server architecture for generating synthetic data without exposing target information. View patent → ▸ Patent · KR Registered · US Pending Method and Data Processing Apparatus for Generating a Synthetic Dataset Containing Multiple Attributes KR Reg. No. 10-2818136 · App. No. 10-2024-0131551 · Registered 2025-06-04 · US Pub. No. US 2026/0017275 A1 (under examination) DTS multi-attribute synthesis patent. Method for generating complex synthetic datasets that span multiple feature columns and attribute types. View patent → ▸ Patent · KR Pending Data Management Method and System for AI Execution Control KR App. No. 10-2026-0053050 · Filed 2026-03-24 · Expedited examination granted 2026-04-08 Core Syntitan patent application. Method and system for controlling and managing data state within AI execution environments. Expedited examination granted. ▸ Patent · KR US Pending Method for Providing Security for On-Device Artificial Intelligence Models KR App. No. 10-2025-0003223 (Filed 2025-01-09) / 10-2026-0000037 (priority, Filed 2026-01-02) · US App. (Ref. PO25-025-US, via export-voucher) Security provisioning method for AI models running on-device. ▸ Patent · KR US Pending Method and Data Processing Apparatus for Validating Synthetic Datasets for Model Training KR App. No. 10-2024-0174041 (Filed 2024-11-28) · US App. No. 19/400,665 (Filed 2025-11-25) · under examination DTS patent application for validating synthetic datasets used to build training models. ▸ Patent · KR Pending Method and Data Processing Apparatus for Filtering Synthetic Datasets for Model Training KR App. No. 10-2024-0174042 · Filed 2024-11-28 · Under examination DTS patent application. Method for filtering synthetic datasets prior to model training. ▸ Patent · KR Pending Method and Inference Apparatus for Building Deep Learning Models Robust to Private Information Exposure KR App. No. 10-2023-0074745 · Filed 2023-06-12 · Office Action response due 2026-07-25 Deep learning model construction robust to private information exposure. Applicant: Ewha Womans University (co-research). ▸ Patent · KR Pending Method and Analysis Apparatus for Building Artificial Intelligence Models that Process Heterogeneous Datasets KR App. No. 10-2023-0013029 · Filed 2023-01-31 · Under examination (response filed 2026-01-14) AI model construction method for heterogeneous datasets. Applicant: Ewha Womans University (co-research). Research The research behind the products. Selected publications by CUBIG founders, from peer-reviewed venues to a survey preprint. The privacy and robustness research behind Syntitan, DTS, and LLM Capsule. Publication · JMLR 2025 Regularizing Hard Examples Improves Adversarial Robustness Hyungyu Lee, Saehyung Lee, Ho Bae, Sungroh Yoon · Journal of Machine Learning Research · 2025 Adversarial robustness method that regularizes hard examples to improve robust generalization. Publication · ICLR 2024 DAFA: Distance-Aware Fair Adversarial Training Hyungyu Lee, Saehyung Lee, Hyemi Jang, Junsung Park, Ho Bae, Sungroh Yoon · ICLR · Vienna, May 2024 Adversarial training method that enforces fairness across subgroups via distance-aware margin adjustment. Publication · Sensors 2024 Evaluation of Malware Classification Models for Heterogeneous Data Ho Bae · Sensors (MDPI) · 2024 Study of malware-classifier explainability on heterogeneous data. Existing explanations fall short, and high accuracy can give a misleading sense of security. Publication · ESORICS 2024 VFLIP: A Backdoor Defense for Vertical Federated Learning via Identification and Purification Yungi Cho, Woorim Han, Miseon Yu, Younghan Lee, Ho Bae, Yunheung Paek · ESORICS · 2024 First backdoor defense specialized for Vertical Federated Learning. It identifies and purifies backdoor-triggered embeddings at inference. Publication · BIBM 2023 Privacy-Preserving Publishing of Individual-Level Medical Data for Cloud Services Ho Bae, Heonseok Ha, Siwon Kim · IEEE BIBM · Istanbul, Dec 2023 Formal privacy-preserving framework for publishing patient-level medical records to cloud services, with emphasis on utility preservation under strict privacy constraints. Publication · ESORICS 2023 FLGuard: Byzantine-Robust Federated Learning via Ensemble of Contrastive Models Younghan Lee, Yungi Cho, Woorim Han, Ho Bae, Yunheung Paek · ESORICS · 2023 Byzantine-robust federated learning that detects malicious clients via an ensemble of contrastive models, strong under non-IID data. Publication · RAID 2023 Exploring Clustered Federated Learning's Vulnerability against Property Inference Attack Hyunjun Kim, Yungi Cho, Younghan Lee, Ho Bae, Yunheung Paek · RAID · 2023 Reveals property-inference privacy risks in clustered federated learning. Publication · IEEE/ACM TCBB 2022 DNA Privacy: Analyzing Malicious DNA Sequences Using Deep Neural Networks Ho Bae, Seonwoo Min, Hyun-Soo Choi, Sungroh Yoon · IEEE/ACM Transactions on Computational Biology and Bioinformatics · 2022 Deep-learning analysis of malicious DNA sequences for security and privacy in genomic data. Publication · BMVC 2022 MPGAN: Membership Privacy-Preserving GAN Heonseok Ha, Uiwon Hwang, Jaehee Jang, Ho Bae, Sungroh Yoon · BMVC · London, Nov 2022 GAN training method that prevents membership inference attacks on generated data, providing formal, provable privacy properties for synthetic outputs. Publication · ACM AsiaCCS 2022 Membership Feature Disentanglement Network Heonseok Ha, J Jang, Y Jeong, S Yoon · ACM Asia Conference on Computer and Communications Security · 2022 Network architecture that disentangles membership-sensitive features from model representations, reducing exposure to membership inference attacks. Publication · IEEE Access 2021 Gradient Masking of Label Smoothing in Adversarial Robustness Hyungyu Lee, Ho Bae, Sungroh Yoon · IEEE Access · 2021 Analysis of how label smoothing induces gradient masking, a false sense of robustness that does not transfer to true adversarial settings. Publication · IEEE TAI 2021 Learn2Evade: Learning-based Generative Model for Evading PDF Malware Classifiers Ho Bae, Younghan Lee, Yohan Kim, Uiwon Hwang, Sungroh Yoon, Yunheung Paek · IEEE Transactions on Artificial Intelligence · Aug 2021 Adversarial generative modeling of malware evasion: learning to produce feature-space perturbations that bypass PDF malware classifiers while preserving functionality. Publication · IEEE Access 2020 Anomaly Detection by Learning Dynamics From a Graph Jaekoo Lee, Ho Bae, Sungroh Yoon · IEEE Access · 2020 Graph-based anomaly detection that learns system dynamics to flag abnormal behavior. Publication · PSB 2020 AnomiGAN: Generative Adversarial Networks for Anonymizing Private Medical Data Ho Bae, Dahuin Jung, Hyun-Soo Choi, Sungroh Yoon · Pacific Symposium on Biocomputing · Hawaii, Jan 2020 GAN-based anonymization of private medical datasets while preserving statistical utility for downstream analysis. Publication · PSB 2019 DNA Steganalysis Using Deep Recurrent Neural Networks Ho Bae, Byunghan Lee, Sunyoung Kwon, Sungroh Yoon · Pacific Symposium on Biocomputing · Hawaii, Jan 2019 Deep recurrent-network method for detecting hidden messages embedded in DNA sequences (steganalysis), applied to genomic data. Preprint · arXiv 2018 Security and Privacy Issues in Deep Learning Ho Bae, Jaehee Jang, Dahuin Jung, Hyemi Jang, Heonseok Ha, Sungroh Yoon · arXiv:1807.11655 · 2018 Comprehensive survey of attack surfaces and defenses in deep learning systems, covering adversarial examples, model extraction, and data poisoning. Awards & Recognition Recognized by government and industry. From government program selections to industry awards at home and abroad: third-party validation of our technology and business. Industry Award Deutsche Telekom T-Challenge 2026 · 2nd Place T-Mobile / Deutsche Telekom · 2026 Placed 2nd in the T-Challenge global open-innovation program (T-Mobile / Deutsche Telekom), recognized for LLM Capsule's context-preserving substitution and local reconstruction. Industry Recognition 2026 Emerging AI+X Top 100 Korea AI Industry Association · 2026 Selected for the 2026 Emerging AI+X Top 100 for its AI-ready data technology. Government Program Selected Supplier · 2026 AI Voucher Program Ministry of Science and ICT · NIPA · 2026 Selected as a supplier for the 2026 AI (Cloud) Voucher program, letting SMEs adopt CUBIG AI-ready data solutions via government vouchers. Government Program Selected Supplier · 2026 Data Voucher Program Korea Data Agency (K-DATA) · 2026 Selected as a Data Voucher supplier, rebuilding restricted enterprise data into AI-ready data with DTS. Government Program Ultra-Gap Startup 1000+ (DIPS 1000+) Ministry of SMEs and Startups · KISED · 2025 Selected in 2025 as a top deep-tech startup (AI / big-data) in the Ultra-Gap Startup 1000+ project, and as a Global ICT Future Unicorn the same year. Global Membership NVIDIA Inception 2024–2025 Member of NVIDIA Inception, the global program for AI startups. Government Award Information Security Product Innovation Award · Minister of Science and ICT Prize Ministry of Science and ICT · 2024 Grand Prize, Information & Physical Security category, at the 2024 H2 Information Security Product Innovation Awards (Minister of Science and ICT Prize). Industry Award Startup World Cup Finalist 2024 Finalist at the global Startup World Cup. Industry Award NextRise Global Innovator 2024 Selected as a NextRise Global Innovator. Accelerator SK Telecom × Hana Bank AI Accelerator SK Telecom · Hana Bank · 2024 Selected for the 2nd SK Telecom × Hana Bank AI startup accelerator (15 of 230 applicants). Partnerships The ecosystem we build with. Cloud and infrastructure partners that the operating layer composes with. AWS Marketplace DTS · LLM Capsule Naver Cloud Platform DTS FAQ Frequently asked questions. Common questions on operational evidence, reproducibility, and sensitive-data handling. What is operational evidence in AI systems? + Operational evidence in AI systems is concrete, verifiable documentation that shows how an AI system behaves in production. It shows what changed between runs, what caused the difference, and whether the same conditions can be replayed. It includes before/after outcomes, state comparisons, and re-run records. What is reproducible AI execution? + Reproducible AI execution means a run can be replayed and verified. The conditions include data state, schema, preprocessing logic, and runtime dependencies. Syntitan achieves this through Release State, Run Binding, and Reproduce. Why is AI execution unstable in production? + AI execution becomes unstable in production when execution conditions change after deployment: data schemas, preprocessing logic, dependencies, data windows. Any of these can cause AI results to drift without any change to the model itself. When execution state is not captured and fixed, it is hard to pinpoint which change caused a production issue. How does CUBIG handle sensitive data in LLM workflows? + CUBIG's LLM Capsule substitutes sensitive values with restorable stand-ins before the data crosses to an external LLM. The original values stay in the protected mapping layer inside the enterprise boundary. The LLM operates on the substituted version; output is reconstructed locally for downstream use. This pattern is the Substitute · Execute · Reconstruct sequence. Can CUBIG be deployed on-premises or air-gapped? + Yes. Syntitan and LLM Capsule are designed to support on-premises and air-gapped deployments for regulated and defense workflows. Detailed architecture available on request for qualified evaluations. The proof is in the execution. Bring one workflow that isn't reproducible today. We'll show, on your data, what AI-ready execution changes. Book architecture review Run a sample proof --- title: "Business Process Automation Starts by Defining the Work" url: "https://cubig.ai/articles/business-process-automation-define-work/" source: live --- Halfway through an automation project, our team needed a simple comparison for a report: how long the manual process took and how long the new system took. We could measure the system. No one had ever produced the manual baseline. Documentation was not the problem. Years of reports described the delays, the workload, and the people involved. Requirements stated the intended outcome. Process tables named phases and assigned roles. None of those records broke the work into units that could be counted, timed, or compared. That gap changed the project. Before we could defend the automation, we had to define the work itself. The distinction matters because process documentation can describe what should happen without showing how the work actually moves from one decision to the next. Business process automation requires that second view. A team must be able to identify each task, the person or system responsible for it, the evidence needed to complete it, and the condition that allows the work to move forward. Why a documented process can still be undefined A requirement gives a team a destination. A phase diagram gives the destination a rough route. Neither necessarily captures the actions and judgments that make the process run. In this project, two authoritative documents described the same workflow at different points in time. An earlier design assigned one step to the people submitting the data. Current operating rules prohibited those people from performing that step because of a stated risk. Each document made sense within its own context. The contradiction appeared only when we placed the work in sequence and assigned an owner to every step. This is why documentation volume is a poor proxy for process definition. A library can be complete at the document level and still leave critical operating questions unanswered: What is the smallest unit of work that can be assigned or measured? Which steps apply rules, and which require human judgment? Who owns each decision and its exceptions? What evidence allows the next step to begin? Which inputs are supported, and why are others out of scope? The Business Process Model and Notation specification provides a standard way to represent activities, events, decisions, and participants. That representation is useful because it makes the workflow inspectable. It does not, by itself, prove that the team captured the real work or measured it correctly. Practitioners still have to expose the decisions they make in practice, including the ones that never reached a formal document. The IEEE Task Force on Process Mining describes process mining as a way to discover and monitor processes from event logs. Logs can reveal the paths that people and systems actually recorded, rather than only the path prescribed by a process diagram. That evidence still has limits: it cannot recover a judgment, exception, or handoff that was never recorded. Event data and practitioner knowledge therefore provide complementary views of the process. Define the work at the level AI must execute We found three missing layers. The first was work units. The existing documents described broad phases such as review. A phase cannot be assigned cleanly to a model, a system, or a person until the team separates it into observable actions. The second was the judgment practitioners applied while doing the work. Experienced operators knew which values to distrust, what they checked first, when they asked a colleague, and which exception stopped the process. That knowledge lived with the people doing the work and varied slightly across the team. The third was measurement. Nobody could write the expression that represented the cost of the manual work. The missing artifact was not the final number. It was the set of variables, units, and owners required to calculate it. Defining those layers changed the design. Automated output became a draft until a person confirmed the judgment that mattered. Unresolved cases stopped instead of flowing downstream. Supported and unsupported input types were separated with explicit reasons. The team placed decisions that had been scattered across the workflow into one ordered review sequence and gave each one a named owner. The table below turns those lessons into a definition check. Process definition evidence checklist Definition elementQuestion to answerEvidence to retainFailure if missing Work unitWhat observable action changes the state of the case?Input, output, completion condition, and time unitA broad phase is assigned to AI without a testable boundary Decision ruleWhich cases follow a rule, and which require judgment?Rule, exception, escalation point, and approval conditionThe system treats an exception as routine work OwnerWho confirms the result or accepts the risk?Named role and handoff recordDecisions move forward without accountable review Input boundaryWhich input types are supported, and why?Typology, inclusion rule, and exclusion reasonUnsupported cases are processed as if they were valid MeasurementWhich variables describe time, cost, quality, or risk?Formula, unit, source, and observation windowSavings or performance claims cannot be reproduced This checklist also reflects ISO's public guidance on the process approach in ISO 9001:2015. The guidance asks organizations to define process inputs and outputs, sequence and interactions, ownership, controls, and measurement. It is not an AI-readiness test, but it reinforces the need for operating evidence beyond a stated objective. Write the formula before estimating the return Once the work units existed, we could write a simple measurement structure: total time = number of tables × table-review time + number of items × item-processing time + number of tables × verification time The expression is illustrative. Its terms came from one project, and the coefficients were still unknown when we first wrote it. That was the point. A formula exposes exactly what must be measured and who can provide each value. The practitioners owned the manual baseline because they performed the existing work. The system team owned the new measurements because it could observe the automated path. A defensible comparison required both. Agreement was not a general approval of a workflow diagram. It meant that each group understood which part of the evidence it owned and stood behind that evidence. This approach also prevents a common measurement error. A metric describes a particular implementation. If a component, decision rule, input boundary, or human handoff changes, the metric may no longer describe the process. The team must define the measurement again when the work changes shape. The NIST AI Risk Management Framework Playbook also places context, assigned roles, measurement, and review within the same risk-management process. It does not prescribe this formula. Its relevance here is narrower: AI systems should be evaluated for their intended context, with responsibilities and measurement defined in advance. Automation and transformation are different projects A before-and-after formula helps distinguish automation from transformation. Automation keeps the expression and reduces one or more coefficients. The same work units remain, the same decisions exist, and ownership stays largely intact. A system performs a step faster or with less manual effort. Transformation changes the expression. Some terms disappear because the old step is no longer needed. Some decisions move from execution to confirmation. New work appears, such as reviewing an unfamiliar input type, maintaining a typology, or recalibrating a metric after an implementation change. A task that was impractical for people may become a new system stage. Neither category is inherently better. The distinction matters because the scope, risk, evidence, and expected return are different. Calling both projects automation can hide new human work and leave owners unprepared. Calling a coefficient reduction transformation can overstate what actually changed. The practical test is straightforward: put the old and new expressions side by side. If the terms stay the same and only their coefficients shrink, the team automated the process. If terms disappear, change owners, or appear for the first time, the team changed the work itself. Current workThe baseline expression Review+Process+Verify Terms and owners establish the baseline. AutomationSame expression Review+Process+Verify Same terms, smaller coefficients. TransformationNew expression Review+Process+Confirm new owner+Exception work Terms are removed, reassigned, or added. AI-ready data depends on the work you define Teams often assess whether data is clean, complete, labeled, or documented before an AI project. Those checks matter, but they cannot answer whether the data is ready for an undefined task. Readiness is evaluated against a target. The team needs to know what the system must do, which input types it will accept, what result counts as usable, who reviews exceptions, and under which operating conditions the result will be trusted. If the work changes during transformation, the readiness test must change with it. This is also the boundary between process definition and data qualification. Process owners define the work and its acceptance conditions. Data and AI teams can then test whether a specific data state supports those conditions. One activity cannot substitute for the other. Use a Proof Run only after the target is clear For this decision, a Syntitan Proof Run becomes relevant after the team can name the task and its success metric. Its current Proof Run workflow uses a selected model task and target metric to compare raw and AI-ready data states. That comparison can answer a bounded question: did the changed data state improve the result under the stated test conditions? It cannot define the business process, produce a missing manual baseline, assign human ownership, or prove that the organization completed a transformation. Those decisions remain with the people who own the work. The sequence matters. Define the task, decisions, boundaries, and measurement first. Then test whether the data state can support them. Otherwise, the team may produce a technically valid comparison for a process it still cannot explain. Run the formula test before the demo Before selecting a model or presenting an automation business case, ask the process owner for the formula. Do not ask for a rough savings estimate. Ask for the variables, units, and owner of each term. Then check whether the team can: break the process into observable work units; identify the rules, judgments, and escalation points; assign an accountable owner to each decision; define supported inputs and explicit exclusions; show where each baseline and system measurement comes from; and explain whether the proposed design reduces coefficients or changes the expression. If the team can answer those questions, it has a process that can be scoped, tested, and measured. If it cannot, the first milestone is not a model. It is a definition of the work. With the work and success metric defined, the next step is to test whether the data is ready for the task. See how Syntitan compares raw and AI-ready data under controlled Proof Run conditions. --- title: "Alation Data Catalog vs Syntitan: Where the Evidence Chain Changes" url: "https://cubig.ai/articles/alation-data-catalog-vs-syntitan/" source: live --- An enterprise AI agent returns an incorrect answer after a source table changes. Using the Alation data catalog within its broader AIOS environment, the team can trace the relevant data, context, agent, and policy, run evaluations, and route the correction to the right layer. The remaining question is narrower: Did the next release use a data state that was qualified for this target, and can the team link the production run to that exact state? Alation has moved well beyond conventional data cataloging. Its current Alation Intelligence Operating System positioning connects data, business context, agents, governance, evaluations, and feedback loops. Syntitan is CUBIG's AI-Ready Data Platform, with a comparison-relevant focus on Core Readiness, target-specific data-state Qualification, and Operating Evidence. The products overlap in data quality, evaluation, governance, and operational feedback. The decision should therefore turn on a specific evidence question: Which data state was judged fit for this target, and which actual run used it? If Alation and the surrounding data platform already answer that question with sufficient evidence, another system may not be necessary. If the evidence chain stops at governed context, agent evaluation, or quality-monitor history, Syntitan may address a narrower data-state gap. Alation is now an intelligence operating system, not only a catalog AIOS is presented as one system spanning data, context, agents, governance, self-improving feedback, and open interfaces. Its foundation includes searchable metadata, ownership, lineage, trust signals, curation, and data quality. The context layer adds governed data products, a marketplace, ontologies, and semantic models. Agent Studio brings agents and evaluations into the same operating model. Alation also applies an active metadata management model. Its Behavioral Analysis Engine uses signals such as popularity, search relevance, and usage recommendations. Bidirectional connectors bring governed metadata into working tools, while lineage and impact analysis show how a change may affect upstream and downstream assets and their users. This broader scope matters because enterprise AI failures are rarely isolated to one component. A wrong answer may begin with stale values, an ambiguous definition, an outdated instruction, a weak tool selection, or an agent that misapplies otherwise correct context. Alation's operating model is designed to help teams identify and correct the responsible layer. Its Data Products App packages governed assets for people and AI systems. Lineage connects sources, transformations, and downstream objects. Data Governance brings policies, workflows, classifications, access controls, and AI asset documentation into the catalog environment. Availability still depends on the deployment. The Data Products App is documented for Alation Cloud Service in the New User Experience, is disabled by default, and requires configuration and commercial access. Lineage depth varies by source, connector, ingestion method, and configuration. Buyers therefore need to evaluate the exact edition and architecture in scope. Alation already produces meaningful evaluation and quality evidence Agent Studio evaluations test an agent against defined cases. Alation describes a continuous loop: Build Agent -> Connect to Data Products -> Define Evaluations -> Test Accuracy -> Improve Metadata -> Test Again An evaluation case can define a natural-language request and an expected output, such as executable SQL, expected search results, or summarization guidance. Each run begins with a fresh chat and compares the output with the defined standard. This gives a team evidence about whether an agent and its knowledge layer answer the selected business questions correctly. Alation also states that human feedback remains part of the process. Alation Data Quality monitors provide another evidence stream. Teams can apply table and column checks, run monitors manually or on a schedule, detect anomalies, notify operators, and create incidents. Run History records timestamps, status, failed or errored checks, check definitions, and observed values. Lineage and catalog history add the upstream sources, downstream consumers, transformations, edited attributes, warnings, endorsements, and deprecations needed for investigation and corrective action. For many teams, this may already be enough. If the surrounding platform also preserves immutable data references, target conditions, release decisions, and actual-run linkage, the evidence chain may be complete without Syntitan. The key difference is what the evidence describes For this comparison, Alation's documented capabilities address a governed system of data, context, agents, and policy. Its evaluations test whether an agent produces the expected result for defined cases. Its data-quality history records whether monitored assets passed specific checks at specific times. Its feedback loop helps the team identify and correct the layer behind a failure. Syntitan focuses on the state of the data used for a selected AI task. It asks whether an identified data state fits a defined Target Profile and whether the released state can be connected to the actual run that used it. An agent evaluation can show that a set of answers met expected standards without identifying every value, transformation, and condition used by a later production run. A quality-monitor record can show that selected checks passed or failed without establishing that the complete data state fit one model or agent task. A governed, certified data product may still need a separate record of the version and conditions approved for a particular release. A target-qualified data state does not supply business definitions, stewardship, policies, semantic relationships, or agent evaluation cases. Syntitan is not a substitute for Alation's wider intelligence and governance role. Syntitan defines a data-state Qualification and Assurance contract Syntitan's current operating model has three connected product areas: Baseline, Qualification, and Assurance. Baseline diagnoses General or Core Readiness across six axes: Usability, Integrity, Context, Consistency, Reproducibility, and Traceability. It can be used before a target exists. A high Baseline score is not a final Target Fit judgment, and every Baseline issue does not need to be resolved before a team can evaluate a specific target. Qualification judges whether a data state fits a defined Target Profile. The sequence is Target Profile, target-specific data refinement, Proof Run Evidence, and Qualification Result. The Proof Run compares evidence under defined AI conditions. The Qualification Result records whether the criteria were met, not met, or remained inconclusive. It does not certify the model or claim universal data quality. Assurance carries the identified state into operations through Release or Release State, Run Binding, Change Event or History, and Requalification. Run Binding connects the Data Release and relevant conditions to an actual AI execution. Release creation, operational-use approval, activation, and actual use remain separate states. This operating model can be summarized as: AI-ready = Core Readiness + Target Fit + Continuous Operating Evidence The formula is a control framework, not a performance guarantee. Release State and Run Binding describe the evidence contract; they do not promise a particular model or business outcome. Compare the evidence chain Comparison of Alation's documented role, the deployment evidence to verify, and Syntitan's defined role Decision pointAlation's documented roleWhat the buyer must verifySyntitan's defined role Enterprise data and contextCatalogs data, lineage, ownership, policy, semantics, and usage; packages governed data productsAre the required assets, definitions, relationships, and controls represented for this use case?Baseline diagnoses General/Core Readiness for the selected data state Agent behaviorEvaluates agents against defined cases and expected outputs; uses failures to improve metadata or guidanceDo the evaluation cases represent the production target and preserve the conditions needed to interpret the result?Proof Run Evidence supports a target-specific Qualification Result Data qualityRuns table and column checks, anomaly monitoring, alerts, incidents, and historical result reviewWhich failures block a release, and does the record identify the complete data state evaluated for the target?Baseline diagnoses common readiness; Qualification addresses gaps relevant to the Target Profile Lineage and changeShows upstream sources, downstream impact, transformations, and catalog historyDoes a detected change trigger the required retest, approval, and evidence update?Change Event or History supports requalification of the affected scope Data product releasePublishes governed, reusable data products through configured marketplace and approval controlsWhich exact product version, data reference, and conditions were approved for this AI release?Data Release and Release State preserve the identified data-state reference separately from use approval Production executionAIOS describes decision traces, feedback loops, and correction across data, context, and agentsCan each actual run be linked to the exact data state and target conditions it used?Run Binding connects an actual AI execution to the identified Release State and conditions Ongoing assuranceRoutes corrections and feedback to the responsible layer and supports continuous improvementCan the team explain what changed, what evidence became stale, and which scope must be requalified?Assurance links actual use, relevant change, and target-specific requalification The third column is the buying test. It does not presume that Alation lacks these controls. It asks whether the exact deployment already supplies evidence at the required level of specificity. Example: an evaluation passes after the underlying data changes Consider a finance agent that answers questions about recognized revenue. Its Alation data product contains the approved tables, business definitions, ownership, lineage, and instructions for handling invoices and currency conversion. The agent has evaluation cases based on recurring finance questions, and the latest suite passes. An upstream regional system then changes its currency format and begins producing more missing invoice dates. Alation Data Quality detects failed checks, lineage identifies the affected assets, and the team updates the relevant metadata or instructions. After the correction, the agent passes its evaluations again. The result is meaningful evidence that the team found the affected layer and restored expected behavior for the evaluation cases. The release team still needs to resolve four questions: Which exact data state was used in the passing evaluation? Did the Proof Run hold the relevant model, prompt, tools, environment, and evaluation conditions constant? Which identified state was approved for operational use? Can the production run be linked to that state and requalified when a relevant condition changes? If the Alation deployment and surrounding data platform already answer all four with auditable records, the existing stack is sufficient for this control objective. If the records show the data product, evaluation, and monitor history but not the target-qualified data state and actual-run linkage, Syntitan can fill that narrower gap. It can preserve the Target Profile, Proof Run Evidence, Qualification Result, Release State, and Run Binding while Alation continues to own governed context, lineage, policy, agent evaluation, and feedback. This coexistence model requires a verified handoff. No native Alation and Syntitan integration was confirmed in the sources reviewed for this Article. The implementation must define how an Alation asset or data product resolves to the Syntitan data-state reference, which identifier persists across systems, who owns approval, and which event triggers requalification. Which platform should lead? Choose Alation as the lead platform when the primary requirement is enterprise data discovery, lineage, governance, reusable data products, semantic context, agent construction, agent evaluation, or cross-layer feedback. It can also satisfy the complete control objective when the surrounding architecture already preserves the required data-state identity, target conditions, release decision, and actual-run evidence. Choose Syntitan for a defined data-state gap when the unresolved requirement is to diagnose Core Readiness, qualify an identified data state for one model or agent target, and carry that result into Operating Evidence through Release State, Run Binding, relevant change, and requalification. Use both with a verified handoff when Alation owns enterprise context, governance, agents, evaluations, and correction workflows while Syntitan owns target-specific data-state Qualification and Assurance. This architecture is justified only if the handoff removes a real evidence gap rather than duplicating records. Keep the existing stack when it already answers the deployment checks end to end. A second platform adds value only when it makes a missing control explicit, repeatable, and auditable. Select one real release process and follow its evidence from the governed Alation data product and evaluation case to the exact data state used by an actual production run. The point where that chain becomes ambiguous should determine the next platform decision. If that ambiguity sits at the data-state level, see how Syntitan connects Core Readiness, target-specific Qualification, and Operating Evidence, then test the framework against the same release process. --- title: "Atlan Data Catalog vs Syntitan: Context Governance or AI-Ready Data State?" url: "https://cubig.ai/articles/atlan-data-catalog-vs-syntitan/" source: live --- An AI agent has passed its context tests and is ready to move into production. In an Atlan data catalog environment, the team knows which definitions, relationships, filters, and source assets the agent should use. It still needs to answer a different question: Is the underlying data state fit for this task, and can the evaluated result be tied to that exact state? That distinction is the most useful way to compare Atlan and Syntitan. Atlan organizes and governs the context around enterprise data and AI. Its current product scope extends well beyond catalog search. It includes lineage, data quality, data products, AI governance, model assets, semantic context, evaluations, and deployment support. Syntitan is CUBIG's AI-Ready Data Platform. For this comparison, its relevant role is to diagnose Core Readiness, qualify a data state for a defined target, and preserve Operating Evidence across release, run binding, change, and requalification. For many organizations, this is not a replacement decision. Atlan can lead when the main requirement is governed discovery and trusted context. Syntitan becomes relevant when the team also needs evidence that a particular data state was fit for a defined model or agent task and that an actual run can be traced to the identified Release State and conditions. The Atlan data catalog already governs more than metadata It would be inaccurate to describe Atlan as a conventional catalog that only documents data at rest. Atlan positions itself as a context layer for AI, supported by an Enterprise Data Graph that connects metadata, semantics, lineage, policies, data products, and usage signals. Its Context Engineering Studio makes that positioning operational. A team can assemble governed Atlan assets into a context repository for a specific use case, define a semantic model, test it with business questions and verified answers, improve it when tests fail, and deploy it to supported targets. Atlan also supports AI model registration, lifecycle stages, linked datasets, evaluation metrics, and governance workflows. This matters because an AI system needs more than access to tables. It needs business meaning, approved relationships, ownership, policy, and a way to determine whether its answers use the intended context. Atlan brings those controls into one metadata-centered operating layer. The comparison with Syntitan therefore starts after acknowledging substantial overlap. Both products address trust, governance, evaluation evidence, and operational use. The difference is not whether one product has governance and the other does not. It is the object being governed and the evidence required at the release boundary. The key distinction is the object being controlled An Atlan context repository references governed assets rather than copying the underlying data into the repository. When the repository is regenerated, it can pull the current state of those linked assets. This keeps context connected to active enterprise metadata and allows changes to definitions, relationships, and policies to flow into the use case. That behavior is valuable for context freshness. It also creates a deployment question that cannot be answered by metadata alone: Which exact data state did the production run resolve to? For example, a repository may correctly identify a customer table, its approved revenue definition, its lineage, and the joins an agent should use. Those controls establish what the data means and how the agent should navigate it. They do not automatically prove that the values presented to a particular run met the task's distribution, completeness, sensitivity, or reproducibility requirements at that moment. Syntitan addresses this second problem through three connected product areas. Baseline assesses General/Core Readiness. Qualification evaluates a data state against a defined Target Profile. Assurance connects the identified Release State to actual runs, relevant changes, and requalification evidence. The distinction can be stated simply: Atlan governs the context that tells an AI system what enterprise data means and how it should be used. Syntitan adds target-specific data-state Qualification and Operating Evidence across release, run binding, change, and requalification. Neither statement makes the other layer optional. Good context cannot repair an unsuitable data distribution. A reproducible data state cannot supply missing business semantics or governance ownership. What Atlan evidence can establish Atlan provides several forms of evidence that matter in an AI control plane. First, lineage can show where an asset came from, what depends on it, and how related metadata propagates. This supports impact analysis when a source, transformation, or policy changes. Second, Data Quality Sensors can run warehouse-native rules for completeness, uniqueness, validity, timeliness, and consistency. Results appear alongside catalog and lineage context, helping teams identify failing assets and understand downstream impact. Atlan describes this capability as post-fact monitoring rather than a pipeline gate by itself, so teams still need to decide how a failed rule affects an AI release. Third, context repositories can carry approved assets and semantic models into a defined use case. Test questions and verified answers provide evidence that the assembled context supports intended business questions. Draft and Active states, plus asset review, create a controlled path toward deployment. Fourth, AI Governance can register models and applications, connect them to datasets, record lifecycle stages and metrics, and apply governance processes. This gives risk, compliance, and platform teams a shared inventory of AI assets and their relationships. Together, these capabilities can answer questions such as: Are the right enterprise assets represented in the use case? Are their meanings, owners, policies, and relationships visible? Does the semantic context answer the business questions it was designed to support? Have monitored data quality conditions or upstream dependencies changed? Which model, application, dataset, and governance records are connected? Those are material controls. A team that already combines them with immutable dataset references, task-specific validation, release approval, and run-level evidence may not need another platform. Where Syntitan's data-state contract begins Syntitan becomes relevant when those final controls are missing or scattered across scripts, pipelines, experiment tools, and approval tickets. Its current framework has three product areas: Baseline, Qualification, and Assurance. Together, they express Syntitan's current operating formula: AI-ready data combines Core Readiness, Target Fit, and Continuous Operating Evidence. Baseline diagnoses General/Core Readiness across six axes: Usability, Integrity, Context, Consistency, Reproducibility, and Traceability. It can be used without a defined target, and improving every Baseline issue is not a prerequisite for Qualification. Qualification assesses Target Fit for a selected model or agent task. Define Target establishes the Target Profile, Target-specific Refinement addresses relevant gaps, Proof Run compares data states under the same AI conditions, and Qualification Result records the judgment from that evidence. It does not claim universal data quality. Assurance preserves Operating Evidence by connecting an identified Data Release and Release State to an actual run through Run Binding, then recording relevant Change Events or Change History and requalifying the affected scope. Release creation and operational-use approval remain separate decisions. This framework addresses questions that are adjacent to Atlan's context controls but not identical: What changed in the actual values or distributions presented to the AI system? Which refinement produced the candidate data state? Did that candidate pass the evidence criteria for this target? Which released state was used by this run? Can the team compare relevant state changes, reconstruct the recorded data state, or requalify the affected scope later? Syntitan should not be positioned as a universal quality score or a replacement for enterprise governance. Its qualification is target-specific, and its evidence is only as meaningful as the target profile, evaluation design, and release policy defined by the team. Compare the operating evidence Comparison of Atlan's documented role, the deployment evidence to verify, and Syntitan's defined role Decision point Atlan's documented role Deployment check Syntitan's defined role Discovery and lineage Connects enterprise assets, ownership, lineage, usage, and metadata in the Enterprise Data Graph Are the intended assets and dependencies visible and governed? Assesses the selected data state for General/Core Readiness and preserves its identity for later target-specific work Business context Builds context repositories from governed assets and semantic models Does the repository represent the definitions, relationships, and questions required by the use case? Uses a Target Profile to define the task, model or agent version, evaluation set, success criteria, and execution conditions Data quality monitoring Runs warehouse-native rules and surfaces results with lineage context Which failures should block, warn, or trigger review for this AI use case? Baseline diagnoses six-axis Core Readiness; Qualification refines gaps relevant to the defined Target Profile AI asset governance Registers models, applications, linked datasets, lifecycle stages, metrics, and policy context Are governance records connected to the deployed system and its approved inputs? Produces a target-specific Qualification Result and preserves separate Data Release and Release State evidence Evaluation Tests context with business questions and verified answers Did the semantic context support the intended questions? Runs a Proof Run under defined AI conditions to generate evidence for the Qualification Result Release Activates reviewed context repositories and deploys semantic models to supported targets Does deployment also resolve linked assets to controlled data references? Records the Data Release and Release State separately from operational-use approval Run evidence Preserves asset, model, lineage, quality, and activity records Can a production result be traced to the exact data state it used? Uses Run Binding to connect an actual AI run to the identified Release State and conditions Reassessment Exposes metadata, lineage, quality, and asset-history changes Which changes require retesting or renewed approval? Compares relevant state changes, records Change Events or Change History, and requalifies the affected scope The third column is the practical buying criterion. It prevents the comparison from becoming a checklist of overlapping features. The real question is whether the existing architecture already answers each deployment check with evidence that is specific enough for the risk and use case. Example: a governed context repository meets changed production data Consider a revenue-analysis agent used by finance leaders. The Atlan context repository contains the approved definition of recognized revenue, certified source assets, relationships between orders and invoices, and verified questions that the semantic model should answer. The repository passes its tests and becomes Active. Before release, one upstream source changes. A regional system begins sending values in a new currency pattern, and the share of missing invoice dates rises. The asset identity and business definition remain correct. The context tests may still pass because the intended relationships and verified answers remain valid for the test set. Atlan lineage and quality monitoring can reveal the upstream change and failed conditions. The team can use that evidence to investigate impact and apply its governance process. If its current platform already blocks the release, snapshots the exact transformed data, reruns target-specific validation, and binds the approved snapshot to the agent run, the operating contract is complete without Syntitan. If those steps depend on manual coordination or cannot be reproduced consistently, Syntitan can fill a narrower gap. The team can diagnose the changed data, refine the candidate state, run the defined Qualification, identify the Data Release and Release State, complete its operational-use approval process, and use Run Binding to connect the approved run to that state. Atlan remains the source of governed context and relationships. Syntitan supplies the data-state evidence required by the release policy. This coexistence model depends on an explicit handoff. Atlan and Syntitan do not have a verified native integration in the evidence reviewed for this article. A team would need to define how an Atlan asset or context repository resolves to the data reference used in Syntitan, how identifiers are preserved, and which system records the authoritative approval status. Which platform should lead? Choose Atlan as the lead platform when the primary need is enterprise discovery, lineage, semantic context, governance, data products, AI asset visibility, or context deployment. It can also be sufficient for the full control objective when the surrounding stack already provides immutable data references, task-specific validation, release gates, and run-level traceability. Choose Syntitan as the lead for data-state Qualification and Operating Evidence when the main gap is determining whether a data state fits a selected model or agent task and connecting its release, actual run, relevant changes, and requalification in an auditable way. This is a narrower decision than selecting an enterprise catalog or governance platform. Use both with a verified handoff when Atlan owns discovery, semantics, lineage, policy, and context repositories while Syntitan owns target-specific data Qualification and Operating Evidence for the relevant data state. Before adopting this design, verify the identifier mapping, approval ownership, immutable reference, failure handling, and reassessment trigger. Keep the existing stack when those controls already exist and can be audited end to end. Adding another platform without removing manual handoffs or closing an evidence gap can create more operational complexity rather than more trust. The most useful evaluation is therefore not a feature count. Select one representative AI use case, follow it from governed context to the exact data state used in a production run, and identify where the evidence chain breaks. That break, if one exists, should determine whether Atlan, Syntitan, both, or neither is the right next step. If your evidence chain breaks at the data-state level, see how Syntitan connects Core Readiness, target-specific Qualification, and Operating Evidence, then test that framework against one real release process. --- title: "When an LLM Evaluation Score Drops, What Actually Failed?" url: "https://cubig.ai/articles/llm-evaluation-failure-attribution/" source: live --- A team is preparing an LLM application for release when its evaluation score drops after the latest change. The release has to pause, but the score does not tell the team where to look. The application code may have broken a field. The model may be responding differently. Retrieval may be supplying different evidence. The test cases, rubric, or judge may have changed. All of those failures can arrive as the same lower number. An LLM evaluation score can flag a problem. It becomes actionable only when the evidence behind it narrows the search to a specific layer. A practical LLM evaluation framework should therefore preserve more than an aggregate result. It should connect each change to the affected cases, traces, data or context, model conditions, and evaluation method. The review should therefore ask more than whether the score went down. It should ask which evidence tells the team what to inspect next. A lower score tells the team to stop, not where to look An aggregate score compresses many observations into one result. That makes it useful as a release gate but weak as a diagnosis. Suppose a retrieval application begins returning answers without required citations after a prompt and pipeline update. A quality score falls. The model may have ignored the instruction, but the pipeline may also have dropped the citation field before the answer reached the evaluator. The retriever may have returned a different set of passages. A scoring prompt may have become stricter. The score alone cannot separate those possibilities. Scoring still matters, but the team also needs a route back to the evidence behind the score. The HELM evaluation framework illustrates why evaluation needs more than one headline number. Its published design covers multiple scenarios and metrics under standardized conditions and releases prompts and completions for inspection. HELM is a research benchmark, not a universal enterprise evaluation recipe. The practical lesson is that a result becomes easier to interpret when the scenario, metric, inputs, outputs, and conditions remain visible. For a product team, a changed score should open an investigation across four possible areas: code and integration, model behavior, data or context, and the evaluation method. If the evidence does not support a narrower conclusion, the honest result is inconclusive. Evaluation score changedPause the release. Open the evidence. ↓ Code and integrationTests, schema, traces, tools Model behaviorIdentity, settings, repeated outputs Data or contextCases, retrieval, source state Evaluation methodMetric, rubric, judge, calibration ↓ Evidence supportsNarrowed causeOr: inconclusive A changed evaluation score should pause the release and open four evidence paths. Code and integration, model behavior, data or context, and the evaluation method are inspected before the team reports a narrowed cause or an explicit inconclusive result. Separate execution integrity from output quality The first check is whether the application completed the intended work at all. Deterministic checks can verify the parts of the system that should not vary. Did every required stage run? Did the response match the expected schema? Were required fields present? Did the retriever return documents? Did a tool call fail? Did the evaluator receive the same output that the user would have seen? These checks do not establish that an answer is relevant, correct, or useful. They establish a basic condition: whether the application delivered a structurally valid candidate for evaluation. Trace-level records can then show where that condition failed. Current MLflow tracing documentation provides one implementation example: traces can preserve inputs, outputs, intermediate spans, tool calls, exceptions, latency, and feedback. Equivalent tracing systems can provide the same evidence. What matters is whether the record lets the team distinguish “the model produced a poor answer” from “the application did not deliver the intended model context or output.” If the integrity checks pass, the team can move beyond basic execution failure. If they fail, the records should narrow the search to code, fixtures, schema handling, retrieval, tool execution, or another application boundary. Compare cases against a pinned reference, not only an average Once the application completed correctly, the team needs to know which cases changed. An average can fall because many cases declined slightly or because one important group failed badly. It can also remain stable while improvements in common cases conceal regressions in rare but consequential ones. Per-case comparisons make those patterns visible. The reference must be identifiable. Record the application revision, model and parameters, evaluation dataset, relevant data or retrieval state, prompt or policy configuration, metric, rubric, judge, and run time. Then compare the changed version against that reference using the same cases and decision rules whenever the comparison depends on them. Official MLflow evaluation-dataset guidance describes datasets that can include inputs, expectations, outputs, and metadata, with production traces or manually curated examples added over time. It also presents golden cases as a way to prevent regressions and compare application versions. This is one current implementation pattern, not a claim that every team needs the same tooling or dataset structure. The comparison may show that a previously passing case now fails. It still does not establish why. The model may have changed, the context may differ, or the judge may have moved. The next step is to inspect the conditions tied to that case rather than treat the difference as proof of causation. Repeated runs can help estimate ordinary variability when the model or scoring method is stochastic. There is no universal run count or acceptable range. The team needs a threshold that reflects its task, observed variability, risk, and cost of a wrong release decision. Treat the judge and rubric as part of the system under test Evaluation methods can introduce their own changes. A rubric defines what counts as acceptable. A judge turns that definition into a label or score. If either one changes, the team is no longer measuring under the same conditions, even when the application output is identical. LLM judges can make large-scale review more practical, but their scores are not ground truth. Chen and colleagues found that human and LLM judges in their study were vulnerable to several tested biases and perturbations. A separate study by Thakur and colleagues found gaps between model judges and humans, along with sensitivity to prompt complexity and length and a tendency toward leniency. These findings belong to the models, tasks, and methods studied. They do not show that every LLM judge is unusable. These findings support a stricter practice: store the judge model and version, the judge prompt, the output it evaluated, the rubric, and any calibration evidence. Compare automated verdicts with human labels on cases that reflect the actual task and risk. Recheck that relationship when the judge, rubric, traffic, or application changes. Human review also needs a stated role. Reviewers can disagree, overlook rare failures, or apply an unwritten standard inconsistently. Both human and model-based review need explicit criteria, and teams should document disagreements rather than hide them in an average. High-risk or uncertain cases should follow the review path the organization has approved. Build an attribution record for every LLM evaluation change The records may live in different platforms, but they still need shared identifiers that connect them to the same comparison and release decision. Evidence boundaries for investigating an LLM evaluation change Suspected layerEvidence to inspectWhat it can narrowWhat it cannot prove Code and integrationRevision, deterministic tests, schema checks, traces, tool and retrieval errorsWhether the intended application path completed and where a recorded failure occurredThat a structurally valid answer is correct or useful Model behaviorModel and parameter identity, repeated case outputs, response distributionWhether outputs changed beyond the comparison range under recorded conditionsWhy they changed when code, context, or evaluation conditions also moved Data or contextEvaluation-set version, retrieved records, source or index state, transformations, permissionsWhether the run received different evidence or a different data stateThat the data difference caused the score change without a controlled comparison Evaluation methodMetric, rubric, judge identity and prompt, human labels, calibration recordWhether the measurement or decision rule changedProduction quality outside the cases and conditions evaluated InconclusiveMissing or conflicting identifiers across the four layersWhich evidence must be recovered before another claim is madeA defensible root cause or release decision Use this table to route the investigation, not to score the team's maturity. Several areas may remain active at once. A retrieval defect can change the context and expose a weakness in the rubric. A model update can produce a real gain while a judge update makes the aggregate score look worse. The record should preserve those possibilities until the evidence narrows them. NIST's AI Risk Management Framework treats test, evaluation, verification, and validation as contextual practices that should be documented and applied across the AI lifecycle. Its Generative AI Profile extends the risk-management context for generative systems. Neither document prescribes this table or one LLM harness. They support the broader requirement to connect measurement to context, evidence, and an accountable decision process. Use a Proof Run only for the bounded data-state question Data and context are one part of the investigation, not the entire explanation. Within CUBIG's operating model, Qualification asks whether a defined data state is ready for a selected model or agent task. A Proof Run then compares results under controlled, recorded conditions to test whether a data-state change affects the target result. That can be useful when the team has already narrowed the unresolved question to data. For example, two evaluation runs may use the same model, application revision, prompt, cases, metric, and judge but receive different retrieval results or dataset states. Comparing those states under the recorded target can show whether the data-side hypothesis deserves further action. The result must remain bounded. A different outcome under a controlled data-state comparison supports data as a plausible explanation, but it does not prove that data was the only cause of the original incident. An unchanged outcome points the investigation toward another layer, but it does not certify the model, code, or judge. Syntitan supports this bounded investigation by preserving the data state used in each test and binding that state to the corresponding run. This makes it easier to compare the data behind two results. It does not replace application traces, model records, rubrics, judge calibration, or human review. A successful data test is still not release approval for the entire LLM application. The evaluation team must separately review the model, application, judge, and operational risk. Start with the smallest evidence set that can change a decision A team does not need to build a complete evaluation platform before reading its first outputs. Start with a small set of cases that reflect the work the application must perform. Write down why each result is acceptable or unacceptable. Record the application version, model, data or context, and evaluation conditions for the run. When a score changes, inspect the affected cases and traces before changing the prompt, model, or retrieval system again. The evaluation record should answer five questions: Did the intended application path complete? Which cases changed against the pinned reference? Were the model, data or context, and evaluation conditions comparable? Does the judge have calibration evidence for this task? What remains unknown, and does that uncertainty block release? Those answers turn an LLM evaluation into a decision aid. They can route the team toward a code fix, a model investigation, a data-state comparison, an evaluation-method correction, or an explicit inconclusive result. The goal is to prevent one lower score from sending every team into the same unfocused investigation, while still allowing an inconclusive result when the evidence is incomplete. If the unresolved question is whether a data-state change affected the result, use Syntitan to run a controlled comparison and preserve the data behind each result. --- title: "Why LLM Benchmark Rankings Do Not Transfer to Your Work" url: "https://cubig.ai/articles/llm-benchmark-rankings-private-evaluation/" source: live --- The model at the top of a public leaderboard can still be the wrong choice for your business. Its score measures performance on a particular test, with a particular answer key and evaluation harness. By itself, it cannot show whether the same ordering will hold for your contracts, support tickets, reports, or agent workflows. Public scores remain efficient screening tools. They help teams remove weak candidates and compare broad capabilities before paying for deeper evaluation. The problem begins when a shortlist becomes the final selection. Use public LLM benchmark results to build a shortlist, then test the finalists on held-out work that reflects your own inputs, labels, and business rules. What an LLM benchmark actually tells you Every LLM benchmark is built around a specific task. Its designers choose the inputs, expected answers, scoring method, prompt format, and aggregation rule. A model's result belongs to that complete setup. A production workload rarely matches it exactly. Requests may be incomplete, domain terms may be ambiguous, and the correct response may depend on internal policy or current records. Inputs can arrive through OCR, exports, retrieval systems, or application fields that do not preserve the clean structure of a benchmark question. The mismatch matters because model strengths are uneven. One model may be excellent at self-contained reasoning questions but weak when it must follow a company-specific response policy. Another may score lower overall yet perform better on the document types and failure costs that dominate your workload. A public average hides that distinction. It combines many tasks and gives each one the weight chosen by the benchmark designer. Your organization needs a different weighting because a rare error in a regulated review can matter more than dozens of correct low-risk summaries. Before comparing ranks, ask a more basic question: how closely does the benchmark resemble the work the model will actually do? Public benchmark 1Candidate A2Candidate B3Candidate C ↓ Task mixInput stateCost of error ↓ Private workload 1Candidate C2Candidate A3Candidate B A public benchmark ranks Candidate A first. After the candidates are tested against the organization's task mix, input state, and cost of error, Candidate C ranks first on the private workload. An LLM benchmark score reflects the test as well as the model Even when the tasks are relevant, the score still reflects how the test was administered. Answer keys can contain errors. In Are We Done with MMLU?, researchers reannotated 5,700 questions across 57 subjects and estimated that 6.49 percent contained errors. In the virology subset they analyzed, 57 percent of the questions were flawed. Those findings apply to that audit, not to benchmarks in general. They show why a precise score can still rest on disputed labels. Administration choices can also move the result. Research on multiple-choice selectors found that models can prefer particular option identifiers, so changing the order of answer choices can affect performance. A separate prompt-design study measured swings of up to 76 accuracy points for LLaMA-2-13B across meaning-preserving prompt formats in tested few-shot settings. The 76-point swing applies only to that model and tested setup. The study shows that spacing, separators, casing, and other harness details can become part of the measured outcome. Record every score with the prompt template, parsing logic, option order, metric implementation, software version, and other conditions that produced it. Without that record, two published numbers may look comparable when the tests were not administered in the same way. Public benchmarks can reward familiarity Once an LLM benchmark is widely available, its questions and conventions can influence model development. That can make the score less informative when the model is moved to unfamiliar work. The SWE-Bench Illusion introduced diagnostic tasks that asked models to identify buggy file paths from issue descriptions without repository structure. The tested models reached up to 76 percent accuracy on repositories included in SWE-Bench and up to 53 percent on repositories outside it. The results raise concern about possible contamination or memorization, but they do not establish that memory explains the entire gap. Repository selection or naming conventions may also contribute. Fresh versions of familiar tests provide another useful check. In the GSM1k study, researchers created new grade-school math problems designed to match the style and difficulty of GSM8K. Several model families lost up to 8 percent, while the frontier models tested showed little change. The results varied by model family, which means a ranking can change when the test changes. Benchmark designers are responding to this problem. LiveBench uses recently released sources and updates questions regularly to limit contamination. That makes it useful for current research comparisons, although it still cannot reproduce a private organization's task distribution, input condition, policies, or cost of error. Use public rankings as a shortlist, not a verdict An LLM benchmark is most useful at the start of model selection. Choose tests that cover the capabilities you need, then review their version, task mix, scoring method, harness, and known limitations. Use the results to narrow the field to a manageable set of candidates. The shortlist identifies which candidates deserve further testing. The final decision depends on which candidate performs best under your conditions, so the next step is a private test built around representative work. Do not reduce the private test to one undifferentiated average. Report the slices that correspond to different tasks, data conditions, user groups, and risk levels. A model that wins overall but fails the highest-cost slice may be the wrong choice. A close aggregate result may also be inconclusive if the test set is too small or ordinary run variation is larger than the observed gap. Build a private transfer test around the decision A private evaluation set should reflect what the model will receive and what the organization needs it to produce. Select cases from real workload categories, then remove or protect sensitive information according to the approved data policy. Keep the conditions in which those inputs actually arrive, including missing fields, OCR noise, inconsistent formats, and retrieval context when they affect the task. Define acceptable outputs with the people who own the business consequence. Where conflicts of interest matter, keep model selection separate from label adjudication. Reserve part of the set for final evaluation rather than routine development. If every case becomes a prompt-tuning target, the private test loses the independence that made it useful. Version the set, document changes, and refresh it when the workload or decision standard changes. Evidence required to transfer an LLM benchmark ranking to a private workload Decision elementWhat the public benchmark showsWhat the private test must addRisk if missing Task mixThe benchmark's published categories and weightsThe organization's actual task distribution and risk weightingThe winning model solves the wrong mix of problems Input stateCurated benchmark inputsDocuments, fields, and retrieval context in their real operating conditionClean-input performance is mistaken for production robustness LabelsPublic answer keys or preference dataTask-specific acceptance criteria and adjudicated examplesThe score rewards an answer the organization cannot accept HarnessPublished prompt, parser, metric, and runtime detailsA fixed application path and evaluation procedureHarness changes are mistaken for model differences Slice resultsBenchmark category scores when availableResults by task, data condition, user group, and failure costA safe average hides a critical weak point Run identityBenchmark version and reported model identityModel, code, data, prompt, judge, environment, and timestamp for every runThe comparison cannot be reproduced or explained The right number of cases, metrics, reviewers, and repeated runs depends on the risk of the decision and the variability observed during testing. If the evidence cannot distinguish two candidates, report the result as inconclusive and run another test. Do not manufacture a winner from a narrow margin. Preserve every run, not only the best one Teams can bias a private evaluation without intending to. They may try several prompts, model versions, retrieval settings, or judge configurations and retain only the most flattering result. Define the comparison plan before the final test. Record every material configuration and result, not just the best one. If the plan changes, document why and treat the next result as a new comparison. This record also makes later revalidation possible. A model, data source, retrieval index, prompt, or evaluation rule may change after selection. The organization needs to know which configuration produced the accepted result before it can decide whether the evidence still supports use. Keep the data side of the test reproducible Model evaluation depends on more than the model and prompt. The data supplied to each candidate must also be identified and held constant. When data preparation is the variable under review, a controlled Proof Run can compare results produced from recorded data states. Syntitan supports this part of the process by documenting the intended task, the data state approved for the test, and the run that used it. This record helps a team determine whether a result changed because the model changed or because it received different data. The evaluation team still owns the benchmark design, expert labels, application settings, and final model choice. If the test changes the data, prompt, model, and judge at the same time, the final score cannot explain which change mattered. Holding the other conditions steady makes the comparison easier to interpret and defend. Public LLM benchmarks are a practical place to build a shortlist. Before making a purchase or deployment decision, test the finalists on held-out work that reflects your inputs, acceptance criteria, and cost of failure. Keep the data, configuration, and results from every run so the decision can be revisited when conditions change. If you need to test how data preparation affects model performance, explore Syntitan. --- title: "Dataiku vs Syntitan: What Evidence Does an AI Run Need?" url: "https://cubig.ai/articles/dataiku-vs-syntitan-ai-data-state/" source: live --- An AI result changes between review and production. The team can inspect its Dataiku project, including the active bundle, Flow, model record, code environment, scenario history, and evaluation results. The remaining question is whether both runs used the same approved data state. That question provides a useful starting point for comparing Dataiku and Syntitan. Dataiku is a broad enterprise platform for building, deploying, monitoring, and governing analytics, machine-learning systems, and AI agents. Syntitan addresses a narrower problem: defining which data state is qualified for a selected AI task, binding that state to a run, and preserving evidence for later comparison and requalification. A well-configured Dataiku environment can preserve substantial operational evidence, including selected data. The comparison therefore depends on what those artifacts prove and whether the operating process requires a task-qualified data state. Dataiku already records more than the workflow Dataiku covers considerably more than AutoML or visual workflow design. Its Flow connects datasets and recipes and records lineage. Git-based project version control provides history, comparison, and revert for project configuration. Deployment bundles move project versions from a Design node to an Automation node, and an earlier bundle can be activated again. Dataiku can also retain evidence across the wider AI lifecycle. Versioned code environments can be linked to bundles, with older versions retained for rollback. Data quality rules preserve results and history, while drift analysis compares input data with a reference dataset. MLflow-based experiment tracking can record parameters, metrics, models, and artifacts. Dataiku also documents agent evaluation, interaction logging, and governance workflows for assessment, traceability, review, and oversight. Depending on how they are configured, these capabilities may already answer most of a team's reproducibility and investigation questions. The remaining comparison should begin with the evidence the deployment already preserves. The boundary is what the configured artifacts preserve A Dataiku project bundle is primarily a deployment and replay artifact. It always contains project metadata. Teams can also include selected datasets, saved models, and managed-folder contents. Versioned code environments can be associated with the bundle. But the contents must be checked, not assumed. Dataiku's documentation states that actual data, persisted models, and global shared code are not all included by default. Reverting a project through Git changes its configuration, not its data. A bundle that includes data can restore that included data, while a bundle that reads from an external source may encounter a later state of that source. Notebook kernels can also change at runtime outside the bundle-linked environment controls. So a Dataiku implementation may provide a strong reproduction record. The strength of that record depends on configuration: Was the relevant feature table included in the bundle? If not, was an external snapshot, table version, or immutable object identifier recorded? Were the model and code-environment versions linked to the deployment? Did evaluation or interaction logs preserve the inputs and outputs needed for diagnosis? Did drift analysis compare the production input with the approved reference? If those answers are complete, the remaining evidence gap may be narrow or may not exist. If they are incomplete, the team should identify the missing object before adding another tool. Syntitan defines a task-specific data evidence contract CUBIG describes Syntitan through three connected domains: Baseline, Qualification, and Assurance. Baseline diagnoses a common data state across six readiness axes. It establishes where the data stands before a specific target is applied. Qualification evaluates data for a selected model or agent task. Its sequence is Target Profile → data refinement → Proof Run Evidence → Qualification Result. The Proof Run tests whether the refined state is fit for the defined target. It does not claim universal data quality. Assurance carries the approved result into operations through Release State, Run Binding, Change Event or History, and Requalification. The purpose is to retain the relationship between a released data state and the AI run that used it, then provide a controlled basis for reassessment when relevant change is recorded. These terms define the intended evidence responsibilities, but they do not prove every implementation detail. CUBIG's materials support Release State and Run Binding as core concepts, with Diff and Reproduce as supporting evidence concepts. They do not confirm current UI or API availability for Diff or Reproduce. They also do not support claims of automatic change detection, one-click full-system restoration, or guaranteed performance. Models, prompts, code, dependencies, seeds, runtimes, and evaluation protocols still require their own records. Dataiku can hold both workflow and data evidence. Syntitan becomes relevant when the operating process requires an independently identifiable, task-qualified data state to be approved, bound to the run, and used as the basis for requalification. Compare the investigation record Dataiku and Syntitan investigation evidence comparison Decision pointWhat Dataiku can provideWhat must be verified in the deploymentSyntitan's defined role Project and logic stateGit history, Flow lineage, project comparison and revertProject revert restores configuration, not dataOutside Syntitan's primary scope Deployment packageProject bundle, prior-bundle activation, selected datasets, models, and managed foldersExact bundle contents and external dependenciesRelease State identifies the data-state reference carried into Assurance Runtime environmentBundle-linked versioned code environmentsNotebook or runtime changes outside enforced versioningOutside the data-state contract Data behaviorQuality-rule history and reference-based drift analysisReference choice, retained values, snapshot IDs, and coverageBaseline diagnoses common readiness; CUBIG describes Diff as a bounded comparison concept Model or agent evidenceExperiment tracking, evaluation records, interaction logs, and tracesWhat was logged and whether row-level inputs and outputs remain availableProof Run Evidence supports task-specific Qualification Approval and accountabilityGovernance workflow, review, and sign-offWhether approval covers the exact data state used by the runQualification Result records the data-fit decision; operational approval remains separate Operational reassessmentMonitoring, recorded changes, and configured rollback pathsWhether a change triggers the required review processRun Binding, change history, and Requalification connect the released state to ongoing use The table is a procurement and operating checklist, not a ranking. Dataiku covers a much broader system. Syntitan's relevance rises only when the missing control is the task-qualified data state and its connection to a run. Diagnose the changed result before expanding the stack Return to the AI result that changed between review and production. First, inspect the Dataiku evidence already available. Confirm the active bundle and project revision. Check the model and code-environment versions. Review scenario history, experiment records, evaluation results, and interaction logs. Compare the production inputs with the approved reference through existing quality and drift controls. Next, establish the data identity. Was the feature table included in the bundle? If the data remained external, does the run record contain a snapshot ID, table version, object version, or equivalent immutable reference? Can that reference recover the values and preparation state that passed review? Only then ask whether Syntitan fills a real gap. If Dataiku's configured artifacts already identify and restore the approved data state, adding another evidence layer needs a different justification. If the team can recover the project, model, and environment but cannot identify the task-qualified data state used by the run, Syntitan's Qualification and Assurance model becomes relevant. The controlled test is straightforward in concept. Restore or select the earlier approved data state while holding the model, code, runtime, prompt, seed, and evaluation protocol constant. If the result returns to the prior range, the data-side explanation gains support. If it does not, move the investigation to the other recorded layers. Neither product can make that test conclusive when the surrounding evidence is missing. The value comes from making the evidence boundary explicit before an incident. Which platform should lead? Dataiku should lead when the organization needs a collaborative platform to prepare data, build and deploy models or agents, automate production work, monitor behavior, and govern an AI portfolio. Evaluate the actual edition and architecture in scope, because the available record depends on what the team configures and logs. Syntitan should be evaluated when the unresolved control is more specific: Is this data fit for this selected model or agent task? Which approved data state did this run use? What changed between the approved and current states? What evidence should trigger requalification? The two roles can coexist. Do not assume an integration until identifiers, ownership, and handoffs have been verified for the proposed deployment. The decision comes down to one question: Which evidence object is still missing when an AI result changes? If the missing object is the project, model, deployment, or governance record, investigate Dataiku's configured controls first. If it is the task-qualified data state bound to the run, examine how Syntitan defines Qualification and Assurance. --- title: "Why AI Model Accuracy Expires and How to Revalidate It" url: "https://cubig.ai/articles/model-performance-degradation-revalidation/" source: live --- An internal model review reports 94% accuracy. The measurement may be sound, but it may describe a data window, customer population, or operating process that no longer exists. A performance score reflects specific conditions: a target, dataset, time window, metric, and acceptance rule. When those conditions change, the score remains a valid record of the original evaluation, but it no longer tells you how the model performs today. Model performance degradation is a measurable decline under the metric and acceptance rule a team uses. It can result from changes to the model, input data, relationship between inputs and outcomes, or evaluation conditions. The right response is not to discard every older result. It is to date the score, preserve the evidence behind it, and define when to test it again. A model score belongs to a point in time Teams often speak about validation scores as if they were permanent properties: this model has 94% accuracy. The more accurate statement is that the model achieved 94% on a specific evaluation set, under a specific metric, at a specific time. That distinction matters because the environment around a deployed model keeps changing. New records arrive. Customer behavior shifts. Equipment and policies change. Outcome labels mature after predictions have already been made. The model artifact may remain fixed while the population represented by the next evaluation window moves further from the population used to establish the accepted score. Vela and colleagues studied temporal quality degradation across 32 datasets from healthcare operations, transportation, finance, and weather using four standard machine-learning model types. They observed degradation in 91% of the 128 model and dataset pairs. This result belongs to their experimental design. It is not a universal failure rate or a forecast for any particular production model. The study also found several forms of degradation, including gradual decline, abrupt deterioration, and increased variability even when average performance appeared relatively stable. The practical conclusion is narrower. As time passes, the evidence needed to support a performance claim changes. Every accepted score therefore needs an as-of date and a clear route back to the conditions that produced it. A recorded model score needs current evidence before it supports a later decision An accepted score is tied to an evaluation window. As time passes, reviewers must recheck the data state, outcome labels, and evaluation rule before deciding whether the score remains supported or requires revalidation. Accepted evaluationScore recorded with an as-of dateThe result is valid for the window and conditions that produced it. Time passes Current evidence checkDoes the score still support today's decision?Data stateOutcome labelsEvaluation ruleCurrently supportedRevalidation due Data movement does not automatically mean model damage An input monitor can show that a distribution moved. On its own, it cannot tell you whether the model's decisions became worse. The two can diverge in either direction. Inputs may change while performance remains within the acceptance rule. Performance may decline because the relationship between inputs and outcomes has changed, even when the observed input distributions move only slightly. A metric may also appear stable because the labels required to calculate current performance have not arrived yet. This is why data drift, concept drift, and model performance degradation should not be used interchangeably. Data drift is a change in the distribution of model inputs. Concept drift is a change in the relationship between inputs and the target. Model performance degradation is a decline in the result the team measures under its current metric and acceptance rule. Rabanser and colleagues studied methods for detecting dataset shift and characterizing whether a detected shift was harmful across their selected datasets and perturbations. Their work supports shift detection as one useful source of evidence. It does not make every alert proof of performance loss. A quiet dashboard does not prove that the model remains accurate either. A drift alert should therefore trigger an investigation, not settle it. The team still needs current evaluation evidence before it can decide whether the score remains fit for use. Build a validity record for every score Every reported score should include enough context for another reviewer to understand what it represents. A slide or dashboard does not need to contain every artifact, but it should point to records that another operator can inspect. Evidence required to judge whether a model score is still valid Evidence fieldWhat must be dated or identifiedReview questionIf it is missing Result identityTarget output, metric, threshold, and acceptance ruleAre reviewers judging the same result?The score can change meaning while keeping the same label Evaluation windowStart and end dates plus population representedDoes this window still match the current decision?An old score can be presented as current evidence Data stateRecords, schema, transformations, split, and released-state referenceWhich exact data produced the score?The team cannot compare or restore the evaluation input Label maturityCoverage, delay, missingness, and cutoff ruleIs current performance measurable yet?A quiet metric can hide incomplete outcomes Executable pathModel, code, preprocessing, and environment versionsDid the same evaluation process run?Two scores may reflect different execution conditions Revalidation triggerPlanned review date and material change eventsWhat requires the claim to be tested again?The score remains in circulation without an expiry boundary Consider a hypothetical churn model approved using data available through March. In June, the team continues to report the same validation score. However, it cannot determine which customer segments have complete outcome labels, whether the feature pipeline has changed, or when the next evaluation on later data is due. This record does not prove that the model is failing. It also does not support the claim that the March score describes performance in June. The defensible status is not “bad model.” It is “current validity unresolved.” That distinction determines the next action. The team does not need to retrain automatically. It first reconstructs the evaluation conditions, identifies any limits created by incomplete labels, and measures performance on a current window using the same acceptance rule. Evaluate a later time window Random train-test splits are useful for many development questions, but they can mix older and newer observations. A production decision needs a test that respects the order of time. The team should train or select the model using what would have been known at one point, then evaluate it on a later window. The Wild-Time benchmark was designed around temporal distribution shifts. It contains five datasets and compares thirteen approaches using fixed-time and streaming evaluation protocols. Across its benchmark settings, the authors reported an average 20% performance drop from in-distribution to out-of-distribution data. That figure describes Wild-Time. It is not an expected decline for every deployed model. The broader lesson is about test design. If a model is expected to operate on tomorrow's cases, its review should include evidence from a later period. A shuffled sample of yesterday's data may answer a development question without showing whether the accepted score still applies to current conditions. A revalidation cycle should preserve four comparisons: the previous accepted score and the conditions under which it was measured; the current score under the same metric and acceptance rule; the material differences between the previous and current data and evaluation states; and the label coverage and uncertainty that limit the comparison. There is no universal revalidation schedule. A fraud model with quickly maturing outcomes and a clinical model with long-delayed outcomes cannot follow the same calendar. The schedule should reflect the risk of the decision, how quickly relevant conditions change, when reliable labels become available, and the cost of relying on stale evidence. Decide what the current evidence supports A review can end in three defensible states. Currently supported. The current evaluation window is sufficiently representative for the decision, the acceptance rule still holds, and no unresolved condition materially weakens the comparison. Revalidation due. Enough time has passed, or a material condition has changed, so the previous score no longer supports the current decision on its own. The next step is to restore the evaluation record and run the agreed test on current data. Claim must be qualified. Incomplete labels, an unresolved prior data state, a changed metric, or another material condition prevents a valid comparison. The team can use the score only within that stated uncertainty. If the decision cannot tolerate the uncertainty, the score should no longer be used for that purpose. These states are more useful than a generic green or red indicator. Each one connects the available evidence to a decision and makes the next action explicit. Where Syntitan fits Part of the validity record belongs to the data layer. A reviewer needs to identify which records, schema, transformations, and released data state reached each evaluation. In CUBIG's operating model, this evidence forms part of Assurance, the record created around actual use and requalification after a material change. Version history can support Assurance, but it is not sufficient by itself. The record must also identify the actual run and the conditions bound to it. Syntitan supports this part of the evidence path. Release State provides a fixed reference for the data state. Run Binding connects an evaluation or AI run to the Release State used for that run. Diff compares two Release States to narrow down what changed on the data side. These controls do not measure live model accuracy, determine whether a shift is harmful, choose a retraining schedule, or replace the systems responsible for model evaluation. They make the data-state comparison traceable. The model artifact, code, metric, labels, and acceptance rule still require records from the systems that own them. This boundary makes the investigation more precise. If a result changes while the data state remains fixed, the team can focus on the model, execution environment, labels, or evaluation process. If the data state changed, the team can identify and test that difference instead of relying on dates and memory. Give every score an as-of date Choose one model score that still appears in a production dashboard, approval memo, or quarterly review. Ask another operator to locate its evaluation window, data state, label maturity, metric definition, acceptance rule, and next revalidation trigger. The first field that cannot be resolved marks the boundary of the claim. Resolve that gap before the score is used in another decision. Make the data state behind each evaluation traceable. See how Syntitan connects released data states to AI runs. --- title: "Same Seed, Different Result: What LLM Reproducibility Requires" url: "https://cubig.ai/articles/same-seed-llm-reproducibility/" source: live --- The rerun should have been routine. The team loaded the same model, set the temperature to zero, reused the saved seed, and ran the same evaluation command. The score still moved. Nothing obvious explains the gap. The code commit matches. The configuration looks complete. Yet the team can no longer tell whether the earlier number will hold in a release, an audit, or even another run next week. The problem is not necessarily a forgotten setting. A seed controls a source of randomness that the software exposes to it. It does not identify the data version, reconstruct preprocessing, pin every dependency, define the evaluation protocol, or control how numerical operations are scheduled across hardware. LLM reproducibility depends on the full set of conditions that produced a result. The seed is one of those conditions, not an identifier for the run. A seed controls a random process, not the whole run Teams use seeds for a good reason. When a framework routes random operations through a seeded generator, the same seed can help reproduce sampling, parameter initialization, data shuffling, or another stochastic step. The exact scope depends on the framework and workflow. Two executions can share a seed while reading different records, applying different preprocessing, loading different library builds, using different evaluation splits, or running on a different serving configuration. Reviews of machine-learning reproducibility trace the problem across experiment design, method, code, data, implementation, and reporting. The implication for an operating team is direct: “same seed” answers only one question. It does not establish that the two executions are comparable. Before comparing results, the team needs to know what the seed actually controlled and which other conditions were held constant. Run state disappears in more than one place Some conditions vanish because nobody wrote them down precisely. Hyperparameters, prompts, preprocessing logic, data splits, metric definitions, stopping criteria, and the exact data state may live partly in a configuration file and partly in a teammate's memory. The file helps only when it resolves the versions and artifacts the execution actually used. Other conditions hide in the environment. Framework versions, tokenizer builds, drivers, operating system libraries, container base images, and compiler settings are easy to treat as background. The application code can remain unchanged while one dependency changes what reaches the model or how the model executes. The least visible conditions sit in the physical and numerical execution path. GPU type and count, batch construction, kernel choice, parallel reduction order, and numerical precision may shape the calculation. These details are not always exposed in the application configuration. With a hosted service, some may never be under the team's control. What the team should do depends on which layer is missing. Record the known experiment conditions, package the software environment where possible, and either control the numerical execution path or state its limits. Greedy decoding can diverge without random sampling Temperature zero makes this limitation easier to see. With greedy decoding, there may be no sampled token for a seed to stabilize, yet outputs can still diverge. Floating-point arithmetic is sensitive to operation order because intermediate results are rounded to a limited precision. Parallel hardware can group reductions differently as batch size, GPU count, or kernel execution changes. The underlying model and input can stay the same while a small numerical difference appears in a token score. Greedy decoding then turns a continuous score difference into a discrete choice. If two tokens have nearly equal scores, a small numerical change can switch which token ranks first. The selected token becomes part of the next input, so an early difference can send the continuation down another path. Yuan and colleagues demonstrated this mechanism in a controlled study of LLM inference. For DeepSeek-R1-Distill-Qwen-7B on AIME’24, BF16 precision and greedy decoding produced an accuracy standard deviation of 9.15 percentage points across 12 runtime configurations. The same experiment reported an average output-length standard deviation of 9,189.53 tokens. These results apply to the tested model, benchmark, precision, and configurations; they are not a universal variance rate for LLMs. The study matters because it shows why a seed cannot be the universal remedy. The relevant difference may come from numerical execution rather than a random draw. System-specific work on deterministic inference therefore moves the control lower in the stack, toward batch-invariant operations and kernels, with engineering tradeoffs of its own. How a small numerical shift can fork greedy decoding The same prompt, seed, and greedy decoding can produce nearly tied token scores. A small numerical shift can reverse their order, select a different first token, and send the continuation down a different path. Same promptSame seedGreedy decoding Nearly tied scoresA small numerical change can reverse the order. Token AToken B Token A selectedThe next step includes Token A. Token B selectedThe next step includes Token B. Consequence: the selected token becomes part of the next input, so the continuations can diverge. Build an evidence envelope around the result A reproducibility record should tell a second operator which conditions must match, which artifacts can be restored, and which conditions remain outside the team's control. Evidence layers, questions, records, and unresolved reproducibility boundaries Evidence layerExact questionRecord or controlBoundary if unresolved Result identityWhich output or metric are we trying to reproduce?Output artifact, metric definition, acceptance rule, timestampTwo runs may be compared against different targets Data stateWhich records, schema, references, and split reached the run?Versioned data reference, manifest, content hash, split identityThe same pipeline name may resolve to different inputs Executable logicWhich model, code, prompt, and preprocessing ran?Model identifier, code commit, prompt or template version, preprocessing artifactA matching model name does not identify the full executable path RandomnessWhich stochastic operations were active, and what did the seed control?Seed, decoding settings, framework-specific deterministic settings“Same seed” may cover only part of the workflow EnvironmentWhich dependencies and system libraries shaped execution?Lockfile, image digest, driver and runtime versionsRebuilt software may not match the original environment Numerical executionWhich precision, devices, batches, and kernels were used?Precision, GPU type and count, batch policy, kernel and determinism settingsIdentical output may be impossible to promise across the boundary EvaluationWhich examples, judge, metric, and aggregation rule produced the score?Evaluation-set version, judge version, metric implementation, aggregationA score can move even when the model output is unchanged Consider a hypothetical support-ticket routing evaluation that runs through a hosted LLM API. The dataset hash, prompt version, seed, and scoring script all match the prior run, but the aggregate score slips. If the provider does not expose the serving batch or hardware path, the team cannot honestly call the executions identical. The team can state that the application-controlled conditions matched while the infrastructure boundary remains unresolved. That conclusion is narrower, but another reviewer can evaluate it. Teams do not need to freeze every variable forever. They do need enough evidence to state what a result proves. If the data, code, environment, execution path, and evaluation all match closely enough for the decision, the rerun is a meaningful reproducibility test. If one layer cannot be matched, the result may still be useful, but the comparison must state that limitation. The evidence should say what changed rather than hiding the difference behind a shared seed. Reproduction does not always mean byte-identical output The required level of sameness depends on the decision. A debugging team may need the same failure to reappear. A model-selection team may need rankings and metrics to remain within a defined tolerance. A regulated review may need a traceable explanation of the inputs and conditions even when the serving provider cannot expose every hardware detail. These are different acceptance rules. Declaring a rerun “reproduced” without naming the rule creates another ambiguity. The team should define the result identity and tolerance before rerunning. It should also separate three outcomes: The result matches under the stated conditions and acceptance rule. The result differs, and the evidence identifies a changed condition that can be tested. The result differs across a condition the team cannot reconstruct or control, so the original claim must remain qualified. This makes “reproduced” a testable claim rather than a blanket promise. Syntitan covers the data-state layer Once the run is treated as a set of linked conditions, responsibility across the stack becomes clearer. Model registries, code repositories, experiment trackers, and serving systems hold different parts of the execution record. The data state needs equally precise versioning. Syntitan addresses that data-state layer. A Release State fixes the approved data state for an AI task. Run Binding connects an execution to that state. Diff identifies changes between released states, and Reproduce restores a recorded state for investigation. This lets the team identify the data state used by the run and isolate changes on the data side. It does not freeze the model, application code, prompt, dependencies, GPU topology, serving batch, or evaluation system. Those records must come from the systems that own them. This boundary is useful during an investigation. If a team restores the same released data state and holds the other material conditions constant, it can test whether a data change contributed to the result. If the output still moves while the data state remains fixed, the evidence directs attention toward the remaining model, software, infrastructure, or evaluation conditions. Neither outcome proves root cause automatically. It gives the team a way to stop treating the run as one opaque configuration and test one evidence layer at a time. The handoff test for an AI result Take one result that matters to a release, customer decision, or internal approval. Give its run record to someone who did not produce it. Can that person identify the target output, restore the data state, resolve the model and code, rebuild the environment, understand what the seed controlled, match or qualify the numerical execution, and apply the same evaluation rule? The first layer that another operator cannot reconstruct defines both the next task and the limit of the original claim. A result becomes defensible when another operator can reconstruct it within a declared boundary, not when one familiar setting happens to match. Keep the seed in the record. Just do not ask it to carry the whole run. Make the data state behind each run part of the reproducibility record. See how Syntitan connects an approved data state to an AI run. --- title: "Databricks Time Travel and Lakebase vs Syntitan: What Each One Proves" url: "https://cubig.ai/articles/databricks-time-travel-lakebase-vs-syntitan/" source: live --- When an AI result changes, “go back to the data that produced the earlier result” sounds like a straightforward instruction. In a Databricks environment, it is not one operation. The relevant history may sit in a Delta table, an operational Lakebase database, or the execution record for the model or agent. Databricks Time Travel, RESTORE, Lakebase point-in-time branching, and MLflow each recover or record a different part of that history. Syntitan addresses a related but separate question: which data state was approved for the AI task, and how was that state connected to the run under review? The distinction matters because historical access and run-level evidence are not interchangeable. Databricks can help a team return to an earlier table or database state. Syntitan is designed to qualify the data for a selected AI task and preserve the path from Release State to Run Binding, Diff, and Reproduce. A production investigation may need both layers. What Databricks Time Travel establishes Databricks table history records operations that modify supported Delta Lake and Apache Iceberg tables. Time Travel can query a table by version or timestamp, allowing a team to inspect an earlier state without replacing the current table. RESTORE performs a different job. It returns a table to a previous version by creating a new current state based on that historical version. Querying an earlier version supports inspection and comparison. Restoring changes the table that downstream work will see. Both operations depend on retained history. Databricks notes that the required transaction log and data files must still exist, and it does not position table history as a default long-term backup strategy. Before relying on Time Travel during an investigation, the team has to confirm that the target version is still recoverable. For an AI workflow, these controls answer a narrow but valuable question: what did this table contain at the selected version or time? They do not identify every other table, file, feature transformation, or reference input used by the AI execution. What Lakebase point-in-time branching establishes Lakebase operates on a different object. A point-in-time branch creates an isolated branch from an earlier moment within the configured restore window. The production branch remains unchanged. That distinction is useful when an application, agent, or model depends on operational records stored in Postgres. A team can inspect the earlier database state, compare it with the present state, and run tests against the branch without rolling production backward. The branch still has a defined boundary. It represents the operational database state available within its restore window. It does not automatically identify external files, Delta tables, feature pipelines, prompts, model artifacts, or evaluation settings that may also have shaped the AI result. Time Travel and Lakebase therefore provide two kinds of historical evidence: Time Travel and RESTORE operate on supported table history. Lakebase point-in-time branching operates on an operational Postgres database. Neither control should be used as shorthand for the complete data state of an AI run. The missing question is whether the state belongs to the run Historical recovery tells a team what data existed. An AI investigation has to establish more. First, the team must know which combination of data was approved for the selected task. A recoverable table version may still be incomplete, inconsistent with a companion source, or unsuitable for the evaluation being performed. Second, that approved state has to be connected to the execution before the result is challenged. Otherwise, the team is reconstructing the likely input after the fact rather than resolving the state that the AI actually received. This is the boundary between Databricks historical controls and Syntitan. Time Travel and Lakebase recover the relevant storage state. Syntitan addresses task-specific data readiness and the evidence that connects an approved data state to an AI run. Historical recovery and run-level evidence answer different questions Databricks Time Travel recovers an earlier Delta table state and Lakebase point-in-time branching exposes an earlier operational database state. Syntitan connects an approved data state to a specific AI run through Release State, Run Binding, Diff, and Reproduce. Historical recovery Delta tableTime Travel restores access to an earlier table state. Lakebase databaseA point-in-time branch exposes an earlier operational state. AIrun Which approved state reached it? Run-level evidence Release State Run Binding Diff Reproduce Recovery shows what existed. Run evidence shows what produced the result. How Syntitan organizes task readiness and run evidence Syntitan begins with the requirements of the selected AI task. It diagnoses the data across six readiness axes: Usability, Integrity, Context, Consistency, Reproducibility, and Traceability. Get AI-Ready addresses identified gaps before the evidence path proceeds through four operating steps: Release State fixes the data state approved for the AI task. Run Binding connects the AI execution to that Release State. Diff compares released states to identify what changed. Reproduce restores the recorded state for a controlled rerun of the investigation. The sequence does not replace Delta table history, Lakebase branching, or MLflow tracking. It adds a task-level record: the approved data state tied to the execution. For a controlled data-side comparison, the team holds the selected task, model or agent, evaluation set, and relevant execution conditions constant while changing the data state. A material change in the result gives the team a reason to examine the data differences. Little or no change is also useful evidence because it shifts the investigation toward the model, prompt, code, tools, or runtime instead of treating data as the presumed bottleneck. These terms describe Syntitan’s current product model. They do not independently prove that every AI result will reproduce or improve. The model, code, seed, dependencies, runtime, prompts, and evaluation configuration still require evidence from the systems that own them. Compare the evidence, not the feature count Databricks historical controls and Syntitan evidence controls. This is not a product quality ranking. Decision questionDatabricks control or recordSyntitan control or recordWhat the evidence supports What did a Delta table contain earlier?Time Travel queryNot a replacement for table historyInspection of a historical table version or timestamp Should the current Delta table return to an earlier state?RESTORENot a replacement for table recoveryAn approved table-level recovery action and its recorded history What did the operational database contain earlier?Lakebase point-in-time branchNot a replacement for database branchingInspection or testing against an isolated historical Postgres state Was the data suitable for the selected AI task?Historical access contributes source evidence but does not establish task readiness by itselfSix-axis diagnosis and Get AI-ReadyA task-specific readiness decision with stated criteria Which approved data state reached the AI execution?Logged dataset references and run records can contribute evidenceRelease State and Run BindingA resolvable connection between the approved state and the run What changed between approved states?Comparison depends on the stored object and the team's implementationDiffThe data changes under review between released states Can the data-side investigation be rerun?Time Travel, RESTORE, or a Lakebase branch can recover historical inputsReproduceA bounded rerun using the recorded data state; full execution evidence remains necessary The table is not a product ranking. It is a control map. Databricks owns the platform and storage operations described in its documentation. Syntitan addresses the task-level qualification and state-to-run evidence that historical access does not establish on its own. Example: one AI result, two Databricks histories Consider a recommendation model that reads a Delta feature table and recent transactions from a Lakebase application database. After both sources are updated, the model produces a different ranking. Time Travel can expose the earlier feature-table version. A Lakebase point-in-time branch can expose the earlier transaction state. Those controls recover two important parts of the input path. The investigation still has to prove that those two historical objects formed the approved combination used by the earlier run. It also needs to determine whether the data met the selected task's readiness criteria. That is where Release State and Run Binding change the investigation from reconstruction to resolution: the team can identify the approved state behind the execution rather than infer it from timestamps alone. In that controlled comparison, Diff identifies the data changes under review, and Reproduce supports a bounded rerun from the recorded state. Neither step captures an execution condition that the team failed to preserve. Where MLflow fits MLflow on Databricks can record parameters, metrics, artifacts, code versions, and datasets for model-development and evaluation runs. That record complements Time Travel, Lakebase, and Syntitan because it covers parts of the execution that a historical data object cannot explain. Its evidentiary value depends on what the team logs. A dataset reference helps only when it resolves to the intended state. An experiment record cannot describe an external input, transformation, or runtime condition that was never captured. A defensible investigation connects three records: the historical table or database state; the execution record for the model or agent; and the task-level evidence showing which approved data state reached that execution. The AI reproducibility evidence boundary remains important. Returning to the same data supports a controlled comparison, but it does not freeze every condition that can affect the result. Use the missing evidence as the architecture decision Start with the question the team cannot answer. Use Time Travel when the investigation needs to inspect an earlier Delta table without changing the current state. Use RESTORE when the team has approved a table-level recovery and understands the downstream implications. Use a Lakebase point-in-time branch when the investigation needs an isolated historical operational database. Use MLflow or the current tracking system to preserve the execution details it can record. Use Syntitan when the unresolved decision concerns task readiness, the approved Release State, or the state-to-run path through Run Binding, Diff, and Reproduce. The stronger architecture does not force one control to perform another control's job. It preserves the boundary between historical storage, execution tracking, and task-level evidence, then connects the records needed for the investigation. For the broader platform boundary, see Databricks vs Syntitan: governing the estate, reproducing the run. The narrower decision here concerns what Time Travel and Lakebase can recover and what Syntitan adds when the team has to explain a specific AI result. Historical recovery is not run evidence Databricks Time Travel can recover a historical table state. Lakebase can expose a historical operational database. Both are valuable when an AI result changes. Syntitan addresses the next question: whether the data was ready for the selected task and which approved state reached the run. That distinction turns historical recovery into one part of an evidence path rather than the final explanation. Know which approved data state produced the result before the next AI decision. See how Syntitan connects AI-ready data to run evidence. --- title: "lakeFS vs Syntitan: Data Versioning Is Not AI Readiness" url: "https://cubig.ai/articles/lakefs-data-versioning-vs-syntitan-ai-readiness/" source: live --- When an AI system produces a different result after a data update, the first challenge is finding out what changed. A team may be able to trace the update and restore an earlier data state, but that does not necessarily explain why the result changed or whether the earlier outcome can be reproduced. lakeFS data versioning addresses the first part of that investigation. It gives data in object storage a repository history, allowing teams to isolate changes, create reproducible checkpoints, compare states, and reverse a bad commit. Explaining the AI result requires a wider record. The team may also need the code version, preprocessing logic, model artifact, parameters, evaluation data, metric definition, and execution conditions associated with the run. This is the practical distinction between lakeFS and Syntitan. lakeFS controls and recovers versioned data states. Syntitan addresses whether a particular state is ready for the intended AI task and how that state is connected to the evidence behind the run. A production workflow may need both layers. Start by identifying the missing evidence An investigation becomes easier when the team separates three questions that are often compressed into the word reproducibility. Can we recover the data? The team needs an immutable reference to the repository state and a history of the changes that produced it. Can we reconstruct the run? The team needs the model or agent configuration, code, preprocessing, parameters, evaluation setup, and other execution records. Can we defend the result? The team needs a bounded comparison that shows which change mattered for the intended task and what the evidence does not prove. These questions depend on one another, but they are not interchangeable. A model run cannot be traced to its input if the data state is unknown. A known data state cannot explain the result if the rest of the run is missing. What lakeFS records and controls lakeFS applies Git-like versioning semantics to data in object storage. Its official documentation defines a branch as an isolated version of the repository, a commit as an immutable checkpoint containing a complete snapshot, and a merge as an atomic update from one branch to another. The architecture matters. lakeFS stores versioning metadata separately while the underlying data remains in object storage. Creating a branch is a metadata operation rather than a full copy of every object. This gives teams an efficient way to test and compare data changes without duplicating the repository for each branch. The lakeFS architecture guide also explains that native clients retrieve version metadata from lakeFS while data operations continue against the underlying object store. That model provides concrete operational evidence: the commit that captured a data state; the branch on which a change was isolated; the differences reviewed before a merge; the history of how the production state was created; and the revert that recorded the reversal of a committed change. The distinction between reset and revert is useful during an incident. In the lakeFS CLI, reset removes uncommitted changes from a branch, while revert creates a new commit that reverses the effect of an earlier commit. Revert therefore preserves an auditable recovery path instead of erasing the history of the committed change. lakeFS can also enforce checks around those transitions. Its Actions and Hooks documentation describes pre-commit and pre-merge validation, including file-format and schema checks, as well as hooks that notify or trigger downstream systems. These capabilities make lakeFS more than a passive archive. They let a team place controls before a changed data state reaches an important branch. This is substantial production infrastructure. It should not be reduced to a simple rollback button. Why the data commit is only one part of an AI run lakeFS describes reproducible data states as useful for debugging and for rerunning machine learning work against earlier data. The limit is not that lakeFS lacks reproducibility. The limit is the object being reproduced. A lakeFS commit identifies the data state. A complete run record may also need the code and preprocessing versions, model or agent configuration, parameters, evaluation setup, runtime conditions, and output artifacts. This broader record is standard experiment-tracking territory. MLflow Tracking separately records code versions, parameters, metrics, datasets, model weights, and other run artifacts. The point is not that every team must use MLflow. Its documented run model provides a neutral illustration of why a data reference alone is not the complete execution record. The result is a layered evidence problem. Data version control can answer, “Which data state did we recover?” It cannot answer, without records from the surrounding workflow, “Did we run the same transformation and evaluation against that state?” Five records needed to investigate a changed AI result To explain a changed AI result, connect the data state, data changes, execution conditions, evaluation setup, and output. Evidence layerQuestion it answersExample recordFailure if missing Data stateWhich exact data reached the run?Repository, branch, commit or released-state referenceThe input cannot be recovered or compared reliably Data changeWhat changed from the prior state?Commit diff, merge history, change manifestThe investigation cannot isolate the candidate change ExecutionHow was the result produced?Code, preprocessing, model, parameters, seed, environmentThe team may restore the data and still run a different experiment EvaluationWhat task and test defined success?Evaluation set, split, metric definition, thresholdTwo scores may look comparable while measuring different conditions ResultWhat output is being reviewed?Metrics, predictions, logs, report artifactsThe team cannot connect the observed change to the recorded run Existing data, orchestration, experiment-tracking, and governance systems may already hold parts of the evidence envelope. The requirement is that the references remain connected closely enough for a reviewer to move from the result back to the exact data state and execution conditions. A rollback test should narrow the cause, not declare it Consider a hypothetical churn model whose recall drops after an April data update. The team uses lakeFS to identify the relevant commit and revert it, then reruns the workflow against the March data state. If recall returns to the earlier level under the same model, preprocessing, split, seed, and metric definition, the rollback provides evidence that the data change contributed to the deterioration. The team can then inspect the data diff and determine which fields, distributions, or class balance changed. If recall does not return, the rollback still creates useful evidence. It rules out a simple explanation in which the data commit was the only material change. The team should then inspect the remaining execution and evaluation records. A preprocessing change, for example, can create preprocessing drift even when the raw data and model version appear unchanged. In both cases, lakeFS improves the investigation because it makes the data intervention controlled and reversible. The conclusion still depends on whether the rerun held the other conditions constant and preserved the resulting evidence. Where Syntitan fits in the evidence path Syntitan approaches the problem from the AI-use side of the boundary. Its current product flow begins with a six-axis readiness diagnosis and Get AI-Ready. The evidence path that follows is organized around four operating steps: A restored data state still needs task-level proof Version control restores a known data state. Syntitan organizes the AI evidence path through Release State, Run Binding, Diff, and Reproduce. Known data state restoredBranch history and rollback recover the selected repository state. Specific AI taskSix-axis readiness check Release StateRun BindingDiffReproduce Result tied to evidenceThe released state, AI run, and bounded test can be examined together. Release State fixes the data state approved for the AI task. Run Binding connects the AI run to that Release State. Diff compares released states to identify what changed. Reproduce restores a prior state and reruns the investigation under controlled conditions. The product page also describes a Change Manifest and Portable Proof Kit as supporting artifacts for Reproduce. The Change Manifest records the relevant data changes, while the Portable Proof Kit packages the before-and-after data, manifest, evaluation configuration, and comparison harness used to review or rerun the bounded comparison. They are supporting artifacts, not separate stages in the operating sequence. Within this workflow, Proof Run has a narrower decision job. The team keeps the AI task, model or agent, evaluation set, and execution conditions fixed while comparing the current data state with an AI-ready Release State. If the result changes materially, data remains a plausible bottleneck and Diff helps isolate the relevant change. If the result does not change materially, the evidence points the investigation toward another layer, such as the model, prompt, code, tools, or runtime. This comparison does not guarantee improvement. It helps determine whether data is the layer that warrants further action. The workflow and artifact names are first-party product claims, not independent proof that every AI result will reproduce. This does not make Syntitan a replacement for lakeFS. lakeFS governs how data changes move through a versioned repository and can run validation around commit and merge events. Syntitan addresses whether a selected state is ready for a defined AI use and how that state is tied to the resulting proof. A team should evaluate overlap and integration requirements in its own architecture rather than assume the products connect automatically. Match the control to the missing evidence If the team cannot identify or reverse the data change that reached production, it needs stronger data version control. The relevant evidence is the branch, commit, comparison, merge, and revert history surrounding that change. If the data state is known but the result cannot be reconstructed, the missing evidence sits at the run level. The team must connect that state to the code, preprocessing, model, parameters, evaluation conditions, and output artifacts used in the execution. If those records exist but the team still cannot determine whether the data was suitable for the intended AI use, it needs task-specific readiness evidence and a controlled comparison. A defensible investigation connects all three layers: the data state, the execution record, and the task-level test. Rollback then becomes a controlled step in the investigation, not the final explanation. Connect every AI result to the data state and evidence behind it. --- title: "Audit Trail for AI Decisions: A Public Sector Proof" url: "https://cubig.ai/articles/public-sector-audit-trail-for-ai-decisions/" source: live --- An audit trail for AI decisions is a record that lets you show, for any specific AI-assisted decision, which data state and execution conditions produced it, and then restore that state so a reviewer can examine the decision as it was made. The stakes for an AI audit trail in the public sector are unusually high. GAO's AI accountability framework for federal agencies is built on four principles: governance, data, performance, and monitoring. In a government context the failure mode is not just a stalled project; it is a challenged decision that no one can reconstruct months later, when the data behind it has already moved on. Representative example The scenario and workflow below are illustrative until you reproduce them on your own model and data. The situation: an AI decision gets challenged Picture an agency that uses an AI workflow to help triage benefit applications, flag cases for review, or allocate limited resources. The workflow runs for months. Then one decision gets challenged, through an appeal, an oversight request, or a freedom-of-information filing. The team pulls up its records and produces a log entry. The entry shows that a decision was made, at a timestamp, with an output. What it cannot show is the data state behind that decision: which records were in scope, which version of the eligibility rules applied, which reference tables the model read, and under which permissions. The log proves the event happened. It cannot reconstruct why. Why a log is not an audit trail for AI decisions A timestamp and an output describe that something occurred. Public accountability asks a harder question: can you reproduce the conditions that produced this specific outcome? Those conditions live in the data state, and by default that state is never captured. Reference tables get refreshed. Eligibility rules get patched. Upstream records get corrected. Each of those changes is reasonable on its own, and together they quietly erase the exact context a reviewer needs. So an inquiry ends where it should not: "the system decided." That is the precise answer public oversight exists to prevent. The gap is not a missing log line; it is a missing ability to rebuild the world as it stood at the moment of the decision. Lawmakers are writing that expectation down: EU AI Act Article 12 requires high-risk AI systems to automatically record events over their lifetime so that results stay traceable. What a real AI audit trail has to prove Strip away the tooling and the requirement is concrete. For any past decision, an accountable agency needs to answer four questions with evidence, not recollection. The NIST AI Risk Management Framework makes the same point in its own language, calling for AI systems to be documented, traceable, and auditable rather than merely logged. Which records and fields were actually in scope when this decision ran? Which version of the rules, thresholds, and reference tables applied at that moment? Would a similar case decided a week earlier have come out differently, and if so, what changed? Can a reviewer see that state restored, not summarized, and inspect it directly? If those four cannot be answered from the system itself, every review becomes a forensic project: interviews, guesswork, and reconstructions that no one can fully trust. Where operating control changes the picture Binding each decision run to a Release State turns a log into a genuine AI audit trail. Instead of recording that a decision happened, the system records the state it happened against, and keeps that state addressable. Comparison of what a plain log and a bound Release State provide for decision review, data provenance, change tracking, restoration, and traceability Question under review What a plain log gives you What a bound Release State gives you Did the decision happen? ✓ Yes, with timestamp ✓ Yes, with timestamp Which data produced it? ✕ No ✓ Yes, via Run Binding What changed since a prior case? ✕ No ✓ Yes, via Diff Can the state be restored? ✕ No ✓ Yes, via Reproduce Is traceability built in? △ Partial ✓ Yes, as a system property Four capabilities carry the weight here. Run Binding connects each decision to the exact data state and conditions that produced it, so the link is captured at decision time rather than reassembled later. Traceability, one of the six readiness axes, makes "which data produced this decision" a standing property of the system. Diff shows whether a rule or reference change between two decisions explains why two similar cases were handled differently. Reproduce restores the decision's data state so a reviewer can examine it as it stood, not as it looks today. For production AI, the question is not only which model ran; it is which data state and execution conditions produced the result. Syntitan, CUBIG's AI-Ready Data Platform, scores enterprise data on six axes, Usability, Integrity, Context, Consistency, Reproducibility, and Traceability, rebuilds what blocks execution, and binds every AI or agent run to a data state you can diff and reproduce. Is your AI decision trail actually auditable? Run this quick test against your current setup. If you cannot answer yes to most of these, your logs are recording events, not evidence. Can you name the exact records and fields a specific past decision read? Can you tell which version of the rules and reference tables applied that day? Can you restore that state and show it to a reviewer, rather than describe it? Can you diff two decisions and point to what changed between them? Is this captured automatically at run time, not rebuilt by hand after a challenge? The outcome: examinable, not just explainable The accountable version lets an agency answer an oversight question with evidence: here is the decision, here is the exact state behind it, and here is that state restored for review. It does not claim the decision was correct, and it should not. It makes the decision examinable, which is the actual requirement in public administration. A citizen, an auditor, or a court can look at what the system saw and judge the decision on its merits. Where it fits An audit trail for AI decisions is one expression of a broader idea: AI is only accountable when its data state is operable and reproducible. This is the operating layer for AI-ready data, where a decision, a backtest, or an agent run all reduce to a Verifiable Data State you can diff and rebuild. The same machinery that answers "reproduce this incident" also answers "reproduce this decision," which is why reproducibility work in one domain pays off in the other. Try it on your data for free. Run a sample proof and see it on your own workflow. You can start by binding one decision workflow to a Release State and reconstructing a past decision from its bound state. For related reading, see operating control versus AI governance frameworks, how to reproduce an AI incident, and public sector citizen service workflows. --- title: "Visual Sensitive Data: Text Inside Images and Scanned PDFs" url: "https://cubig.ai/articles/visual-sensitive-data/" source: live --- Visual sensitive data is confidential information that lives inside an image or a scanned PDF rather than in machine-readable text: the part numbers annotated on an engineering drawing, the names and figures on a signed contract someone scanned back in, the handwriting on an intake form. It is the material a plain-text pipeline never sees, because there is no text to see, only pixels arranged to look like text. The scale of the blind spot is not small. Even privacy law is written for text: HIPAA recognizes exactly two de-identification paths, Expert Determination and Safe Harbor, and Safe Harbor works by removing 18 specified identifiers, a mechanism that only helps when the identifiers exist as characters a pipeline can find. Turning those pages into characters is itself lossy: the Donut researchers note that OCR errors propagate downstream and degrade every step that depends on the extracted text. Much of what an enterprise actually holds was never typed into a field. It was scanned, photographed, or drawn. When a team plans for sensitive data, they picture a name in a database column or a figure in a spreadsheet cell. Meanwhile the contract that actually governs the deal is a PDF someone signed and scanned, and the specification that carries the real intellectual property is a drawing with tolerances written in the margins. To a text pipeline those pages read as empty. They are anything but. Why plain-text handling misses visual sensitive data A text-based approach works by scanning for character spans and substituting them. Feed it a scanned contract and it finds nothing to substitute, because there are no characters, only an image of them. The page passes straight through. From the pipeline's point of view the document is clean and ready to send. In reality the most sensitive thing in the building has just left as a picture, and nothing flagged it. This is the quiet failure mode that makes it dangerous. It does not throw an error or raise a warning; the system did exactly what it was told, which was to look for text. A scanned PDF of a signed agreement, dense with names, prices, and terms, goes to a model verbatim because the scanner looked for characters and saw an image instead. The gap stays invisible right up until the moment it costs something. Plain masking has the same limitation from a different angle. Even where masking does fire on extracted text, it deletes or blacks out the value and leaves a hole, which is why people reach for it as a safety measure and then find the document no longer works. Neither approach was built for a page where the sensitive content and the structure that gives it meaning are fused into one raster surface. Where visual sensitive data hides in the enterprise It turns up in more places than most teams expect, and the pattern is consistent: wherever the authoritative version of a record is a physical artifact, the sensitive content ends up inside an image. Engineering and manufacturing run on drawings whose annotations carry the real value, including tolerances, materials, and supplier part numbers. Legal and finance archives fill up with scanned originals, because the version that counts is the one that was physically signed and stamped. Operations teams photograph forms, meter readings, and handwritten logs as the fastest way to capture what happened on the floor. Structure matters here as much as the values do. An annotation on a drawing means something because of where it points; a figure on a scanned invoice belongs to the line it sits on. Strip the layout and you inherit the same problem covered in document layout preservation: the arrangement was carrying meaning, and flattening it destroys the meaning along with any privacy you were trying to add. Real confidential business context lives in the relationship between a value and its position, not in the value alone. How a structure-preserving boundary handles it The handling principle is the one that runs through this whole approach, applied to a harder surface. Locate the sensitive text inside the image or scanned page, substitute it in place with a stand-in, and leave everything around it, the drawing, the form fields, the layout of the contract, exactly where it was. The model then receives a page that still reads as the document it is, with the confidential values swapped out. When the result comes back, the originals are reconstructed inside your environment, so the workflow closes on real data even though the model never touched it. The approach also lines up with GDPR Article 32, which requires technical and organizational measures appropriate to the risk and names pseudonymization explicitly, alongside the ability to restore access and to test the measures regularly. Put plainly: send the model the structure of the work, not the raw values, then rebuild the result locally. The rule has to hold for pixels as firmly as it holds for a database row, which is the demanding part. Detecting a name in a text field is one problem; finding the same name handwritten at an angle on a scanned form, substituting it, and keeping the form legible is a different order of difficulty. That cross-format consistency is the job of a Context-Preserving Data Layer for AI, in CUBIG's case LLM Capsule, which runs on the CUBIG Syntitan platform. A quick self-diagnostic Run this test on your own document pipeline before you trust it with a scanned archive. If you answer no to more than one of these, visual sensitive data is probably slipping through: If you drop a scanned, image-only PDF into your redaction or masking step, does anything actually get detected, or does the page pass as clean? Can your pipeline find sensitive text that is handwritten or set at an angle, not just clean typed characters? After handling, does the drawing, form, or contract still read as the same document, with layout intact? When the model returns a result, are the real values reconstructed inside your environment rather than left as stand-ins? Does the same boundary cover images and scanned pages, or only the plain-text parts of a mixed document? Text versus visual sensitive data at the boundary The contrast below shows why a boundary built only for text leaves the riskiest material exposed, and what a structure-preserving boundary does differently. Comparison of plain-text handling and a structure-preserving boundary across scanned documents, handwriting, layout retention, document usability, and local reconstruction At the boundary Plain-text handling Structure-preserving boundary Scanned, image-only PDF ✕ No ✓ Yes Handwriting and angled text ✕ No ✓ Yes Keeps layout and position △ Partial ✓ Yes Document still usable after △ Partial ✓ Yes Real values reconstructed locally ✕ No ✓ Yes The point is not that text handling is wrong; it works well on text. The point is that a scanned archive is mostly not text, so a boundary that stops at characters covers the easy half and misses the confidential half. Where visual sensitive data fits Visual sensitive data is one face of the broader multimodal AI data boundary, which holds the same principle across every format in a mixed document, from a table to an embedded chart to a scanned attachment. Together they extend sensitive AI workflow enablement to the documents your teams actually scan, photograph, and draw, rather than only the ones they type. The goal throughout is the same: let people run confidential work through an AI model in place, reconstruct the result locally, and close the workflow on real data without the raw values ever leaving the boundary. LLM Capsule extends the same boundary to what a model sees, not only what it reads. The visual structure reaches the model while the sensitive content in the image stays inside your environment. --- title: "Data Readiness Diagnosis: The Diagnose Step in DTS" url: "https://cubig.ai/articles/the-diagnose-step-in-depth/" source: live --- The Diagnose step is the first stage of DTS, where the engine measures why a dataset cannot be used to train or run AI, scoring what is restricted, what is scarce, and what is structurally broken before any data readiness diagnosis leads to a rebuild. When RAND interviewed engineers and data scientists about why AI efforts collapse, it put the failure rate above 80%, double what conventional IT projects suffer, and problems with data quality and availability ranked among the five root causes named most often. Most teams read that as a modeling problem. It is closer to a diagnosis problem: they never measured what was wrong with the data before they tried to train on it. The Diagnose step exists to catch that at the start, so the rebuild that follows is aimed at a real defect rather than a guess. What data readiness diagnosis actually measures DTS is CUBIG's AI-ready data transformation engine, and it works in three moves: Diagnose, Transform, Rebuild. The Diagnose step comes first because you cannot rebuild what you have not measured. Its job is narrow and concrete: read the dataset as it stands and report, field by field, why a model cannot learn from it in its current form. Three failure modes dominate. Data can be restricted, meaning regulation or contract forbids the raw values from reaching a training pipeline. Data can be scarce, meaning the events you most need to model, fraud, rare defects, unusual cohorts, appear too few times for a model to generalize. And data can be structurally broken, meaning schema drift, inconsistent encodings, or mismatched joins have quietly corrupted the relationships a model depends on. Diagnose produces a profile that separates these, because each one calls for a different rebuild. The output is not a pass or fail stamp. It is a readiness profile with a defect map attached: which columns are blocked, which patterns are underrepresented, and where the structure has broken down. That map is what the Transform step reads next. Why teams skip diagnosis, and what it costs The usual sequence in a stalled AI project is to blame the model, then the features, then the labels, and only much later the data itself. By then the team has spent weeks tuning around a defect that a proper data readiness diagnosis would have surfaced on day one. Redman's often-cited estimate, published through Harvard Business Review, put the cost of bad data in the US at roughly $3.1 trillion a year, and most of that waste is downstream of a data quality assessment nobody ran. Skipping diagnosis also hides which problem you have. A team that assumes its data is merely scarce will reach for augmentation and get nowhere, because the real blocker was a restricted field they were never allowed to use. A team that assumes a compliance wall will over-anonymize data that was structurally fine, destroying the signal in the process. Naming the defect correctly is half the fix. The three checks inside the Diagnose step Diagnose runs against the raw data through DTS's non-access architecture, so the data profiling reads statistical shape and structure without exporting the sensitive values themselves. Three checks run in parallel, and each one feeds a different part of the rebuild plan. Comparison of the questions answered by restriction, scarcity, and structure checks and what each check contributes to the data rebuild Check Question it answers What it feeds into the rebuild Restriction check Which fields are blocked by regulation or contract from reaching a model? Structure-Preserving Synthesis of the blocked fields Scarcity check Which important patterns are too rare for a model to generalize? Rare-pattern augmentation of the underrepresented cases Structure check Where has schema drift or bad encoding broken the data’s relationships? Repair of joins, types, and consistency before synthesis The restriction check maps every field against what is allowed to leave the boundary. The scarcity check measures class balance and the density of the rare events you actually care about, not overall row count, because a million rows can still hold only forty fraud cases. The structure check compares the data's current shape against what the schema promises and flags where they diverge. How diagnosis hands off to Transform and Rebuild Diagnosis is only useful if it changes what happens next. The readiness profile is a set of instructions for the rest of DTS. Restricted fields route to Structure-Preserving Synthesis, which rebuilds them into values a model can learn on while keeping the distributions and relationships intact. Scarce patterns route to rare-pattern augmentation, which increases the density of the underrepresented cases without inventing signal that was never there. Broken structure routes to repair first, because synthesizing on top of corrupted joins would only launder the corruption. This is why the Diagnose step is a measurement stage and not a cleaning stage. It does not fix anything on its own. It tells Transform and Rebuild exactly where to act, and it records the before-state so the improvement can be checked afterward. DTS runs on the Syntitan platform, which keeps that before-and-after state so a rebuilt dataset can be compared against what it replaced rather than trusted on faith. Run this quick diagnostic on your own dataset Before you commit to a rebuild, you can run a rough version of the Diagnose logic by hand. If you answer no to any of these, your data is not yet ready for the model you have in mind: Can you name every field that regulation or contract forbids from reaching a training pipeline? Do you know the raw count, not the percentage, of the rarest event your model must predict? Does the data's current schema match what your pipeline expects, with no silent type or encoding drift? Have you measured class balance on the specific outcome you care about, rather than assuming it from row volume? Can you tell whether your blocker is restriction, scarcity, or broken structure, rather than lumping them together as "bad data"? Most teams stall on the second and fifth questions. They know the data is a problem but cannot say which problem, and that is exactly the gap a data readiness diagnosis closes. Regulation now assumes this measurement exists: Article 10 of the EU AI Act holds the training, validation, and testing data behind high-risk systems to documented data-governance practices and to explicit quality criteria, meaning data that is relevant, sufficiently representative, and as close to complete and error-free as possible. Where the Diagnose step fits in CUBIG's operating layer DTS is CUBIG's AI-ready data transformation engine, and Diagnose is its opening move: measure the defect, then Transform and Rebuild against it. The rebuild uses Structure-Preserving Synthesis and rare-pattern augmentation, and the whole engine runs through a non-access architecture so restricted values are never exported to be profiled. DTS runs on the Syntitan platform, so a dataset that Diagnose flagged and Rebuild repaired can be diffed against its earlier state instead of being taken on trust. The practical takeaway is that diagnosis is not overhead you add to a data project; it is the step that decides whether the rest of the project is aimed at the right target. Run it first, and the transformation that follows has something specific to fix. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "Synthetic Data for Fraud Detection: Rebuilding Rare AML Patterns with DTS" url: "https://cubig.ai/articles/dts-for-fraud-aml/" source: live --- Synthetic data for fraud detection rebuilds the rare fraud and money-laundering patterns that real datasets barely contain, so a model can learn on enough labeled examples of the behavior it is supposed to catch. Fraud and anti-money-laundering (AML) teams sit on huge transaction histories, yet the events they most need to detect are a rounding error in the data. Confirmed fraud often runs well under one percent of transactions, and in AML transaction monitoring, confirmed laundering typologies are rarer still. That scarcity is not a minor inconvenience, and regulators have taken it up directly: the UK Financial Conduct Authority's Synthetic Data Expert Group devoted its report to synthetic data across fraud controls, model testing and validation, and data sharing in financial services. A fraud model trained on almost no positive examples is exactly the kind of project that stalls after a promising demo. Why fraud and AML data breaks machine learning The core problem is class imbalance layered on top of access restrictions. A model that sees fraud in 0.3% of rows can score 99.7% accuracy by labeling everything as clean, and it will still miss almost every real case. Rebalancing by throwing away legitimate transactions loses signal; naive oversampling of the few fraud rows teaches the model to memorize a handful of individuals rather than the shape of fraudulent behavior. The access problem compounds it. Transaction records carry account numbers, counterparties, amounts, and timing that fall under banking secrecy, privacy law, and internal risk controls. The data scientists who could improve the model frequently cannot touch the raw rows, and the raw rows cannot leave the regulated environment. So the scarce positive examples that do exist are also the hardest to work with. How DTS builds synthetic data for fraud detection DTS, CUBIG's AI-ready data transformation engine, treats a scarce fraud class as a structure to be rebuilt rather than a quota to be filled. It follows three steps: Diagnose, Transform, Rebuild. Diagnose measures where the data blocks learning, which for fraud usually means the joint structure of the rare class: how amounts, timing, sequence, and counterparty relationships co-occur in confirmed cases. Transform applies structure-preserving synthesis so the rebuilt records carry those correlations forward instead of sampling each field on its own. Rebuild produces a fuller labeled set the model can actually learn the minority pattern from. Two properties matter for regulated fraud work. First, this is rare-pattern augmentation, not blanket duplication: DTS concentrates on the parts of the distribution the model is starving for, and leaves the abundant clean traffic mostly intact. Second, it runs on a non-access architecture, so the engine learns the statistical structure of the sensitive transactions without analysts reading raw customer records. Differential privacy can bound how much any single real account influences the rebuilt output, which keeps a synthetic row from becoming a copy of one person's history. Structure-preserving synthesis versus plain resampling The distinction that decides whether a fraud model improves is whether the rebuilt minority class preserves relationships or just adds volume. Synthesizing the minority class has deep roots: SMOTE, the JAIR 2002 method, standardized the generation of synthetic minority-class examples to correct imbalance. Classic resampling and interpolation techniques add rows, but they tend to smear fraud examples toward the clean majority or invent points that violate how a real transaction behaves. Structure-preserving synthesis keeps the conditional structure intact, so a rebuilt laundering sequence still looks like a plausible chain of movements rather than a bag of independent numbers. Comparison of dropping clean rows, naive oversampling, and structure-preserving synthesis across rare-class structure preservation and restricted-data use Approach Preserves rare-class structure Works on restricted data Drop clean rows to rebalance ✕ No △ Yes, but loses signal Naive oversampling / interpolation △ Partial △ Yes, but risks memorizing individuals Structure-preserving synthesis (DTS) ✓ Yes ✓ Yes, non-access architecture How this changes AML transaction monitoring AML transaction monitoring has its own scarcity problem: known AML typologies such as layering, structuring, or mule networks appear in tiny numbers, and new typologies appear with essentially zero labeled history. Teams that rely only on rule thresholds drown in false positives, and teams that try to train supervised models rarely have enough confirmed cases per typology to generalize. Rebuilding rare typologies with structure-preserving synthesis gives the model more examples of the sequence-level behavior that distinguishes laundering from ordinary movement, which is where flat, per-transaction features fall short. Because DTS runs on the Syntitan platform, the rebuilt data does not sit off to the side as a one-time export. The augmentation step is part of a data state you can point a run at, so an AML transaction monitoring model retrained this quarter can be tied back to the exact rebuilt set it learned from. That matters when a regulator or an internal audit asks why an alert fired and what the model was trained on. Banks already answer to this standard on risk data: BCBS 239 requires risk data to be aggregated accurately and completely, largely by automated means, precisely to keep the probability of error down. A quick self-diagnostic for fraud and AML teams Run this short test before your next model retrain. If you answer "no" to two or more, scarcity is probably capping your detection rate: Do you have at least a few hundred confirmed positive examples per fraud or laundering typology you want to catch? When you rebalance, are you preserving the joint structure of the rare class rather than duplicating rows? Can your data scientists improve the model without reading raw customer transactions? Can you tie a deployed detection model back to the exact dataset it was trained on? Do rebuilt or augmented examples stay statistically representative until you reproduce results on your own held-out data? Where it fits in CUBIG's operating layer DTS handles one specific job: rebuilding restricted, scarce, or unusable data into data a model can learn on. For fraud and AML, that job is the difference between a model that has seen the behavior and one that has only seen its absence. DTS diagnoses where the rare class blocks learning, applies structure-preserving synthesis to rebuild it, and does so without exposing the underlying records. It runs on the Syntitan platform, so the rebuilt data belongs to a state you can reproduce rather than a loose file no one can account for later. The honest framing is this: synthetic augmentation makes a fraud model trainable, not omniscient. Performance you see on rebuilt data is representative until you reproduce it on your own held-out cases. That is the standard we hold DTS to, and the one your risk and audit functions will hold you to as well. For related depth, see structure-preserving synthesis, how to validate that synthetic data is good enough, and the pillar overview of what DTS is. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "Structure-Preserving Synthesis: Rebuild Restricted Data for AI" url: "https://cubig.ai/articles/structure-preserving-synthesis/" source: live --- Structure-preserving synthesis is an AI data transformation method that rebuilds restricted, scarce, or unusable enterprise data into a fully usable dataset while holding the statistical structure, column relationships, and rare patterns of the original intact, so a model can learn from the rebuilt data as if it had trained on the source. Most enterprise AI projects do not stall on the model. They stall on the data feeding it. Synthesizing data to repair a weak training set is not a new idea: SMOTE, the canonical technique from JAIR 2002, generates synthetic minority-class examples to correct class imbalance, but imbalance is only one of the ways enterprise data fails. When the training set is locked behind privacy rules, too small to cover the cases that matter, or too messy to feed a model, the project quietly dies. Structure-preserving synthesis exists to reopen that path. What structure-preserving synthesis actually preserves The word "synthetic" carries a lot of baggage. People hear it and picture random rows that look plausible but teach a model nothing. That is not what this method does. Structure-preserving synthesis is defined by what it keeps, not by what it invents. It preserves three layers of the source data. First, the marginal distributions: each column keeps its real shape, its range, its skew, its outliers. Second, the joint relationships: if income and default risk move together in the source, they move together in the rebuilt set. Third, and hardest, the rare patterns: the fraud case that shows up once in ten thousand rows, the defect that appears on one part in a batch, the cohort that only a handful of patients belong to. Generic synthetic generators tend to smooth these away because they are statistically inconvenient. A rebuild that erases the rare cases produces data a model can fit but not learn the hard problem from. DTS, CUBIG's AI-ready data transformation engine, runs this as a three-step loop: Diagnose, Transform, Rebuild. It first diagnoses where the source data blocks execution, then transforms the restricted or scarce parts, then rebuilds a dataset that carries the original's structure forward. The engine uses a non-access architecture, so the rebuild happens without a person or an external system reading the raw records, and differential privacy can be applied as a technique inside the transform step when the source demands a formal privacy budget. Why plain masking and generic synthetic data fall short Two older approaches try to solve the same problem and leave a gap. Masking, or anonymization, strips or scrambles identifying fields. It protects individuals, but it damages exactly the relationships a model needs, and re-identification research keeps showing that masked data is less anonymous than teams assume: Rocher et al. estimated that 99.98% of Americans would be correctly re-identified in any dataset using 15 demographic attributes, and that even heavily sampled anonymized datasets are unlikely to satisfy GDPR anonymization standards. Generic synthetic data goes the other way: it fits a broad generator to the data and samples new rows, which is fine for volume but tends to flatten the tails where the valuable, rare signal lives. Structure-preserving synthesis sits between those failures. It does not just hide values and it does not just manufacture volume. It diagnoses the structure worth keeping and rebuilds toward it. The table below lays out the difference. Where the method earns its keep: rare-pattern augmentation The clearest payoff shows up when the signal you care about is rare. Fraud, financial crime, equipment defects, and uncommon clinical presentations all share the same shape: the interesting cases are a tiny fraction of the data, and that fraction is often the part locked down hardest for privacy or compliance reasons. A model trained on data where fraud is one row in ten thousand will happily predict "not fraud" every time and score well on accuracy while being useless. Rare-pattern augmentation rebuilds more examples of the rare class that stay faithful to how those cases actually behave, rather than duplicating rows or injecting noise. The model then has enough real structure to separate the hard cases. This is why teams reach for DTS on fraud and anti-money-laundering data, on clinical cohorts too small to train on, and on manufacturing lines where a defect is precious precisely because it almost never happens. A fair caveat: rebuilt data is representative until you reproduce the result on your own source. The point of the method is to get a model that works on the real distribution, so validation against held-out real data is part of the process, not an afterthought. How to tell if you need structure-preserving synthesis Run this quick self-check against a stalled or blocked project. If two or more of these are true, generic tooling is probably the wrong fit. Your best training data sits behind a privacy or regulatory wall your model pipeline cannot cross. The cases you most need the model to catch are rare, and accuracy hides how badly it misses them. You tried masking and model quality dropped, because the relationships the model needed went with the masked fields. You tried an off-the-shelf synthetic generator and the rebuilt data lost the outliers and edge cases. You need a formal, defensible privacy statement for the data you train on, not just a promise. You cannot grant raw-data access to a transformation tool, so the rebuild has to happen without anyone reading the records. Where AI data transformation fits in an AI-ready platform Structure-preserving synthesis is one engine, not the whole system. For production AI the question is not only which model ran but which data state and execution conditions produced the result. DTS runs on the CUBIG Syntitan platform, so a dataset it rebuilds becomes a versioned input the rest of the platform can score, bind to a run, and reproduce later. That connection matters: a rebuilt dataset that no one can point back to is a liability, while a rebuilt dataset tied to a recorded data state is an asset an auditor can accept. In practice a team diagnoses a blocked source with DTS, rebuilds it into trainable data, and then treats that rebuilt set as a first-class data state on the platform rather than a one-off export — AI data transformation as a recorded pipeline step, not a side script. The rebuild engine solves the "we cannot use this data" problem; the platform around it solves the "prove what you used" problem. If you want the broader picture of what makes a dataset trainable in the first place, start with our overview of what AI-ready data means, then see how DTS works end to end and how rare-pattern synthesis applies to fraud and AML. A short worked example Consider a bank with two years of transaction records that cannot move out of a regulated environment, where confirmed fraud is roughly 0.1% of rows. The team wants a model that flags fraud in real time. Masking the account identifiers breaks the network patterns that make fraud detectable. Exporting the raw data is not allowed. A generic synthetic tool produces a clean-looking dataset that misses the coordinated, rare fraud rings entirely. With structure-preserving synthesis as the AI data transformation step, DTS diagnoses the transaction graph and the fraud tail, rebuilds a trainable dataset that keeps both the everyday spending structure and the rare coordinated patterns, and applies a differential-privacy budget so the output carries a defensible privacy statement. A budget alone is not the whole story: NIST SP 800-226 provides guidelines for evaluating differential privacy guarantees and documents the common implementation pitfalls it calls privacy hazards. The team trains on the rebuilt set, then validates on held-out real records inside the regulated environment. The number that matters is not how good the rebuilt data looks; it is whether the model reproduces its catch rate on the real distribution. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "CUBIG Operating Layer vs AI Governance Platforms: Policy vs Proof" url: "https://cubig.ai/articles/cubig-operating-layer-vs-ai-governance-platforms/" source: live --- AI governance platforms describe and police AI at the policy layer, while the CUBIG operating layer acts at the data-and-run layer, so a decision is not just documented as compliant under AI data governance but can actually be reproduced and proven from records. Buy an AI governance platform and you get policies, model inventories, risk registers, and documentation. That is real value, and for many obligations it is exactly what a regulator wants to see. The gap opens when someone asks you to prove that a specific decision was what your documentation says it was. A policy record describes intent; it does not rebuild the run. The CUBIG operating layer covers that second half. Governance itself is standardizing: ISO/IEC 42001 arrived as the first AI management system standard, with scope spanning AI lifecycle management, risk and impact assessment, and supplier oversight. Certifying how you govern is still not the same as rebuilding what a run did. What AI governance platforms do An AI governance platform organizes the oversight of AI across an organization. It maintains a model inventory, tracks risk and bias assessments, records approvals, maps controls to frameworks, and produces the documentation an auditor or board expects. This is necessary work, and doing it in one place beats scattering it across spreadsheets. The expectation reaches the intergovernmental level: the OECD AI Principles, which 47 governments adhere to, count transparency, accountability, robustness, security and safety among their tenets, and a governance platform is where that documented practice gets produced. What these platforms record is largely about the system and the process: which model exists, who approved it, what policy applies. That is the policy layer. It answers what you intend and what you have declared, and it does so well. Where the policy layer runs out Documentation describes a decision; it does not reconstruct one. When a regulator points at a single output from three months ago and asks you to justify it, a governance record can show the model was approved, and the policy was in force. It cannot show the exact data the model read that day, nor can it rebuild the run to demonstrate the result. The evidence stops at the paperwork, one layer above where the decision was actually made. This is not a flaw in governance platforms; it is their scope. They were built to govern, not to reproduce. Reproduction requires the data state from each run, which lives a layer down in the execution itself. Regulation is already writing requirements at that depth: Article 10 of the EU AI Act holds the training, validation, and testing data of high-risk AI systems to quality criteria under documented data-governance practices, data that is relevant, sufficiently representative, and as error-free and complete as possible. Comparison of an AI governance platform and the CUBIG operating layer across system layer, primary output, governance evidence, run reproducibility, and audit evidence Dimension AI governance platform CUBIG operating layer Layer Policy and oversight Data state and run Primary output Policies, inventory, documentation Release State, run binding, reproduce Answers “was this governed?” ✓ Yes △ Partial Answers “can you rebuild this run?” ✕ No ✓ Yes Audit evidence Records of intent Reproducible result Two layers, not two products to choose between The operating layer and the governance layer are complementary, and treating them as either-or leads to a program that appears compliant but cannot prove it. A governance platform tells the story of how AI is managed; the operating layer supplies the evidence that the story is true, run by run. Keep the AI data governance platform for policy, inventory, and oversight, and add the operating layer so that any documented decision can be resolved back to the data state that produced it and replayed on demand. Put plainly, governance says the decision was allowed; the operating layer shows the decision, rebuilt. Auditors increasingly want the second, because the first is a claim and the second is proof. Is your AI data governance backed by reproducible runs? If a regulator names one past decision, can you rebuild the exact data it used, or only show it was approved? Does your evidence stop at documentation, or reach the run itself? Can you diff two runs to explain why their outcomes differed? Could you replay a six-month-old decision and reproduce the same result? If your governance stops at records of intent, the operating layer is the half that turns those records into proof. Where the CUBIG operating layer fits The CUBIG operating layer runs on the Syntitan platform. It scores enterprise data on six readiness axes, Usability, Integrity, Context, Consistency, Reproducibility, and Traceability, and binds every AI or agent run to a data state you can Diff and Reproduce. Governance platforms sit above it at the policy layer; the operating layer gives their records something reproducible to stand on. For more, see what AI-ready data means and how operating control compares to AI governance frameworks. For a public-sector example of rebuilding and examining a past decision, see an audit trail for AI decisions in the public sector. Any figure you see is representative until you reproduce it on your own data. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "LLM Capsule vs AI DLP: Block the Data or Finish the Work" url: "https://cubig.ai/articles/llm-capsule-vs-ai-dlp/" source: live --- AI DLP and LLM Capsule both keep confidential data from leaking to an AI model, but they do opposite things with the workflow: AI DLP detects and blocks the data, stopping the task, while LLM Capsule substitutes the values and lets the task finish. As enterprises wire LLMs into real work, two confidential AI tools show up in the same conversation and get mistaken for each other. AI DLP extends data-loss prevention to AI traffic. LLM Capsule enables confidential workflows to run on AI. Both address sensitive data meeting a model, and that shared surface hides a basic difference in what happens to the work. The stakes were set early: after employees leaked sensitive internal data to ChatGPT in 2023, Samsung temporarily banned generative AI tools on company devices, and a ban is the bluntest form the blocking instinct takes. What AI DLP does AI DLP applies the data-loss-prevention model to AI traffic. It inspects prompts and payloads headed for a model, matches them against policy, and blocks, redacts, or alerts when sensitive data is detected. This is genuinely useful, and it is the right tool for its job: enforcing a rule that certain data must not leave, and creating an audit trail when someone tries. For egress control and policy enforcement, AI DLP does exactly what it should. The limit is in the verb. AI DLP is built to stop things. When it catches sensitive data in a workflow that actually needs that data, the workflow stops or the data gets stripped into something the model can no longer use. The policy held, and the work did not get done. What LLM Capsule does LLM Capsule starts from the opposite intent: the work should get done, and the values should stay home. Instead of blocking the sensitive data, it substitutes each value with a structure-preserving stand-in, lets the model execute on that version, and reconstructs the result inside your environment. The model never receives the raw values, and the task still completes, because what crossed the boundary was the structure of the work rather than the identities. This is sensitive AI workflow enablement. Where AI DLP asks whether data is allowed to leave, LLM Capsule arranges for the work to proceed without the data leaving at all, so there is nothing to block. Comparison of AI DLP and LLM Capsule across primary action, workflow effect, model input, returned output, and best fit Dimension AI DLP LLM Capsule Primary action Detect and block Substitute and reconstruct Effect on the workflow Stops or strips it Completes it What reaches the model Nothing, or redacted text Structure-preserving stand-ins What the team gets back A blocked request A finished result Best fit Egress policy enforcement Running confidential work on AI They are not competing for the same job Because both sit between sensitive data and a model, they look like alternatives. They are closer to different layers. AI DLP is a control that says no when a rule is violated, and every enterprise needs that control. LLM Capsule is an enabler that lets a confidential workflow run in the first place, which no amount of blocking can provide. A team can run both: DLP guarding the traffic that should never move, Capsule enabling the traffic that has to move as work rather than as data. The control side has no shortage of justification, since the AI Incident Database now counts more than 1,500 real-world AI failures. The failure mode to avoid is using a blocker where you needed an enabler. If your AI initiative keeps stalling because DLP correctly refuses to let the necessary data through, more blocking will not unstick it. That is the point where enablement is the missing piece. Which one does your situation call for? Do you need the sensitive data to never leave, or to be usable by the model without leaving? Is the goal to stop a risky action, or to complete a confidential task? When DLP blocks a request, does the work you needed simply not happen? Do you want a policy verdict back, or a finished business document back? If your answers point to completing work on data that cannot leave, a blocker alone will not get you there. Where LLM Capsule fits in Confidential AI LLM Capsule is a Context-Preserving Data Layer for AI that runs on the CUBIG Syntitan platform, enabling confidential AI workflows that a control layer can only permit or deny. It complements egress controls rather than replacing them. For sensitive environments, the reference point now exists: NSA and CISA, with Five Eyes partners, published joint guidance on deploying AI systems securely, spanning the confidentiality, integrity, and availability of AI systems and their data. For the fuller picture, see sensitive AI workflow enablement and why an AI gateway alone is not enough. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "DTS vs Data Anonymization vs Differential Privacy: How to Choose" url: "https://cubig.ai/articles/dts-vs-anonymization-vs-dp/" source: live --- DTS, data anonymization, and differential privacy all let you work with sensitive data, but they trade privacy for utility in different ways: anonymization strips detail, differential privacy adds noise under a formal budget, and DTS rebuilds the data so its structure survives for AI to learn from. Every method for using sensitive data makes the same bargain in a different way: give up some fidelity to gain some privacy. The three approaches teams weigh most often are classical data anonymization, differential privacy, and structure-preserving synthesis through DTS. They are not interchangeable, and choosing by reflex is how a project ends up with data that is private and useless, or useful and exposed. The exposed half has been measured: Rocher and colleagues estimated in Nature Communications that 15 demographic attributes are sufficient to correctly re-identify 99.98% of Americans in any dataset, and that even heavily sampled anonymized datasets are unlikely to meet the GDPR's anonymization standard. What data anonymization does Classical data anonymization removes or generalizes identifying detail: mask a name, bucket an age into a range, drop a rare field, coarsen a location. The appeal is that it is well understood and easy to explain to a reviewer. The cost is that it works by destroying information, and the fields it coarsens are often the ones a model needed. Push anonymization hard enough to resist re-identification, and the rare patterns quietly disappear, which is exactly where fraud, defects, and rare cohorts live. What differential privacy does Differential privacy takes a formal route. It adds calibrated noise so that any single individual's presence or absence cannot be inferred beyond a measured bound, tracked as a privacy budget. Its great strength is that the guarantee is mathematical rather than a matter of judgment. The trade-off is the noise itself: strong guarantees degrade accuracy, and on small or skewed datasets the noise can swamp the signal you were trying to keep. It is a powerful tool when you can afford the budget, and the data is large enough to absorb it. Implementation is its own hurdle: NIST SP 800-226 exists to help teams evaluate differential privacy guarantees, and it catalogs the common implementation mistakes it labels privacy hazards. What DTS does differently DTS, CUBIG's AI-ready data transformation engine, does not strip detail or blur it. It rebuilds the dataset using structure-preserving synthesis, generating new records that preserve the joint distribution and correlations of the original while not reproducing any real individual. Differential privacy can be applied during that synthesis, so the two are not opposed: DTS can reconstruct data and still carry a formal guarantee. The aim is a dataset that stays learnable, including in the rare classes anonymization tends to erase. The rebuild route has a published case in medicine: a paper in npj Digital Medicine argues for patient-centric synthetic data generation because it removes re-identification risk from analyzing real patient records. How to choose Match the method to what you are protecting and what you need back. For a low-risk data release where losing detail is acceptable, anonymization is simple and defensible. For aggregate statistics or queries over a large population where a provable bound matters most, differential privacy earns its keep. For training data that has to stay useful, especially when rare classes carry the signal, structure-preserving synthesis keeps the utility that the other two spend down. These lines blur in practice, which is why DTS treats differential privacy as an option it can apply rather than a rival. A quick self-diagnosis Do the fields you would anonymize carry the signal your model needs? Is your dataset large enough to absorb differential-privacy noise without losing the pattern? Are the rare cases, not the averages, what you are trying to model? Do you need the output to train a model, or only to publish a statistic? If the signal lives in the detail and in the rare cases, stripping or blurring it is the wrong bargain, and rebuilding it is the one that keeps the data usable. Where DTS fits DTS runs on the CUBIG Syntitan platform and is the option when the goal is data that stays learnable rather than data that is merely released. For the mechanics, see what DTS is, how structure-preserving synthesis holds the statistics, and how it differs from a generic synthetic data generator. Any figure you see for utility or privacy is representative until you reproduce it on your own data. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "I Came for Synthetic Data. Do I Actually Need LLM Capsule?" url: "https://cubig.ai/articles/i-came-for-synthetic-data-do-i-need-llm-capsule/" source: live --- If you came looking for synthetic data, the real question is whether your problem is a training dataset you cannot use or a live workflow you cannot run on confidential data, because the first calls for synthetic data and the second calls for LLM Capsule. Plenty of teams arrive asking for synthetic data when what they actually need sits one step away. The phrase has become shorthand for "let us use our sensitive data with AI," which covers two different privacy-preserving machine learning problems that need different tools. Getting the diagnosis right saves a quarter of building the wrong thing. The obvious shortcut, training directly on the raw records, has a documented failure mode: Carlini et al. showed that training-data extraction attacks recover verbatim sequences from LLMs, including names, phone numbers, and email addresses, and that larger models are more vulnerable. What synthetic data actually solves Synthetic data solves a training problem. You have a dataset that is too sensitive, too imbalanced, or too scarce to train or share, so you rebuild it into a version a model can learn on. In CUBIG's stack, this is the job of DTS, which diagnoses the dataset, then rebuilds it with structure-preserving synthesis so the statistics that matter survive. The output is a dataset. You take it into your training pipeline and build a model. If your blocker is at model-building time- that a training set cannot leave a boundary, or a rare class is too thin to learn- synthetic data is the right answer, and you can stop reading here. What LLM Capsule solves LLM Capsule solves a workflow problem. Your model already exists, often a frontier LLM or an agent, and the blocker is at runtime: the documents, records, and context you need the model to act on are confidential and cannot be sent to it. Here you are not building a training set; you are trying to complete a task today on live sensitive data. That exposure is routine in practice: Cyberhaven measured ChatGPT use across 1.6 million workers and found that 11% of what employees paste into ChatGPT is confidential data, from source code to client records. LLM Capsule handles this by sending the model the structure of the work rather than the raw values, then reconstructing the result inside your environment. The data stays home; the work still gets done. This is sensitive AI workflow enablement, and it is a different mechanism from rebuilding a dataset. Privacy-preserving machine learning: when you need both The two are not rivals, and some teams need each for a different stage. You might rebuild a training dataset with synthetic data to get a model that performs, then use LLM Capsule to run that model on live confidential documents in production. One tool prepares the data a model learns from; the other governs the data a model acts on. They meet in the same goal of privacy-preserving machine learning: getting value from data you could not previously put in front of AI. The cost of skipping the workflow stage is on record: Samsung temporarily banned generative AI tools on company devices in 2023 after employees leaked sensitive internal data to ChatGPT. A quick self-diagnosis Answer these about the problem in front of you right now: Are you trying to build or improve a model, or to complete a task with a model you already have? Is the sensitive data a training set, or the live input to a workflow? Do you want a dataset back, or a finished document back? Is the pain that a class is too rare, or that confidential context cannot reach the model? Lean toward the first answer, and you want synthetic data. Lean toward the second and you want LLM Capsule. If both describe you, you have a training stage and a workflow stage, and you will use each in turn. Where this fits in the CUBIG stack Both privacy-preserving machine learning tools run on the CUBIG Syntitan platform, which is why the decision is a fork rather than a wall. If your problem is training data, start with what DTS is and how structure-preserving synthesis rebuilds a usable dataset. If your problem is a live workflow, start with sensitive AI workflow enablement. Diagnosing which one you need is the whole point of asking the question before you build. If you arrived from the other direction, having come for LLM Capsule and wondering whether you also need synthetic data, the mirror-image guide I came for LLM Capsule, do I need DTS? walks the same fork from that starting point. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "Do I Need DTS If I Already Have LLM Capsule?" url: "https://cubig.ai/articles/i-came-for-llm-capsule-do-i-need-dts/" source: live --- If you came to CUBIG for LLM Capsule, you need DTS only when your data itself cannot teach or answer the model, not when your worry is sending confidential data to an LLM; LLM Capsule solves the second problem, DTS solves the first. Most teams arrive with a concrete blocker: legal will not let a model see raw customer records, so an AI project that looked ready in a demo stalls the moment it touches production data. That is the LLM data privacy question, and it is what LLM Capsule answers. The DTS question is quieter and shows up later, when the data you are allowed to use turns out to be too thin, too skewed, or too locked-down to train or evaluate anything reliable. The first blocker is well grounded: Rocher et al. estimated that 99.98% of Americans would be correctly re-identified in any dataset using 15 demographic attributes, and concluded that even heavily sampled anonymized datasets are unlikely to satisfy GDPR anonymization standards. In practice those two failure modes, no access and no usable data, account for most stalled AI projects. Knowing which one you actually have keeps you from buying the wrong tool. LLM data privacy: what LLM Capsule does, and where it stops LLM Capsule is a Context-Preserving Data Layer for AI, running on the CUBIG Syntitan platform. It sits between your workflow and the model so the model receives the structure of the work, not the raw sensitive values. The pattern is Substitute, Execute, Reconstruct: substitute confidential fields with placeholders before anything leaves your boundary, execute the model on that safe surface, then reconstruct the real answer locally from a protected mapping layer that never left your side. That closes a specific LLM data privacy gap. Your analysts can run a frontier model over contracts, tickets, or patient notes without the underlying values crossing the LLM data egress line. The work completes in place, and you keep an operational artifact you can hand to an auditor. The gap itself is not small: Cyberhaven's analysis of ChatGPT use across 1.6 million workers found that 11% of what employees paste into ChatGPT is confidential data, from source code to client records. Here is the boundary of what it does. LLM Capsule assumes the data you hold is already good enough for the model to do the job once it can see the structure. It does not improve the data. If your records are too sparse to represent the case you care about, or a whole category of examples is legally off-limits, protecting them on the way to the model changes nothing, because the model still has too little to work with. What DTS does that a data boundary cannot DTS is CUBIG's AI-ready data transformation engine. Where LLM Capsule governs how data crosses a boundary, DTS rebuilds the data itself into something a model can learn from. Its loop is Diagnose, Transform, Rebuild: it diagnoses what makes a dataset unusable, transforms it through structure-preserving synthesis so the statistical shape survives, and rebuilds the parts that were missing, including rare-pattern augmentation for the cases that barely appear in your real records. It uses a non-access architecture, and differential privacy is available as a technique when the source is restricted. That technique deserves care in its own right: NIST SP 800-226 provides guidelines for evaluating differential privacy guarantees and documents the common implementation pitfalls it calls privacy hazards. The tell is simple. If your problem is that the data exists and is fine but must stay confidential, DTS is not your answer. If your problem is that the data you are allowed to touch cannot support training or evaluation, no boundary product will fix that, and DTS is exactly the tool for it. DTS also runs on the Syntitan platform, so adopting one does not strand you from the other. LLM Capsule vs DTS: which blocker are you solving? The two products answer different questions. One is about access, the other about substance. This table lays them side by side so you can place your own situation. Side-by-side comparison of LLM Capsule and DTS across the blocker each solves, the situation each fits, what each does to the data, how each works, where the data sits, and the shared platform   LLM Capsule DTS The blocker it solves Access — you cannot send the data to the model. Substance — the data itself cannot teach or answer the model. Your situation The data already exists and is accurate, but must stay confidential. The data you are allowed to use is too thin, too skewed, or too locked-down. What it does to the data Governs how data crosses your boundary; it does not improve the data. Rebuilds the data itself into something a model can learn from. How it works Substitute, Execute, Reconstruct. Diagnose, Transform, Rebuild. Where the data sits Raw values never cross the egress line; the protected mapping layer stays on your side. Non-access architecture, with differential privacy available when the source is restricted. Platform Runs on the Syntitan platform. Runs on the Syntitan platform. When you came for LLM Capsule but actually need both Plenty of teams need both, and the order matters. A bank that wants a model to read confidential deal memos starts with LLM Capsule, because the immediate blocker is egress. Months later the same bank tries to train a fraud model and finds it has only a few dozen real fraud cases, far too few to learn a reliable pattern. That is a DTS problem, and no amount of boundary protection creates the missing examples. The reverse also happens. A team adopts DTS to augment a rare clinical cohort, gets a workable dataset, then wants to run a frontier LLM over the real patient notes for a separate task. Now the access question returns, and LLM Capsule handles it. Treat the two as answers to two independent questions rather than as competitors, and the sequencing takes care of itself. There is a third thread worth naming, because production teams hit it fast. Once a model runs on rebuilt or boundary-protected data, the question is not only which model ran but which data state and execution conditions produced the result. Syntitan, CUBIG's AI-Ready Data Platform, scores enterprise data on six axes, rebuilds what blocks execution, and binds every AI or agent run to a data state you can diff and reproduce, so an answer you got last quarter is still explainable this quarter. A quick self-test before you choose Run through these. If most of your yes answers land in the first cluster, you came to the right place for LLM Capsule. If they land in the second, you are looking at a DTS problem instead. The data I need already exists and is accurate, but compliance will not let it reach the model. (LLM Capsule) My analysts want to use a frontier LLM on confidential records without those records leaving our boundary. (LLM Capsule) The case I most need to detect barely appears in my real data. (DTS) A whole category of records is legally off-limits for training, so my model never sees it. (DTS) My evaluation set is too small or too skewed to trust the numbers it produces. (DTS) I need both: run a model on confidential data now, and rebuild a scarce dataset for a separate task later. (Both, sequenced) Where it fits in CUBIG's operating layer LLM Capsule and DTS are two entry points into the same operating layer for AI-ready data, and they meet on the Syntitan platform. LLM Capsule keeps confidential work inside your boundary; DTS rebuilds data that cannot yet teach a model; Syntitan holds the reproducible AI-ready state underneath both, so the run you approved is the run you can reproduce. You do not have to decide the whole architecture on day one. You do have to name the blocker in front of you correctly, because the right first product follows directly from that. If you are still unsure which side of the line you are on, the fastest way to find out is to look at the data you would hand the model tomorrow and ask whether the problem is that you cannot send it or that it cannot answer. For the deeper split between rebuilding data and protecting it, our guide on whether synthetic data buyers actually need LLM Capsule takes the same decision from the other direction, and what DTS is covers the rebuild engine in full. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "Rare Defect Synthesis for Manufacturing: Rebuilding Rare Defect Data" url: "https://cubig.ai/articles/dts-for-manufacturing/" source: live --- Rare defect synthesis is the practice of rebuilding scarce, restricted, or unusable manufacturing defect data into a balanced, learnable dataset that preserves the statistical structure of the real production line, so a machine vision inspection model can actually learn to catch the failures that almost never happen. On a mature production line, a critical defect might show up in fewer than one part in ten thousand. That scarcity is exactly what makes the defect dangerous and exactly what starves the model meant to detect it. The class-imbalance problem is old enough to have a canonical answer: SMOTE, published in JAIR in 2002, corrects imbalance by generating synthetic examples of the minority class. In manufacturing quality, the data is rarely missing. It is imbalanced, locked to the shop floor, or too sparse in the classes that matter. Why rare defects break machine vision inspection models A machine vision inspection model learns from what it sees. When 9,999 of every 10,000 images are good parts, the model can score 99.99% accuracy by predicting "pass" every single time and still miss the one failure that triggers a recall. The metric looks excellent. The model is useless. Class imbalance is the surface problem. Underneath it sit three harder ones. Real defect images are often tied to a specific customer's product and cannot leave the plant. Historical defect records are thin because good engineering made the defects rare in the first place. And the defects that do appear cluster around specific process conditions, so a model trained on last quarter's failures does not generalize to a new material lot or a retooled station. That decay is measurable: a Scientific Reports team aged four standard model types on 32 industry datasets and watched 91% of the combinations degrade over time. Simple fixes fall short. Oversampling the minority class just shows the model the same handful of images repeatedly, which teaches it to memorize rather than generalize. Generic augmentation (rotate, flip, adjust brightness) changes pixels without adding a genuinely new defect. What a quality model needs is more true variation in the rare class, not more copies of what already exists. What rare defect synthesis actually rebuilds Rare defect synthesis produces new, plausible examples of the failure modes that are underrepresented, while keeping the relationships that define a real defect intact: the way a crack propagates along a grain boundary, the co-occurrence of a surface scratch with a specific coating thickness, the sensor signatures that accompany a solder void. The point is not to draw more pictures. The point is to give the model examples that behave, statistically, like defects it will meet in production. This is where DTS, CUBIG's AI-ready data transformation engine, does its work. DTS treats the rare class as a distribution to be rebuilt, not a folder of images to be copied. It uses Structure-Preserving Synthesis to hold the correlations that make a defect a defect, then applies rare-pattern augmentation to expand the thin class into a range a model can actually learn from. Differential privacy can be applied during synthesis so the rebuilt set does not leak any single real part or customer record, which is what lets the data move beyond the plant at all. Diagnose, Transform, Rebuild on defect data DTS runs the same three steps on manufacturing data that it runs anywhere, and the sequence matters because you cannot rebuild well what you have not measured. Diagnose. Before touching anything, DTS profiles the dataset: which defect classes are present, how badly they are imbalanced, which features carry signal, and where the gaps are. On a real line this is often the first time a quality team sees, in numbers, just how thin their worst classes are. Transform. DTS then rebuilds the scarce classes using Structure-Preserving Synthesis, so the synthetic defects respect the joint distribution of the originals rather than inventing artifacts a physicist would laugh at. This is a different thing from a generic synthetic data generator that optimizes for realistic-looking pixels; here the target is statistical fidelity to the failure mode. Rebuild. Finally DTS assembles a balanced, model-ready dataset that combines real good parts, real rare defects, and rebuilt rare defects, ready to hand to your training pipeline. Because DTS uses a non-access architecture, it operates on the structure of the data without pulling raw parts or images out of the controlled environment. The caution fits the territory: a production line is operational technology, which NIST covers with a dedicated security guide, SP 800-82 Rev. 3, built around the performance, reliability, and safety constraints the domain runs under. Rebuilt synthetic data versus the usual workarounds Quality teams have tried several ways to feed a starving model. Here is how rare defect synthesis compares to the common alternatives on the things that decide whether a defect model ships. Comparison of approaches for adding rare-class variation, preserving defect structure, and safely moving data outside the plant Approach Adds real variation to rare class Preserves defect structure Can leave the plant safely Do nothing (raw imbalanced data) ✕ No ✓ Yes ✕ No Oversample / duplicate minority class ✕ No △ Partial ✕ No Generic image augmentation △ Partial △ Partial ✕ No Generic synthetic data generator ✓ Yes ✕ No △ Partial Rare defect synthesis (DTS) ✓ Yes ✓ Yes ✓ Yes The column that usually surprises people is the third one. A model that never leaves a single plant is a science project. Because DTS can apply differential privacy during synthesis, the rebuilt defect data can be shared across sites, vendors, or a central AI team without exposing the underlying parts, which is often what turns a pilot into a program. Is your defect data ready to train on? Run this quick check against your own quality dataset before you spend another sprint tuning a model that cannot see enough failures. Does your worst defect class have fewer than a few hundred genuine examples? Would your model score above 99% accuracy simply by predicting "pass" on everything? Are your defect images or records locked to one plant because of customer or regulatory constraints? Do new material lots or retooled stations produce failures your current model has never seen? Are you oversampling or duplicating the minority class to force class balance? Two or more yes answers means the bottleneck is your data, not your architecture, and no amount of model tuning closes that gap. Where rare defect synthesis fits in the CUBIG stack DTS is the engine that rebuilds restricted, scarce, or unusable data into something a model can learn on. Rare defect synthesis is DTS pointed at the hardest version of that problem: a class so thin the model has almost nothing to study. On manufacturing data, DTS diagnoses the imbalance, rebuilds the rare classes with Structure-Preserving Synthesis and rare-pattern augmentation, and hands back a balanced dataset, all under a non-access architecture so raw parts stay where they belong. DTS runs on the Syntitan platform, which is where the rebuilt data connects to the rest of your AI-ready workflow. If you want the deeper mechanics, start with what DTS is, then read how Structure-Preserving Synthesis keeps the statistics honest, how the Diagnose step measures your gaps first, and how DTS validation confirms the rebuilt set behaves like the real one before you train on it. A note on expectations: a model trained on rebuilt defect data performs representatively until you reproduce the result on your own line. That is the point of running a proof on your data rather than trusting a benchmark from someone else's factory. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "Private LLM Workflows for Defense on Classified Data" url: "https://cubig.ai/articles/defense-air-gapped-llm/" source: live --- Air-gapped LLM workflows let a defense team run a private LLM on classified material inside an accredited enclave by sending the model the structure of the work, not the raw values, running the analysis in place, and reconstructing the finished product against the real data behind the boundary. Defense pushes the confidential-data problem to its hardest edge. There is no "send it out carefully" option, because there is no outside. And the bar is written down now: NSA and CISA, joined by their Five Eyes partners, have published joint guidance on deploying AI systems securely, organized around the confidentiality, integrity, and availability of AI systems and their data in sensitive environments. The tooling is catching up to the constraint: research toolkits like OnPrem.LLM exist specifically to apply large language models to sensitive, non-public data in offline or restricted environments. A classified enclave is exactly the kind of environment that guidance describes, and the first obstacle there is blunt: the data cannot reach the model. The analyst who synthesizes by hand Picture a defense analyst with a stack of intelligence reports to synthesize before a morning briefing. The reports are dense, cross-referencing, and full of the kind of detail a model could pull together in minutes. But the work happens on an air-gapped network with no path to the outside, and the strongest models the analyst has heard about live on an internet the enclave cannot reach. The data cannot come to the model, and the model cannot come to the data. So the analyst synthesizes by hand, again, while the capability everyone else uses stays on the far side of the gap. Multiply that by every shift and every desk, and the cost is not one late briefing; it is a standing tax on how fast the mission can reason. Why the obvious answers do not hold The first instinct is to bring a smaller private LLM inside the enclave and run everything locally. That genuinely helps, and many programs do it, but it trades away the capability gap: the in-enclave model often sits well behind the frontier, and on hard synthesis tasks that distance shows in the output. The second instinct is to sanitize the reports down to a level cleared for an outside system. In practice that means stripping so much that the model is synthesizing a hollowed-out version of the material, and the analyst still has to redo the real work by hand. Neither path lets a strong model reason over the actual intelligence. Both approaches share a hidden assumption, that the classified content and its structure are one indivisible thing that either stays locked in or gets gutted. That assumption is where they break, because the two come apart. Substitute, Execute, Reconstruct inside the boundary The enablement approach lets the structure of the work be processed while the classified values stay put. Three steps carry it. Substitute. Sensitive entities, designators, locations, and identifiers are swapped for consistent stand-ins, while the shape of each report is preserved: the references between documents, the sequence of events, the analytic relationships. The same real value maps to the same stand-in every time, so cross-document reasoning still holds. Execute. The model synthesizes against that structure. Where it runs depends on the program's accreditation; in a fully air-gapped posture the whole flow executes on approved infrastructure inside the boundary, with nothing crossing the gap at any step. Reconstruct. When the model returns a draft, that draft is rebuilt against the real values behind the boundary using a protected mapping layer, so the analyst reads a finished product with the classified detail back in place. This closes the workflow: the analyst gets an operational artifact, not a redacted sketch they still have to finish. For production defense AI the question is not only which model ran, but which data state and execution conditions produced the result. The point of the design is that classified values never become the thing that has to travel; only the work's structure moves, and only inside walls that are already accredited to hold it. A private LLM compared to the alternatives The three paths a program can take diverge on one axis: does a strong model ever reason over the real material, and does the classified data stay put while it does. Approach Strong model on real material Classified data leaves the enclave Small in-enclave model Partial No Sanitize and send out No Partial Substitute, Execute, Reconstruct Yes No The first row keeps the data safe but caps the reasoning. The second row reaches for a better model but hands out a gutted version of the work. Only the third row lets a frontier-grade model reason over the true structure while the sensitive values stay inside the boundary. A quick self-check for your enclave Run these five questions against any AI workflow you are weighing for classified or restricted work. If you answer "no" to more than one, the workflow is either capping your model or leaking your data. Does the classified value ever have to leave the accredited boundary for the model to work? Can the model you actually want to use reason over the real analytic structure, not a stripped-down copy? Does the same real entity map to a consistent stand-in across every document in the set? When the draft comes back, is it reconstructed into a usable product inside the enclave, or does an analyst finish it by hand? Can the security authority point to a concrete reason the data-handling rules are satisfied? Where it fits This is defense's instance of sensitive AI workflow enablement, under the strictest version of the constraint. The same separation of structure from values that lets a telecom NOC reason over a live network without exposing it lets an analyst synthesize intelligence without exposing the source, and the same pattern shows up in confidential financial documents. The substitution and reconstruction are performed by a Context-Preserving Data Layer for AI, in CUBIG's case LLM Capsule, embedded in the environment the program already runs rather than bolted on as a new external service. It runs on the CUBIG Syntitan platform. Because the classified values never leave the boundary, the workflow fits the handling rules the environment is accredited under, and the security authority has a concrete basis to approve. The posture parallels how NIST treats operational technology in SP 800-82 Rev. 3: certain systems carry performance, reliability, and safety constraints so particular that the security guidance has to be written around them. That approval matters more here than anywhere, and it is still a by-product. The point was to put a frontier-grade tool in the analyst's hands on the real material, and to see it hold up on your own workflow before you trust it. LLM Capsule lets a private LLM run the same pattern inside an air-gapped environment. The model executes in place on substituted data, so nothing crosses the boundary and the mission data never leaves. --- title: "AI on Telecom NOC Data: Topology and Subscriber Records" url: "https://cubig.ai/articles/telecom-noc-workflows/" source: live --- Running AI on telecom NOC data means letting AI root cause analysis correlate live network topology and subscriber records to speed up incident triage, without either of those confidential inputs ever leaving the operator's environment. A carrier network sits closer to operational technology than to ordinary IT; the terrain NIST maps in SP 800-82 Rev. 3, its guide to securing OT, is written for systems where performance, reliability, and safety constraints come first. In a telecom network operations center, the blocker is rarely model quality. It is that the two most useful inputs, the topology and the subscriber records, are exactly the two the operator cannot send outside. So the work that a model could do at 2 a.m. gets done by a tired engineer instead, while the outage keeps running. Why the NOC is a hard case Picture a regional outage spreading across a dozen nodes. The on-call engineer has alarms firing in some order, a topology map showing how those nodes connect, a handful of recent change tickets, and a stream of affected-subscriber records. A model could correlate the alarms against the recent changes and propose a likely root cause faster than a person working by hand. Two kinds of confidential data sit inside that one incident. The subscriber records are personal data the operator is legally bound to protect. The network topology is operator-sensitive infrastructure that both competitors and attackers would value. Neither is cleared to leave for an external model, so the correlation stalls. Why anonymizing the incident stalls too The obvious workaround is to scrub the incident and send the anonymized version. In a NOC that breaks fast, because the relationships are the diagnosis. Which node connects to which, the order in which alarms fired, which change ticket preceded the failure: strip those specifics and the model can no longer trace cause and effect through the network. The topology is not context around the problem. It is the problem. The other common answer is to keep everything inside and forbid the external model, which is where most operators sit today. That leaves the engineer to correlate by hand while the clock runs, and it also caps what a NOC can do with the strong general models that live outside the operator. Both routes rest on the same assumption, that the sensitive identifiers and the structure of the incident cannot be separated. In a network, where structure carries almost all of the meaning, that assumption is the thing worth challenging, because if you can hand a model the shape of the incident while the real values stay home, the tradeoff between speed and confidentiality mostly disappears. How AI on telecom NOC data actually runs The enablement pattern is Substitute, Execute, Reconstruct. You send the model the structure of the incident, not the raw records and the literal topology. The connectivity graph keeps its shape, so the model can still reason over how a fault propagates, while node identifiers, subscriber details, and the specifics that reveal the real network are swapped for consistent stand-ins. The alarm and change sequence stays in order. The model correlates against a faithful structure and proposes a root cause: it might flag the change ticket that preceded the first alarm, or the node whose failure best explains the pattern of downstream symptoms. That proposal is reconstructed inside the operator's environment, where the real node and subscriber identities return, and the engineer reads an analysis that points at actual equipment. The triage ran. The subscriber data and the topology never left the operator. A Context-Preserving Data Layer for AI, in CUBIG's case LLM Capsule, holds the real values inside and rebuilds the analysis on the way back. It runs on the CUBIG Syntitan platform. Preserving the graph rather than flattening it is what lets the model reason about propagation at all; a protected mapping layer keeps the link between stand-ins and real identities inside the operator's boundary, so reconstruction happens locally rather than at the model. What crosses the boundary, and what does not Incident element Sent to the model? How it is handled Connectivity graph shape Yes Structure preserved so propagation can be reasoned about Real node identifiers No Consistent stand-ins; real values held in a protected mapping layer inside the operator Subscriber records No Substituted before egress, restored locally on reconstruction Alarm and change sequence Yes Order kept intact so cause and effect stay traceable Root-cause analysis Returned Reconstructed inside the operator with real identities A quick test for your own NOC Run this check against an incident-response workflow you would like a model to help with: Would the incident contain subscriber records, and are you bound to keep them inside? Does the diagnosis depend on the topology, the connection order, or the change sequence? If you anonymized the incident, would the model lose the relationships it needs to reason? Can your reconstruction step return real node and subscriber identities inside your environment? Does the workflow close, meaning the engineer gets an analysis pointing at actual equipment? If most of these are yes, the barrier is data movement, not model capability, and enablement is the pattern that fits. Where it fits This is the telecom view of sensitive AI workflow enablement. The same logic carries into other regulated operations. In the public sector the protected file is a citizen case file rather than a network incident, and the method holds: send structure, restore values locally. The concern that a retrieved record or a tool result quietly carries raw identifiers out is the same risk described in agent tool-output leakage. All of it runs on the CUBIG Syntitan platform. Because neither subscriber data nor topology crosses the boundary, the workflow maps onto the telecom confidentiality and personal-data duties an operator already carries. The direction also matches what security agencies now expect: NSA and CISA, with their Five Eyes partners, published joint guidance on deploying AI systems securely, centered on the confidentiality, integrity, and availability of AI systems and their data in sensitive environments. The threat load is real and rising: ENISA counted 188 telecom security incidents across the EU in 2024, more than a fifth above the year before. Security and compliance get a basis to approve, but that is the by-product. The point was to shorten the outage. LLM Capsule is how CUBIG runs AI root cause analysis on NOC workflows without subscriber or topology data leaving. The network structure reaches the model as stand-ins while the real identifiers stay inside your environment. --- title: "Clinical AI Workflow on PHI: Run Models Without Exposing the Chart" url: "https://cubig.ai/articles/healthcare-phi-workflows/" source: live --- A clinical AI workflow on PHI is HIPAA-compliant AI: it runs a language model on real charts and notes by sending the structure of the case, keeping protected health information inside the hospital, and restoring the identifying values to the finished document locally. The pressure to use models on clinical text is real, and so is the rulebook standing between the two. HHS guidance recognizes exactly two de-identification paths under HIPAA: Expert Determination, or Safe Harbor, which strips the 18 listed identifiers and demands no actual knowledge that what remains could point back to a patient. The clinical stakes are recognized at the top: the World Health Organization warns that large multi-modal models in health carry cybersecurity risks that could endanger patient information. In healthcare the data problem has a sharper edge: the most useful text a model could touch is the text a hospital is least allowed to move. That single constraint stalls more clinical AI than any model limitation does. Picture the discharge summary that needs writing. A patient spent nine days on the ward. There are progress notes from four clinicians, a medication list that changed twice, two imaging reports, and a free-text nursing handover that mentions the patient's name, date of birth, and the room of the patient next door. A model could draft that summary in seconds, and draft it well, because the note is dense, repetitive, and exactly the kind of text language models compress cleanly. Every line of it is protected health information, and the chart is not allowed to leave the hospital. The most useful thing the model could do is the one thing the data won't permit. Why stripping the chart and banning the model both fail Hospitals reach for one of two responses, and a clinical AI workflow on PHI breaks on both. De-identify the chart first, and you learn quickly that clinical meaning lives in the details you just removed. The temporal sequence of events, the specific dosages, the way one clinician's note references another's finding from two days earlier: scrub aggressively enough to feel safe and the model is summarizing a case that no longer hangs together. The draft comes back plausible and subtly wrong, which in a clinical setting is worse than no draft at all. Lock everything down and refuse the external model, and the clinician keeps writing summaries by hand after a twelve-hour shift. The promised time savings never arrive, and some staff, under pressure, quietly paste chart text into a consumer chatbot to get through the backlog. That last outcome is the one every privacy officer fears most, and a blanket ban tends to produce it rather than prevent it. The trap is treating the patient identifiers and the clinical reasoning as one inseparable thing. They are separable, and the whole workflow turns on that fact. HIPAA compliant AI applied to the PHI workflow The enablement pattern sends the model the structure of the case rather than the chart itself. The nine-day timeline stays a timeline. The medication that was started, escalated, then stopped stays that exact arc. The four clinicians' notes stay distinct and stay in order. What gets substituted is the identifying layer: the patient's name and date of birth, the neighbor mentioned in the handover, the record numbers, all replaced before anything reaches the model, while the medical structure underneath is preserved faithfully. HIPAA compliant AI drafts its discharge summary against that preserved structure. The draft is then reconstructed inside the hospital's systems, where the real patient identity returns to the document and the summary becomes something a clinician can sign rather than rebuild from scratch. The protected data never left. The drafting ran inside the environment the hospital already trusts, on-prem or in its existing cloud, not in some new place the chart had to travel to. A Context-Preserving Data Layer for AI, in CUBIG's case LLM Capsule, does the holding and the rebuilding. It runs on the CUBIG Syntitan platform. Because the substitution preserves structure instead of blanking fields, the model sees a coherent case rather than a chart full of black bars. That distinction is the difference between a draft the clinician edits and signs and a draft the clinician has to write over. Three moves: Substitute, Execute, Reconstruct The workflow closes in three moves, and it helps to name them. Substitute. Identifying values are swapped for stand-ins before anything leaves the boundary, with the originals held in a protected mapping layer inside the hospital. Execute. The model drafts against the substituted case, which still carries the timeline, dosages, and cross-references it needs to reason well. Reconstruct. The stand-ins are swapped back for the real values locally, producing an operational artifact the clinician can sign. The table below shows how this compares to the two habits it replaces. Dimension De-identify the chart Ban the model Context-preserving substitution Clinical detail the model sees Degraded None Preserved PHI leaves the hospital Partial No No Usable draft returned Partial No Yes Clinician still types by hand Often Yes No Is your PHI workflow ready for a model? Run this quick self-check against any clinical task you want a model to take on. If you answer no to more than one, the enablement pattern is likely what is missing. Can you identify which fields carry patient identity and which carry clinical meaning, separately? Does your current approach preserve the case timeline and dosage arcs when it hands text to a model? Does the drafting run inside an environment your security team already trusts? Do the real identifiers return to the finished document without a human retyping them? Can your privacy officer trace, for any output, that no protected health information crossed the boundary? Where it fits This is healthcare's version of sensitive AI workflow enablement. The same pattern appears in finance, where the protected document is a confidential operational document rather than a chart, and the logic is identical: send the structure, restore the values locally. The underlying mechanics, substitution and local reconstruction, are covered in Substitute, Execute, Reconstruct. All of it runs on the CUBIG Syntitan platform. Because no protected health information crosses the boundary, the workflow maps cleanly onto HIPAA-style obligations for handling PHI, and it fits the direction clinical AI oversight is already moving: the FDA treats AI in software as a medical device as something to supervise across the total product lifecycle, with its Good Machine Learning Practice principles underneath. The privacy officer and the security team get a concrete basis to sign off. That approval is the by-product. The point was to let the clinician stop typing summaries at midnight. LLM Capsule is how CUBIG runs HIPAA-compliant AI on clinical workflows without PHI leaving the environment. Patient identifiers are substituted before anything reaches the model and restored locally, so the work runs and the records stay put. --- title: "Financial AI Workflow on Confidential Documents" url: "https://cubig.ai/articles/financial-confidential-documents/" source: live --- A financial AI workflow runs AI document analysis on contracts, credit files, and pricing memos by sending the model the structure of the work, not the raw values inside it, so the analysis runs while the confidential operational documents stay inside the institution. The gap between what a model can do and what a bank will let it touch is where most of this stalls. Even so, in regulated finance the block is rarely the model. It is that the work lives in documents no one is allowed to send anywhere. That constraint now has its own tooling: research toolkits like OnPrem.LLM exist to run models on sensitive, non-public data in offline or restricted environments. Picture a credit team with forty renewal files to summarize before the committee meets on Friday. Each file is a stack: the signed facility agreement, the counterparty's last three quarters of management accounts, an internal pricing memo with the spread the desk is willing to hold, and a risk note that names two other clients for comparison. Reading that stack, surfacing the covenants that changed, and flagging where pricing sits against the book is exactly what a capable model does well. It is also the most sensitive paper the institution owns. So the file sits there, and someone summarizes it by hand at 9pm. AI document analysis on confidential financial documents The two dead ends in a financial AI workflow When a financial AI workflow meets a confidential operational document, teams usually reach for one of two moves, and both end badly. The first is to clean the document before sending it: strip the counterparty name, blank the pricing, redact the comparison clients. What comes back is a summary of a document that no longer says anything. The spread was the point. The named comparisons were the point. Remove them and the model reasons about a hollow shell, and the analyst has to reinsert every removed fact from memory before the output is usable. The second is to forbid the external model entirely and wait for an internal one that may never match it. Meanwhile the analysts who have a Friday deadline paste excerpts into whatever chat window is open on a personal machine, and now the institution carries an exposure it cannot even see. Both paths share one hidden assumption: that the raw values and the analytical work are the same object. They are not. How AI document analysis runs on a confidential document The move is to send the model the structure of the work and reconstruct the result locally. The renewal file keeps its shape. Counterparty A stays a single consistent entity across all four documents. The pricing keeps its relationships: this spread wider than that one, this covenant tighter than last year's. What changes is that the real name, the real basis points, and the real identities of the comparison clients are substituted before anything leaves, so the model reasons over a faithful skeleton of the deal rather than the deal itself. The model returns its committee summary against that skeleton, and the institution reconstructs the result inside its own systems, slotting the real counterparty name and the real numbers back where the placeholders were. The analyst opens a finished summary, not a puzzle. None of the original values made the trip out, and the whole flow runs inside the environment the bank already operates: the cloud tenant or on-prem cluster that already holds the data. This runs on a Context-Preserving Data Layer for AI, in CUBIG's case LLM Capsule, which holds the structure intact on the way to the model and rebuilds the business meaning on the way back. The substitution is structure-preserving, which is what separates it from plain masking. A masked contract loses its relationships; a substituted one keeps them and only swaps the values. LLM Capsule runs on the CUBIG Syntitan platform. What the model actually receives It helps to see what AI document analysis actually receives. The table below contrasts the three ways a credit file reaches a model, using the same renewal stack in each row. Comparison of what the model sees and whether the output remains usable across three financial AI workflow approaches Approach What the model sees Is the output usable? Raw upload Real names, spreads, comparison clients △ Yes, but the values left the boundary Plain masking / redaction Blanks where the facts were; relationships broken ✕ No, analyst rebuilds it by hand Structure-preserving substitution A faithful skeleton: same entities, same relationships, substituted values ✓ Yes, and values never left The third row is the one that lets AI document analysis run. The model gets enough structure to reason well, and the confidential operational documents never leave the institution's control. A quick test for your own files Run this check against a document your team currently will not send to a model: If you removed the sensitive values, would the model still have enough to do the work? If not, redaction is not your answer. Does the same entity, say one counterparty, appear across several documents that must line up in the output? Do relationships between values carry meaning, such as one spread relative to another? Would the finished output need the real names and numbers put back before anyone could use it? Can the whole flow stay inside the cloud or on-prem environment you already run? If most of these are yes, you are looking at a workflow that structure-preserving substitution can open, not a redaction problem. Where it fits This is one industry view of sensitive AI workflow enablement, the broader practice of running AI on data you are not allowed to expose. The same pattern carries into other regulated settings; see how it plays out for clinical workflows on PHI, where the document is a chart instead of a credit file, and how the boundary changes when the sensitive text lives inside images and scanned pages in visual sensitive data. One more thing falls out of running the workflow this way. Because the raw contract values and customer identities never leave the boundary, the bank's own data-privacy and financial-confidentiality obligations have a clean answer, and security and legal reviewers get a concrete basis to approve. Supervisors point the same way: BCBS 239 expects banks to aggregate risk data accurately and completely, leaning on automation rather than manual steps to keep the probability of error low, a standard a workflow that ends in retyping was never going to meet. The systemic view lines up: the Financial Stability Board lists model risk alongside data quality, governance, and cyber risk among the AI vulnerabilities it watches most closely. That approval is a by-product of how the work runs, not the reason to run it. LLM Capsule is how CUBIG runs AI document analysis on financial documents without the figures leaving. Statements and schedules are substituted in place, so the model reads the real structure while the numbers stay inside your environment. --- title: "Operational Artifact: The Real Unit of AI Delivery" url: "https://cubig.ai/articles/operational-artifact/" source: live --- An operational artifact is the real unit of AI delivery: a finished business document, reconstructed inside your own environment with the real values in place, that a person can send, file, or sign without rebuilding it first. It is not raw model text, and it is not a chat reply someone still has to clean up. It is the thing the work was actually for. Most conversations about AI output stop at the model's response, and that framing quietly sets the bar too low. Regulation already refuses that framing: Article 12 of the EU AI Act obliges high-risk AI systems to record events automatically across their entire lifetime, precisely so that results stay traceable after the fact. The cost of stopping short is measurable: S&P Global Market Intelligence found the share of companies abandoning most of their AI initiatives jumped to 42%, up from 17% a year earlier. Deloitte's enterprise survey points to why: regulatory-compliance concern has become the top barrier to deploying generative AI. When you measure delivery by the answer instead of the finished artifact, the work that turns an answer into a usable document stays invisible until it starts eating the time the AI was supposed to save. The gap between a response and a result A model returns text, and often the text is good. But a response is not yet a result. The renewal letter the model drafted still carries placeholder names where the customer's real details belong. The reconciliation it produced references stand-in figures, because the real numbers never left your systems. Someone on the team then has to take that response and turn it back into the document the business needed. That last mile of reassembly is invisible in a demo and very visible in production. It is manual, it is easy to get wrong, and it scales badly: every run of the workflow inherits the same cleanup. The demo looked finished; the workflow was not. Why the gap exists at all The gap is structural, not a bug in the prompt. When you run AI on confidential data the right way, the model never sees the real values, so it works on substitutes and its output is naturally keyed to those substitutes. A response built on substitutes is honest about what it is; it is simply not done. The model still understands the task because you send it the work's structure, not the raw values. For why the model works on substitutes in the first place, see confidential business context. The point of the approach is to substitute, execute, and then reconstruct. The substitution keeps the real data in place. The reconstruction is what closes the distance between a plausible answer and a finished artifact. What makes an operational artifact operational An artifact is operational when three conditions hold at once. First, it carries the real values, restored from the substitutes the model worked on. Second, it lands in the format the team already uses: the actual letter, report, or record, not a block of text to reformat. Third, it is complete enough to act on, so the next step is to send it, file it, or sign it rather than rebuild it. Getting there takes an explicit reconstruction step. The model's response comes back, and inside your environment the stand-ins are reversed to the real values while the output is reassembled into its proper form. What leaves the boundary is structure; what returns to the user is an operational artifact. That reconstruction is performed by a Context-Preserving Data Layer for AI, in CUBIG's case LLM Capsule, which runs on the CUBIG Syntitan platform. Response versus operational artifact The distinction is easier to hold once you put the two side by side. A response is what the model hands back; an operational artifact is what the workflow is measured on. Dimension Model response Operational artifact Values shown Substitutes Real values, restored locally Format Raw text to reformat The document the team already uses Next action Reassemble by hand Send, file, or sign Where it forms At the model Inside your environment What you measure A good-looking answer Work that actually closed Why naming the unit of AI delivery changes what you measure Naming the unit changes the definition of success. If the deliverable is "a model response," success is a good-looking answer, and the reassembly work hides in someone's afternoon. If the deliverable is an operational artifact, success is a document that closes the task, and you can see immediately whether the workflow actually finished. That is the difference between AI that demos and AI that delivers, and it is what workflow closure is built to deliver. It also sharpens how you plan enterprise AI adoption. Teams that budget for the artifact, not just the answer, stop being surprised by the hidden cost of the last mile, because they measured for it from the start. A quick self-diagnostic Run your current AI workflow through these questions. If you answer "no" to any of them, you are shipping responses, not artifacts. When the model finishes, does a real, usable document come out, or does someone still have to reassemble it? Does the output already carry the real values, or does it still show placeholders and stand-in figures? Is the deliverable in the format the team actually uses, or is it text waiting to be reformatted? Can the next person send, file, or sign it, or do they have to rebuild it first? Do you track the finished artifact, or only whether the answer looked right? Where it fits in sensitive AI workflow enablement The operational artifact is the output side of sensitive AI workflow enablement. Structure crosses the boundary on the way out, as described in LLM data egress, and an operational artifact comes back on the way in, reconstructed locally so the real values never had to travel. All of it runs on the CUBIG Syntitan platform, which lets a regulated team complete confidential work with an AI model without sending the model anything it should not hold. LLM Capsule is what turns model output into an artifact your team can use. The result is reconstructed inside your environment with the real values restored, so what comes back is a finished document rather than a placeholder. That finished document—not the model's reply—is the real measure of AI delivery. --- title: "Synthetic Data for Clinical Research: Rare Cohort Augmentation with DTS" url: "https://cubig.ai/articles/dts-for-clinical-research/" source: live --- Synthetic patient data for clinical research is data rebuilt to preserve the statistical structure of a real patient population, including its rare cohorts, so a model can learn on it when the original records are too restricted, too scarce, or too imbalanced to use directly. The hard part of clinical research is rarely the common case. It is the rare cohort: the patients with an uncommon comorbidity, the responders who react badly to a standard therapy, the subgroup that shows up a few dozen times across a decade of records. Those cases carry the clinical signal that matters, and they are exactly the cases a model sees too rarely to learn. They are also among the most closely guarded records an institution holds, which is why an npj Digital Medicine paper argues for patient-centric synthetic data generation in biomedical work: it takes the re-identification risk of analyzing real patient records off the table. Why clinical research runs short on the cases that matter Three constraints tend to collide in a research dataset. Access is limited because the records are protected health information governed by HIPAA and institutional review; a modeler often cannot see raw values at all. The sanctioned routes are narrow, too: HIPAA recognizes exactly two de-identification paths, Expert Determination and Safe Harbor, with Safe Harbor requiring the removal of 18 specified identifiers and no actual knowledge that whatever remains could re-identify a patient. Volume is limited because a rare condition may appear in a fraction of a percent of patients, so even a large cohort yields too few positive examples to train on. And balance is skewed, because the healthy majority swamps the clinically interesting minority. You can wait for more data, but rare cases do not accumulate on a schedule. You can oversample the minority, but naive duplication teaches a model to memorize a handful of patients rather than learn the underlying pattern. Neither path gives a research team a dataset it can actually model, and both leave the rarest, most valuable cohorts underrepresented right when a study depends on them. What DTS does with clinical data DTS, CUBIG's AI-ready data transformation engine, works in three steps: Diagnose, Transform, and Rebuild. It diagnoses where a dataset blocks modeling, then rebuilds the blocked parts into data a model can learn on, using structure-preserving synthesis rather than masking or simple copying. DTS runs on the Syntitan platform, so the rebuilt dataset carries a data state a research team can reference later. Diagnose. DTS profiles the cohort and finds the constraints: which fields are too sparse, which subgroups are underrepresented, where class imbalance will bias a model, and which correlations between variables have to survive for the data to stay clinically meaningful. Transform. Using a non-access architecture, DTS learns the joint distribution of the population without a modeler ever handling raw identifiers. Differential privacy bounds how much any single patient can influence the result, which matters when a cohort is small enough that one record could otherwise stand out. Rebuild. DTS generates records that keep the statistical structure of the original, including the correlations between age, lab values, comorbidities, and outcomes, and it augments the rare cohorts so the minority pattern is represented well enough to learn. This is rare-pattern augmentation: not copying the few real rare patients, but reconstructing the shape of the rare pattern so a model generalizes from it. Synthetic patient data versus plain masking It helps to separate structure-preserving synthesis from the older tools it is often confused with. Plain masking hides values but leaves the same thin, imbalanced dataset behind, so the rare-cohort problem is untouched. Anonymization removes identifiers and, in doing so, frequently damages the correlations a clinical model depends on. Structure-preserving synthesis takes a different route: it rebuilds the population so the structure holds and the rare cases are learnable, without exposing any real patient. Comparison of data privacy approaches by rare-cohort learnability, correlation preservation, and raw PHI exposure Approach Rare cohort learnable? Correlations preserved? Raw PHI exposed? Plain masking ✕ No △ Partial △ Partial Anonymization ✕ No ✕ No ✓ No Naive oversampling △ Partial △ Partial ✕ Yes Structure-preserving synthesis ✓ Yes ✓ Yes ✓ No The distinction is practical, not academic. A masked dataset still fails a rare-disease model for the same reason the raw one did: there are not enough positive examples. A rebuilt dataset changes that, because the rare pattern is represented at a density a model can actually learn from. A worked example: augmenting a rare responder cohort Consider a study of an adverse drug response that shows up in roughly one patient in a thousand. A hospital has ten years of records and perhaps forty confirmed cases, spread across different ages, baseline conditions, and dosages. Forty examples will not train a reliable classifier, and the forty are protected, so they cannot leave the hospital's environment. DTS diagnoses the cohort in place, learns the joint distribution that links the baseline factors to the response, and rebuilds a larger synthetic cohort in which the responder pattern appears often enough to model, with the age and dosage correlations intact. The research team trains on the rebuilt data, then validates against the real forty cases held back for the test. Results are representative until the team reproduces them on its own held-out data, which is the standard any clinical claim should meet before it moves forward. Here, synthetic patient data does not invent a signal—it rebuilds the distribution the rare responders already imply, at a volume a model can learn from. How to tell if your cohort needs rebuilding Run this quick check against a study that is stalling: Your target condition or outcome appears in under a few percent of the cohort, and the model keeps predicting the majority class. You cannot get raw record access cleared for modeling within the study's timeline. Oversampling or class weights improved training metrics but the model failed on new patients. The rare subgroups you care most about have too few examples to split into train and test. An anonymized extract broke the correlations your clinicians said were essential. You need to explain later how the training data for a result was produced. If three or more of these are true, the blocker is the data, not the model, and rebuilding the cohort is the step that unblocks the study. When those signs line up, synthetic patient data is usually the fastest route to a dataset a model can actually train on. Where it fits in AI-ready clinical work DTS handles one specific job: turning restricted, scarce, or imbalanced clinical data into data a model can learn on. Because it runs on the Syntitan platform, the rebuilt dataset is not a loose file with no history. The platform records the data state behind each rebuild, so a team can reference which version of the data trained a given model and reproduce that state later, which is what auditors and review boards tend to ask about first. The regulatory trajectory points the same way: the FDA approaches AI in software as a medical device with oversight spanning the total product lifecycle, grounded in Good Machine Learning Practice guiding principles. For teams weighing rebuilding against protection-only tools, our comparison of PHI masking versus synthetic data lays out where each one fits, and what DTS is covers the engine end to end. The related question of reproducing a clinical result is covered in reproducible cohort analysis. That is what makes synthetic patient data usable for a model, not just safe to store. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "The DTS Transform Step in Depth: How Structure-Preserving Synthesis Rebuilds Your Data" url: "https://cubig.ai/articles/the-transform-step-in-depth/" source: live --- The DTS Transform step is the stage where CUBIG's AI-ready data transformation engine rebuilds restricted, scarce, or unusable data into a synthetic dataset a model can actually learn on, preserving the statistical structure of the original without ever exposing its raw records. The gap between a promising pilot and a model that stalls in production is usually the data, not the algorithm, and regulation keeps the sanctioned routes narrow. HIPAA recognizes just two de-identification paths for health data: a formal Expert Determination, or Safe Harbor, which strips 18 enumerated identifiers and asks only that you hold no actual knowledge that what remains could identify someone. The DTS Transform step exists to close that gap: it takes the diagnosis of what blocks your data and produces a rebuilt dataset that keeps the patterns a model needs while dropping the exposure and access problems that kept the original locked away. Where the Transform step sits in DTS DTS, CUBIG's AI-ready data transformation engine, runs three stages in order: Diagnose, Transform, Rebuild. Diagnose maps what makes your data unusable, whether that is regulatory restriction, a class that appears too rarely to train on, or fields you cannot move off-premises. Transform is the middle stage that acts on that map. Rebuild is the validation and delivery stage that confirms the output holds up. Transform is where the actual synthesis happens. It reads the structure the Diagnose stage recorded, then generates new records that reproduce the joint distributions, correlations, and rare patterns of the source, without copying any individual row. Because the engine works from structure rather than from access to the raw values, the sensitive originals never have to leave their environment. This is what we call a non-access architecture, and it is the reason regulated teams can run synthesis on data that compliance would otherwise keep off the table. What Structure-Preserving Synthesis actually does Plain synthetic data generators sample from a learned distribution and stop there. That is fine for a demo and thin for training, because the relationships that carry signal, the way a fraud flag co-occurs with a transaction sequence or the way a lab value tracks a diagnosis, tend to wash out. Structure-Preserving Synthesis is DTS's answer to that problem. The Transform step preserves the multivariate structure of the data: the correlations between columns, the conditional distributions, and the tail cases that a model has to see in order to generalize. Two mechanisms do most of the work. The first is joint modeling, which learns columns together rather than one at a time, so a synthetic record is internally consistent instead of a bag of independently plausible values. The second is rare-pattern augmentation, which deliberately reproduces and, where useful, amplifies the low-frequency cases that matter most for classification but appear too seldom in the raw data to train on reliably. Differential privacy is available as a technique inside the Transform step for teams that need a formal privacy guarantee on the output. It bounds how much any single source record can influence the synthetic result, which lets you tune the balance between privacy strength and how faithfully the rebuilt data tracks the original. Transform versus plain synthetic data generation The difference is easiest to see side by side. A generic generator optimizes for records that look real one at a time. The DTS Transform step optimizes for a dataset that behaves like the original when a model trains on it, which is a harder and more useful target. Property Plain synthetic data generation DTS Transform step Column relationships Often lost Preserved (joint modeling) Rare and tail cases Under-represented Reproduced, augmented on demand Access to raw records Usually required Non-access architecture Formal privacy option Rarely built in Differential privacy available Success measure Looks realistic Trains a comparable model How the Transform step runs, step by step In practice a Transform run moves through a short, repeatable sequence. The engine reads the structural profile from Diagnose, fits a joint model to the source patterns, generates candidate synthetic records, and applies any privacy budget you set. It then checks the candidates against the recorded structure before handing them to the Rebuild stage for full validation. Rare-pattern handling deserves a note here, and the lineage is long: SMOTE, the 2002 JAIR technique, established synthetic minority-class examples as the standard answer to class imbalance. If your fraud data carries positives at a fraction of a percent, a naive sample will hand your model almost none of them. The Transform step can hold those cases in view and rebuild them at a density the model can learn from, so the signal survives into training instead of being rounded away. Is your dataset a candidate for the Transform step? Run this quick check against the data blocking your project. If two or more of these are true, the Transform step is worth a trial. Regulation or contract keeps you from moving the raw records to where the model trains. The class you care about, such as a defect or a fraud case, is too rare to train on reliably. Column relationships carry the signal, so shuffled or masked data loses the point. You need a dataset you can share with a vendor or a partner without exposing customers. You want a formal privacy guarantee you can show an auditor, not just a promise. Where it fits in the CUBIG operating layer The Transform step is one stage of DTS, and DTS runs on the Syntitan platform, CUBIG's AI-Ready Data Platform. The division of labor is clean: DTS rebuilds data that was restricted, scarce, or unusable into something a model can learn on, and the platform around it scores enterprise data for readiness and keeps the rebuilt output governable. Because the Transform step records how it produced each synthetic dataset, the output is not a one-off artifact you lose track of; it is a data state your team can point to when someone asks where a model's training data came from. For high-risk systems in Europe that question carries legal force, since Article 10 of the EU AI Act expects training, validation, and testing data to sit under documented data-governance practices and to meet quality bars for relevance, representativeness, and completeness. For teams comparing options, the honest framing is that Transform is a rebuild engine, not a masking tool. Masking removes values; Transform reconstructs the patterns those values carried so a model still has something to learn from. If you want the full picture of the engine, start with the pillar on what DTS is, then read how the Diagnose step feeds it and how DTS validation confirms the output is good enough to trust. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "DTS Validation: How Do You Know Synthetic Data Is Good Enough?" url: "https://cubig.ai/articles/dts-validation/" source: live --- DTS validation is the acceptance step that proves rebuilt data is good enough for AI: it checks that synthetic data preserves the statistical structure a model needs to learn while holding a measured privacy guarantee, before that data is trusted for training or evaluation. Teams rebuild restricted data, then use it on faith. That is the gap DTS validation closes. Even a formal guarantee needs its own inspection: NIST wrote SP 800-226 to guide the evaluation of differential privacy guarantees, and the document catalogs the recurring implementation mistakes it calls privacy hazards. If you cannot state, in numbers, why the synthetic data is trustworthy, you have not finished the transformation. You have only started it. What DTS validation actually measures DTS is CUBIG's AI-ready data transformation engine. It runs a three-move loop, Diagnose, Transform, Rebuild, to turn restricted, scarce, or unusable data into data a model can learn on, using a non-access architecture and structure-preserving synthesis. Validation is the check that sits after Rebuild and before release. Good enough is not one number. It is an answer to two questions that pull against each other. Does the rebuilt data carry the patterns a model needs, and does it keep the privacy guarantee that let you use the original in the first place? Push fidelity too hard and you leak; push privacy too hard and the data goes flat. DTS validation reports both sides so the tradeoff is a decision you make on evidence, not a hope you hold. In practice we group the checks into four families: statistical fidelity, whether the shape of the data holds; utility, whether a model trained on synthetic performs when tested on real; privacy, whether an attacker can recover a real record; and coverage, whether the rare but important patterns survive. A dataset that clears one family and fails another is not validated. It is a warning. The four families of checks, in plain terms Statistical fidelity compares the rebuilt data against the original distribution: per-column marginals, pairwise correlations, and the joint structure that generic tools tend to smear. For DTS the bar is higher than "the averages match," because structure-preserving synthesis exists to keep the relationships between fields, not just each field on its own. Utility is the test that matters to the people who will actually ship a model. Train on synthetic, test on real, and see whether the score lands close to a model trained on the real data. This is often called TSTR, train-on-synthetic test-on-real, and it is the most honest signal that the rebuilt data is fit for its job. Privacy asks the adversary's question. Run membership-inference and nearest-neighbor distance checks to confirm that no synthetic record is a thin disguise for a real one. When DTS applies differential privacy during transformation, the privacy budget gives you a stated, auditable bound rather than a vibe. Coverage asks whether the rare but important patterns survived the rebuild. Fraud, rare defects, and unusual clinical cohorts are exactly the cases a model most needs, and they are the first to vanish when a generator smooths the distribution. DTS treats rare-pattern augmentation as part of the rebuild, so validation confirms those minority classes and edge cases are still present rather than assuming they are. How the validation metrics compare The table below lays out the check families, what each one answers, and the failure it catches. Read it as a checklist, not a menu: a release should clear all of them, at thresholds you set before you look at the results. Check family Question it answers Failure it catches Statistical fidelity Does the rebuilt data keep column shapes and correlations? Flattened joints, lost rare combinations Utility (TSTR) Does a model trained on synthetic hold up on real data? Data that looks right but does not train Privacy Can an attacker link a synthetic record back to a person? Memorized rows, re-identification Coverage Are rare but important patterns still present? Dropped edge cases, minority classes Why one good number is a trap A high fidelity score with no privacy check is the classic failure. The rebuilt data matches the original so closely that some rows are near-copies, which means an attacker who holds one real record can confirm it sits in your training set. The opposite trap is just as common: a clean privacy budget on data so noised that a model learns nothing useful, and the utility test quietly tanks. Coverage is the quiet third failure. Fraud, rare defects, and unusual clinical cohorts are exactly the patterns a model most needs, and they are the first to vanish when a generator smooths the distribution. The UK FCA's Synthetic Data Expert Group report walks through the financial-services uses that lean on those very patterns: fraud controls, model testing and validation, and data sharing. DTS treats rare-pattern augmentation as part of the rebuild, so validation checks that those patterns survived rather than assuming they did. Run this self-diagnostic before you trust rebuilt data If you are about to hand synthetic data to a training team, walk this list first. Any "no" is a reason to hold the release. Did you set fidelity, utility, and privacy thresholds before generating, not after seeing the scores? Do you have a TSTR result, trained on synthetic and tested on a held-out real set? Is there a stated privacy bound, for example a differential-privacy budget, tied to this specific dataset? Did you run a membership-inference or nearest-neighbor check for leaked records? Did you confirm the rare classes and edge cases you care about are still represented? Can you reproduce every one of these numbers from the same inputs on demand? Where DTS validation fits in an AI-ready pipeline Validation is not a certificate you file and forget. It is the moment where rebuilt data either earns a place in production or gets sent back. For production AI the real question is not only which model ran but which data state and execution conditions produced the result. DTS runs on the Syntitan platform, CUBIG's AI-Ready Data Platform, which scores enterprise data on six axes, rebuilds what blocks execution, and binds every AI or agent run to a data state you can diff and reproduce. That binding is what turns a validation report from a one-time score into something durable. When a regulator or an internal reviewer asks why a model made a decision, you can point to the exact rebuilt data state, its validation numbers, and the run that used it. Science already knows what happens without that discipline: in Nature's survey of 1,576 researchers, more than 70% reported trying and failing to reproduce a peer's experiment. A validated, reproducible data state is how DTS keeps its own evidence out of that category, at the data layer rather than after the fact. If you want the deeper mechanics, see how structure-preserving synthesis keeps field relationships intact, read the Transform step in depth, or start from the pillar, what DTS is. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "Structure-Preserving Synthesis vs Generic Synthetic Data: 6 Differences That Decide Production" url: "https://cubig.ai/articles/structure-preserving-synthesis-vs-generic-synthetic-data/" source: live --- Structure-preserving synthesis is a data rebuild method that reproduces the statistical relationships, rare patterns, and schema of a restricted source dataset, so a model can learn on it, whereas generic synthetic data optimizes for surface realism and often loses the correlations and edge cases that decide whether the model works. The gap between the two shows up late, in production, where it is expensive. Synthetic data has been correcting class imbalance since long before today's generators: SMOTE, the canonical JAIR 2002 technique, generates synthetic minority-class examples for exactly that purpose. A synthetic dataset that looks right on a dashboard but drops the rare pattern your fraud model needs is quietly failing the assignment that technique was built for. This article compares the two approaches on the terms that matter for a model in production: whether the rebuilt data carries joint distributions, whether it holds the rare cases, whether the schema survives, and whether you can defend the privacy claim. If you are choosing between a general synthetic data generator and a structure-preserving rebuild, the table below is the decision. What structure-preserving synthesis actually preserves Generic synthetic data answers a narrow question: does each column, looked at on its own, resemble the real thing? Age distributions match, transaction amounts fall in a plausible range, category frequencies look familiar. That is enough to demo, and it is enough for many privacy checkboxes. It is often not enough for a model. A model does not learn from columns in isolation. It learns from the relationships between them: the correlation between income and default, the sequence of events that precedes a machine failure, the layout of a form that tells a document parser where the total sits. Structure-preserving synthesis is built to reproduce those relationships, not just the marginals. Three things it holds that generic generators tend to lose. First, joint distributions, meaning the way variables move together rather than each variable's shape alone. Second, rare patterns, the low-frequency events like fraud, defects, or unusual clinical presentations that a model exists to catch. Third, the schema and its constraints, so foreign keys, valid ranges, and required fields still hold after the rebuild. Lose any one of these and the synthetic set trains a model that scores well offline and fails on the cases you care about. Structure-preserving synthesis vs generic synthetic data, side by side Here is the comparison on the dimensions that change a production outcome. "Generic synthetic data" here means a general-purpose generator tuned for realism; "structure-preserving synthesis" is the rebuild approach used by CUBIG's DTS. Dimension Generic synthetic data Structure-preserving synthesis Per-column realism Yes Yes Joint distributions across columns Partial Yes Rare pattern retention No, often smoothed away Yes, augmented on purpose Schema and constraints preserved Partial Yes Privacy claim you can defend Varies by tool Differential privacy budget, non-access build Fit for training and validation Weak on edge cases Built for the edge cases The row that decides most enterprise use is rare pattern retention. A generic generator learns the common shape of the data and reproduces it well, which means it reproduces the majority class well and the minority class poorly. For a fraud, anti-money laundering, or defect model, the minority class is the whole point. Why the rare pattern gap is the expensive one Consider a card fraud model. Fraud is well under one percent of transactions. A generic synthetic generator, trained to reproduce the overall distribution, will faithfully reproduce that: less than one percent of its rows will be fraud, and the fraud rows it does produce will be blurred toward the average because the model had few examples to learn the pattern from. Structure-preserving synthesis takes the opposite stance on the rare class. Because it models the structure that separates fraud from normal activity, it can rebuild the rare pattern faithfully and, where you ask for it, augment it: generate more valid rare cases so the downstream model has enough signal to learn from. This is rare-pattern augmentation, and it is the difference between a model that flags novel fraud and one that only recognizes what it has already seen. Financial supervisors are studying the same ground: the UK FCA's Synthetic Data Expert Group report examines synthetic data across fraud controls, model testing and validation, and data sharing in financial services. The same logic runs through clinical research, where the cohort you need to study is small by definition, and through manufacturing, where the defect you want to catch happens once in ten thousand parts. In each case the rare rows are the reason for the project, and a generator that smooths them away has quietly deleted the project's value. Privacy is not the same as realism Both approaches promise privacy, and it is worth being precise about what each one actually offers, because a demo of realistic looking rows tells you nothing about re-identification risk. That risk is precisely what a paper in npj Digital Medicine argues patient-centric synthetic data generation removes in biomedical work, where the alternative is analyzing real patient records. A defensible privacy claim rests on two things: a formal privacy guarantee and a build process that never exposes the raw records. Differential privacy provides the first, a mathematical bound on how much any single individual's data can influence the output, tunable through a privacy budget. A non-access architecture provides the second, meaning the rebuild runs without a human or a downstream system reading the sensitive rows directly. Structure-preserving synthesis as CUBIG implements it in DTS combines both. Many generic generators offer neither in a form you can put in front of an auditor. Run this self-diagnostic before you trust a synthetic set Does the synthetic set keep the correlation between your two most predictive columns, or only their individual shapes? If you filter to the rare class you actually care about, are there enough realistic rows to train on? Do foreign keys, valid ranges, and required fields still hold, or did the rebuild produce impossible records? Can you state the privacy guarantee as a number, or only as "it looks anonymized"? Was the raw data ever read directly during the build, or was it a non-access process? If you cannot answer two or more of these, you are likely looking at generic synthetic data wearing a structure-preserving label. Run the checks before you train. Where it fits in CUBIG's operating layer Structure-preserving synthesis is the method behind DTS, CUBIG's AI-ready data transformation engine. DTS diagnoses what makes a dataset unusable for AI, transforms it while holding the structure that carries the signal, and rebuilds restricted, scarce, or unusable data into data a model can actually learn on. It applies differential privacy and a non-access architecture so the rebuild is defensible, not just realistic. DTS runs on the Syntitan platform, CUBIG's AI-Ready Data Platform, alongside the rest of the data-readiness work an enterprise needs before a model goes to production. If your blocker is not scarcity or restriction but running AI on confidential data you cannot move, that is a different tool in the same family; the comparison in DTS vs anonymization vs differential privacy walks the boundary, and this decision guide helps you tell which one your project needs. For the readiness foundation underneath both, see what AI-ready data means. How to choose If your model works on common, high-frequency behavior and you mainly need to de-risk a demo or a data-sharing agreement, a generic synthetic data generator may be all you need. If the value of your project lives in the rare class, in the correlations between fields, or in a privacy claim you will have to defend to a regulator, structure-preserving synthesis is the approach that survives contact with production. The honest test is not which one looks more realistic on a chart; it is which one still holds up when you train on it and check the cases that matter. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "Confidential Business Context: The Data Beyond PII That Blocks AI" url: "https://cubig.ai/articles/confidential-business-context/" source: live --- Confidential business context is the sensitive information that has nothing to do with personal data: pricing logic, contract terms, trade secrets, internal metrics, and operational detail that describes how a company competes, and it quietly blocks more enterprise AI projects than personal data ever does. When people hear "sensitive," they picture personally identifiable information: names, account numbers, health records. Two decades of privacy law trained everyone to look there. Yet the data that actually stalls an enterprise AI project is usually not personal at all. Samsung's 2023 episode is the canonical case: the company temporarily banned generative AI on company devices after employees leaked sensitive internal data to ChatGPT, and what leaked was corporate material, not customer records. The pattern holds across enterprises: a Harmonic Security analysis found that customer records, legal and finance material, and security details, not just classic PII, make up most of the sensitive data employees paste into AI tools. And the gap is systemic: Stanford HAI's 2025 AI Index notes organizations acknowledge these risks while their mitigation efforts lag. Projects like these stall for a reason no privacy tool addresses: the data the model needs is commercially sensitive, and no one can safely hand it over. What confidential business context actually covers Walk through one contract and it comes into focus. The counterparty name might sit near personal data. The negotiated discount, the renewal terms, the volume commitments, and the penalty clauses do not. None of that is personal, and every line of it is information the company would never want in a rival's hands or in a model's training set. Pricing strategy, margin assumptions, the internal metric that reveals which product line is carrying the quarter: this is the material that decides who wins the deal. Confidential business context. Sensitive information whose value is commercial rather than personal, where the sensitive part and the useful part are the same part. That last property is what makes it different. Redact a customer name and you can still summarize the document. Redact the pricing and the contract logic, and nothing worth summarizing remains. The sensitivity lives inside the exact content that carries the meaning, so removing it does not just create privacy risk, it destroys the task. Why it blocks AI projects more than PII does Personal data has a mature playbook. Teams mask it, tokenize it, or drop the column, and legal knows those techniques cold. Confidential business context has no equivalent, because it is not a field you can strip. It is woven through the numbers, the relationships between clauses, and the phrasing of a single term. There is no clean seam between the sensitive part and the useful part, so the usual reflex of removing the sensitive column leaves you with a document a model cannot reason about. The project then stalls in a familiar spot. A data team wants to run a model across the renewal portfolio. Legal reviews what would leave the building during LLM data egress and declines. The compromise, a version with commercial terms stripped out, produces output too generic to act on. Everyone agrees the AI would help, and nobody can get the data to it. The blocker was never personal data; it was the business context the whole exercise depended on. How plain masking falls short here To see why enablement differs from masking, compare what each does to the same confidential contract. Approach What the model receives Task survives? Send raw document Real pricing and terms leave the environment Yes, but egress is blocked Plain masking / redaction Blanked values, broken logic No Strip commercial fields A generic shell with no numbers No Structure-preserving substitution Faithful stand-ins, real relationships intact Yes Plain masking treats every sensitive token as noise to be blotted out. That works for a stray Social Security number. It fails on a pricing table, because the model still needs the shape of that table, the way one clause references another, and the fact that a penalty triggers on a specific condition. Blank those out and you have not protected the work, you have removed it. Operating on it without exposing it The resolution applies the same enablement pattern that works for personal data to a harder target. You do not have to choose between exposing commercial terms and gutting them. You send the model the structure of the work rather than the raw values: the contract keeps its clauses, its tables, and its internal references, while the real pricing and terms are represented by faithful stand-ins. The model reasons over a coherent document, and the real values are reconstructed locally so the result is usable. What makes this fit confidential business context specifically is that structure is preserved rather than blanked. A masked contract loses the logic the model relied on. A structure-preserving copy keeps the relationships, so this price maps to that term and this penalty follows that condition, while the literal values stay inside your environment. The model gets a real problem to solve, and the company's commercial detail never crosses the boundary. Performance you see on a substituted document is representative until you reconstruct and confirm it on your own workflow. A quick self-diagnostic Run this test on any AI project that keeps stalling in review: If you removed every name and account number, would the remaining data still be too sensitive to send out? Does the value the model needs live in pricing, terms, margins, or internal metrics rather than in personal fields? When you strip the sensitive parts to get approval, does the output come back too generic to act on? Has legal blocked a project even though no personal data was involved? Do your current masking tools have no clean way to handle a document where the numbers are the point? Say yes to two or more, and the blocker is confidential business context, not privacy. That distinction changes which tool you reach for, because a privacy tool is built to remove sensitive values while the work here depends on keeping their structure. Naming the category correctly is the first step toward unblocking the project instead of watering it down. Where it fits Treating confidential business context as a first-class category is central to sensitive AI workflow enablement. It explains why the boundary cannot be solved by personal-data tools alone, and why reconstruction matters: the output has to return as a finished operational artifact with the real terms restored, which is how the workflow closes. This is also the step beyond simple redaction that separates enablement from plain PII masking. The substitution and reconstruction run through a Context-Preserving Data Layer for AI, in CUBIG's case LLM Capsule, which runs on the CUBIG Syntitan platform. LLM Capsule treats this business context as first-class, not just personal data. Pricing, terms, and internal figures are substituted before the prompt leaves, so the model reasons over the real shape while the values stay inside. --- title: "Multimodal AI Data Boundary: Running AI on One Whole Document" url: "https://cubig.ai/articles/multimodal-ai-data-boundary/" source: live --- A multimodal AI data boundary is a structure-preserving data boundary for inputs that mix more than one format in the same document (typed text, images, scanned PDFs, and tables), substituting the sensitive values inside each part while keeping the document whole, so an AI model can read across all of it at once and you can reconstruct the originals in place when the result returns. Enterprises are not short on ambition for multimodal AI; they are short on data their models can actually touch. NIST's Generative AI Profile names data privacy and information security among twelve risks it treats as unique to or amplified by generative AI, and mixed-format documents concentrate both risks in one document. Reading such a document at all takes multimodal models: Microsoft's LayoutLMv3 works by modeling text and image together, the unified view a format-by-format pipeline breaks apart. And converting formats to plain text is lossy in its own right: the Donut researchers note OCR errors propagate downstream into every step that follows. A single claim file might hold typed notes, a scanned form, a photo of a damaged part, and a spreadsheet of line items, with confidential values scattered across every one of them. Why single-format inputs are the easy case Plain text is the case everyone solves first. You find the sensitive spans, swap them for substitutes, run the model, then restore the originals. It works because text is one-dimensional: a sentence is a line of tokens, and a span has a clear start and end. The trouble starts the moment the input stops being a tidy paragraph. An image has no spans in that sense; sensitive text lives in pixels, positioned somewhere on a page. A scanned PDF is an image wearing the costume of a document. A table carries meaning in its grid, where the link between a cell and its header is part of the content, a point we cover in document layout preservation. Each format hides confidential data in a different place, and each one breaks in a different way when you handle it carelessly. Why splitting the document apart backfires The tempting shortcut is to pull each format out, process it on its own, and reassemble the pieces afterward. It rarely holds up, because the meaning often lives in the seam between the parts rather than inside any one of them. The damage photo only makes sense next to the claim note that references it. The signature image matters because of the contract clause sitting directly above it. Separate the pieces and you lose the relationships that made the document worth analyzing in the first place. A model reading the photo alone, with no idea which claim it belongs to, is doing a weaker and different job. So the boundary has to hold the document together on its way to the model, not quietly take it apart. What a multimodal AI data boundary does A multimodal AI data boundary applies one principle across every format at the same time. It follows a Substitute, Execute, Reconstruct pattern: sensitive values are substituted wherever they live (in the body text, inside the image, on the scanned page, within the table), the model executes on the whole coherent document, and the original values are reconstructed in place when the result comes back inside your environment. The arrangement that ties the parts together never leaves. Because the substitution is consistent, the model receives a document that still reads as one thing. The note still points at the right photo, the clause still sits above the right signature, and the table rows still line up under their headers. Sensitive text baked into an image gets the same treatment as text in the body, which is its own subject in visual sensitive data. That consistency is the whole point: a boundary that protects the typed paragraph but leaks the scanned page next to it is not a boundary, it is a gap with good intentions. How it differs from plain masking Plain masking answers a narrower question. It blacks out or redacts sensitive characters and hands the model a document with holes in it, and it usually treats each format in isolation. The redacted output is safer to store, but it is also worse to work with: the model loses the very values it needed to reason, and the cross-format links dissolve because each part was handled on its own track. A multimodal AI data boundary is built for enablement instead of redaction. It sends the model the structure of the work rather than the raw confidential values, keeps every part connected, then rebuilds the originals locally once the answer returns. The comparison below sets the two side by side. Approach to the document Answer quality on the document Keeps the document connected Plain masking or redaction Partial, the document loses the values the model needs No, each format is handled in isolation Multimodal AI data boundary Full, the model sees the real structure Yes, parts stay linked and originals rebuild locally How to tell if your inputs need a multimodal boundary Not every workflow needs this. If your data is uniform text, a simpler approach is fine. Run this quick self-check against a representative sample of the documents your teams actually send to AI: Does a single file routinely combine typed text with images, scanned pages, or tables? Does confidential data appear inside images or scanned PDFs, not only in editable text? Would the analysis lose meaning if you split the file into separate pieces? Do you need the original values back, in place, after the model runs, rather than a permanently redacted copy? Must the confidential values stay inside your environment while the work still gets done? Two or more yes answers is a strong signal that per-format masking will leave gaps, and that a boundary spanning all formats at once is the fit. It also reframes what leaves your network: the concern shifts from files to LLM data egress, the substituted stream that actually crosses the line to the model. Where it fits The multimodal boundary is the version of sensitive AI workflow enablement you reach for once your inputs stop being clean paragraphs and start looking like the documents your teams handle every day. In practice this is delivered by a Context-Preserving Data Layer for AI, in CUBIG's case LLM Capsule, which keeps the structure of every part intact on the way to the model and reconstructs the original values, in place, when the result returns. It runs on the CUBIG Syntitan platform, alongside the rest of the AI-ready data pipeline, so a mixed document does not force you to choose between using AI and keeping confidential values where they belong. LLM Capsule holds this boundary across formats, not just text. Tables, PDFs, and images travel through as structure while the sensitive content stays inside your environment. --- title: "Retrieved-Context Leakage in RAG: Where It Happens" url: "https://cubig.ai/articles/rag-context-leakage/" source: live --- Retrieved-context leakage in RAG is when a retrieval-augmented workflow injects confidential source text straight into the model prompt, so sensitive values leave your environment inside every request the retriever helped answer. The stakes here are not abstract. The OWASP Top 10 for LLM Applications ranks prompt injection first and sensitive information disclosure second, and disclosure is precisely what an ungoverned retrieval step produces: teams cannot put their real, useful documents in front of a model without losing control of what those documents contain. Researchers have shown this concretely: ConfusedPilot documents how a retrieval step can make an enterprise assistant surface a confidential document, even one already deleted, through the content it retrieves. A separate study by Zeng and colleagues demonstrated that RAG systems can leak the private documents held in their own retrieval database. Retrieval-augmented generation is where that tension shows up first, because RAG only works when it reaches into the material an enterprise most wants to protect. Where retrieved-context leakage actually happens in RAG A team builds a retrieval assistant over its internal knowledge base, and it works beautifully in the demo. Then security asks a plain question: when a user query pulls three chunks from the contract repository and the HR folder, where do those chunks go? The honest answer is that they get pasted into the prompt and sent to the model in full, every time. Most people treat retrieval as the safe part of the stack and reserve their worry for the model. That instinct is backwards. Retrieval is exactly where confidential text gets staged to leave the building. The leak does not live in the vector database, and it does not live in the embedding. It lives in the join, the moment the retriever stitches relevant passages into the context window next to the user's question. Whatever lands in that window is what the model receives. If the source documents hold salaries, customer names, or contract terms, the raw values travel with the request. Retrieval did its job well, and that is precisely the problem: the more relevant a chunk is, the more sensitive it usually is, because the answer lives in the specific clause, not the boilerplate around it. Why filtering at retrieval time breaks the assistant The common first fix is to filter at retrieval time and drop any document above a sensitivity label. This guts the assistant. The contract repository was the reason to build the thing in the first place, so excluding it leaves you with a system that answers everything except the questions people actually asked. Redaction after retrieval fails differently. If you strip the values out of a chunk before it reaches the model, you also strip the meaning: a clause with the numbers blacked out no longer supports a real answer, and the model either guesses or refuses. Plain masking treats the sensitive value as noise to remove, yet in a RAG workflow that value is often the signal. Closing the retrieve-to-inject boundary The better move is to treat the moment between retrieval and injection as a boundary rather than a passthrough. The retrieved chunks are real and relevant, so you keep them. What the model needs from them is their structure and the relationships between facts, not the literal sensitive values. So the chunks are substituted before they enter the prompt. Names, figures, and identifiers are replaced with consistent stand-ins while the surrounding language and the links between facts stay intact. The model then executes over a coherent passage; it simply never reasons over the raw value. When the model returns an answer keyed to those stand-ins, the answer is reconstructed against the original values inside your environment, so the user sees a real, complete response. This Substitute, Execute, Reconstruct pattern turns retrieval from a leak into a Restorable AI Data Boundary, and it is handled by a Context-Preserving Data Layer for AI, in CUBIG's case LLM Capsule, sitting on the retrieve-to-inject path. The retrieval quality does not drop. The window that used to carry confidential text now carries a structure-preserving version of it, and the confidential business context stays inside your perimeter instead of crossing it with the request. Approach on the RAG path Answer quality on sensitive docs Sensitive values leave the boundary Send chunks as-is Full Yes Filter sensitive docs out of retrieval Poor, the key documents are gone No Plain masking or redaction before inject Partial, meaning is lost No Substitute, execute, reconstruct Full No A quick self-diagnostic for your RAG workflow Run this test against your own pipeline before you assume the retrieval step is safe: Can you name every repository your retriever is allowed to reach, and would you be comfortable if a chunk from each one appeared verbatim in an external model log? When a chunk is injected, does the raw text of that chunk leave your environment, or a substituted version of it? If you removed your most sensitive sources from retrieval, would the assistant still answer the questions it was built for? When the model returns an answer, is it reconstructed against real values inside your perimeter, or does the real value only ever exist outside? Can you show an auditor exactly what crossed the boundary for a given query, without reconstructing it from memory? If any of those answers make you hesitate, the retrieve-to-inject boundary in your workflow is open. Where it fits in CUBIG's operating layer Retrieved-context leakage is one shape of a broader issue we cover in LLM data egress: sending data to a model is itself a boundary crossing, so the question is what form the data takes when it crosses. The agentic version of the same leak, where a tool result rather than a retrieved chunk carries the values out, is covered in agent tool-output leakage. Closing this boundary is what makes RAG usable under sensitive AI workflow enablement, and LLM Capsule runs on Syntitan, CUBIG's AI-Ready Data Platform. LLM Capsule closes this gap inside the retrieval step. Passages are substituted before they reach the model, so the context your pipeline sends carries structure, not raw confidential values. --- title: "Agent Tool-Output Leakage in Agentic AI Workflows" url: "https://cubig.ai/articles/agent-tool-output-leakage/" source: live --- Agent tool-output leakage is what happens when an agent calls a tool, receives real data such as a database row or an API response, and folds that result into its next reasoning step, sending confidential values to the model that no person ever typed into a prompt. Enterprises are betting on autonomy, and the bet is fragile. OWASP's AI Agent Security guidance names data exfiltration through tool calls, API requests, and agent outputs as a first-class risk, and it is the one agents feed quietly, because these systems touch data no one budgeted for governing. Microsoft's AI Red Team maps the same failure in agents: untrusted content that flows into an agent's memory can become a pivot point to exfiltrate data. An agent does not wait for a human to paste in the sensitive material. It reads the material itself, at runtime, from whatever tool the task requires, then carries it forward. That is the specific failure this article is about. What agent tool-output leakage looks like in practice Picture an agent handed a routine task: reconcile this week's failed payments and draft the customer emails. It calls a tool to query the billing database. Back come rows of real customer names, card fragments, amounts, and failure codes. The agent passes that result straight into its next reasoning step, which means straight into the model's context window. Nobody wrote a line of code that says "send customer records to the LLM." The agent did it on its own, because reasoning over tool results is exactly how agents work. Most of the attention in agentic systems goes to what the agent is allowed to do: which tools it can call, which actions it can take, which permissions it holds. The quieter risk is what the agent reads back and then carries into its own context. The action is visible in an audit log. The data riding along inside the reasoning trace usually is not. Why the tool-to-model hand-off is the leak point A single-shot prompt has one boundary crossing you can reason about: the text goes out, an answer comes back. An agentic AI workflow has many crossings, and they are dynamic. Every tool call can return real data, and the whole design of the agent is to take that return value and reason over it in the following step. Query a database, read a file, hit an internal API: each result becomes context for the model. The leak is not the tool call itself, and it is not the final answer. It is the moment the tool's output is folded back into the model's working memory. That is the crossing teams forget to govern, because it looks like an internal step rather than an egress. What makes this slippery is that the path is decided at runtime. You cannot enumerate in advance every record an agent will pull, because the agent chooses based on the task in front of it. The usual reflex, locking the agent out of the sensitive tools, produces the same dead end it produces everywhere: the billing database was the point of the task. An agent that cannot read it cannot reconcile the payments, so you have traded a leak for a system that no longer does the work. How this differs from a normal prompt leak The reason agent tool-output leakage slips past controls that catch ordinary prompt exposure is that the sequence is different at several points. A prompt-based check assumes a human author, a single moment, and a readable input. Agentic AI breaks all of those assumptions at once. Dimension Direct prompt leak Agent tool-output leakage Who introduces the data A person types or pastes it The agent reads it from a tool When it happens Once, at the prompt At every tool-to-model hand-off Where it shows up In the written prompt Inside the reasoning trace Path known in advance Yes No, decided at runtime Does blocking the tool fix it Not applicable No, it breaks the task Because no person types the data in, the leak never appears as a written prompt for a reviewer to flag. It appears as an internal reasoning step, which most monitoring treats as trusted. That is why a team can pass a prompt-hygiene review and still be exposed. The exposure is not theoretical: the AI Incident Database has cataloged more than 1,500 real-world AI failures. Closing the tool-to-model hand-off The durable approach treats each tool result as a boundary crossing on the way back into the model, the same way a careful team treats data going out. When a tool returns rows of customer records, those records are substituted before they reach the model's context: names, card fragments, and identifiers become consistent stand-ins, while the shape of the data stays intact. Which row owes what, which payment failed and why, all of that structure survives. The agent reasons over a faithful structure and never holds the raw values. When the agent produces its output, the draft emails or the reconciliation report, the result is reconstructed against the real records inside your environment, so the final artifact is complete and correct. The agent keeps its tools and its autonomy. The values simply stop riding along through every step. In this model the middleware does three things in order: it substitutes the tool output before the model sees it, it lets the agent execute over the faithful stand-ins, and it reconstructs the real values into the finished artifact locally. The component that sits on that hand-off is a Context-Preserving Data Layer for AI, in CUBIG's case LLM Capsule. It intercepts each tool result, applies the substitution, and rebuilds the output at the end. Because the stand-ins are consistent across steps, an agent can reference in step two the same entity it saw in step one and still reason correctly, without the underlying values ever entering the model. A quick self-diagnostic for your agents Run this check against any agentic AI workflow you already have in production or in a pilot. If you answer no to more than one, the tool-to-model hand-off is likely an open leak path. Can you name every tool your agent can call that returns confidential or regulated data? When a tool returns real records, do those values enter the model's context unaltered? Could you reconstruct, after the fact, which sensitive records passed through a given run? Does your current control depend on blocking a tool the task actually needs? If a reviewer read only the prompts, would they see the sensitive data the agent handled? Where it fits in CUBIG's operating layer Agent tool-output leakage is the agentic cousin of retrieved-context leakage in RAG workflows, where the carrier is a retrieved chunk rather than a tool result. Both are instances of LLM data egress, the broader point that sending data to a model is a perimeter crossing whether a human or an agent initiates it. Closing the hand-off is what lets multi-step agents run under sensitive AI workflow enablement, which runs on Syntitan, CUBIG's AI-Ready Data Platform. LLM Capsule holds the boundary at the tool call, not just the prompt. A tool result is substituted before it flows back to the model, so the agent finishes the task without raw values re-entering the context. --- title: "Workflow Closure: When AI Actually Completes the Work" url: "https://cubig.ai/articles/workflow-closure/" source: live --- Workflow closure is the point at which an AI step returns a finished business artifact rather than a draft, because its output has been reconstructed back into the original context, which is what turns a model response into work that is actually done rather than work someone still has to reassemble by hand. Most enterprise AI pilots stop one step short of workflow closure. The model produces something impressive in the demo, and then a person quietly spends twenty minutes putting the answer back into the document it came from. The demo closes the loop; the production workflow does not. Pilots that stall at this step rarely fail on model quality at all. They fail because the last mile, the reassembly, was left to people. The gap between a response and a result Consider what actually happens when a confidential document runs through a large language model. To get useful output without exposing the real values, the team substitutes or strips the sensitive parts first, then sends the model the structure of the work rather than the raw values. The model returns its answer against those stand-ins. Now a person has to map the placeholders back to the real entities, drop the text into the right cells, fix the formatting, and confirm nothing slipped. The model did the thinking; a human did the reassembly. That reassembly is where the value leaks back out, because it is slow, it is easy to get wrong, and it scales badly across hundreds of documents a week. A workflow that needs a person to finish every AI step is not really an AI workflow, it is a manual workflow with a model bolted into the middle. The tidy numbers from the pilot deck rarely survive contact with that reality. What closing the loop actually means Closure happens when the output comes back already mapped to the original context. The placeholders resolve to the real values, and the result lands in the structure it belongs to: the same table, the same clause numbering, the same layout the source had. What the reviewer opens is the finished document, not a kit of parts shipped with assembly instructions. Picture a renewal contract sent out for an obligations summary. An open workflow returns prose that refers to "Counterparty A" and "the agreed figure." A closed workflow returns the summary with the actual counterparty named and the real amount in place, sitting in the document where legal expects to read it. One is a draft; the other is done, and the difference between them is reconstruction back to the original context. Why closure depends on the data boundary Closure and confidentiality turn out to be the same problem solved well. The reason output usually comes back unfinished is that the real values were removed so the model could run at all. Restore those values locally, inside your own environment, and you get two results at once: the workflow closes, and the confidential data never had to cross the boundary during LLM data egress. This is why a substitution that preserves structure matters far more than one that simply blanks the data out. If the model worked on a coherent copy, with tables intact, relationships intact, and the confidential business context represented faithfully by stand-ins, then reconstruction is a clean mapping back. If the model worked on a redacted mess, there is nothing coherent to reconstruct, and closure stays impossible no matter how strong the model is. Plain masking optimizes for hiding values; enablement optimizes for finishing the work. Open loop vs closed loop, side by side The distinction is easiest to see when you put an open workflow and a closed one next to each other on the same task. Read down the right column and you have a working definition of AI workflow completion: the artifact is usable the moment the model returns it, and the sensitive values never left the room to make that happen. A closed loop is also a recordable one. Each pass through it can be logged automatically end to end, which is what Article 12 of the EU AI Act expects of high-risk AI systems: events recorded over the system's lifetime so the results can be traced. Closure is the exception, not the rule: an MIT study reported by Fortune found roughly 95% of enterprise generative AI pilots reach no measurable business impact, most stalling before the work ever finishes. Others scrap the work earlier still: S&P Global Market Intelligence found the average organization abandoned 46% of its AI proofs-of-concept before they reached production. A quick self-diagnostic Run your current AI workflow against this short test. If you answer "yes" to more than one of these, you have an open loop, not workflow closure: After the model responds, does a person still copy the output into the real document? Does the answer come back referring to placeholders like "Party A" instead of real names? Does someone reapply the original formatting, tables, or clause numbering by hand? Would doubling the document volume roughly double the human hours spent finishing? Do the real values get pasted back in a spreadsheet or editor outside any controlled step? Where workflow closure fits Workflow closure is the second half of sensitive AI workflow enablement. The first half sends the structure of the work across the boundary and lets the model run in place; closure brings the result home as a usable document. What the workflow delivers at the end is a reconstructed operational artifact, which is the real unit of AI delivery, not a chat transcript. The substitute, execute, and reconstruct mechanism is handled by a Context-Preserving Data Layer for AI, in CUBIG's case LLM Capsule, which runs on the CUBIG Syntitan platform. LLM Capsule is what makes closure the default rather than the exception. The work runs end to end on substituted data and the finished result is reconstructed inside your environment, so the workflow actually finishes. --- title: "From PII Masking to Workflow Enablement" url: "https://cubig.ai/articles/from-pii-masking-to-workflow-enablement/" source: live --- From PII masking to workflow enablement is the shift from hiding sensitive data so a model cannot misuse it, to substituting that data so a model can still complete the work, with the original values reconstructed inside your own systems afterward. Masking is built to make sure nothing sensitive is exposed. Enablement is built to make sure the work still runs. Those are not the same goal, and treating them as if they were is why many confidential-data AI projects stall. Regulated teams did not invent the removal habit; it is codified. HIPAA recognizes exactly two de-identification paths, Expert Determination and Safe Harbor, which removes 18 specified identifiers and requires no actual knowledge that the remainder could re-identify someone. NIST frames the underlying tension plainly: de-identification attempts to balance the contradictory goals of using and sharing data while protecting privacy, which is exactly the balance plain masking gives up. The alternative keeps utility: a peer-reviewed study by Vakili and colleagues found pseudonymized data can train and fine-tune models end to end without harming performance. When the data a regulated team is allowed to send a model has been stripped of the very structure the model needs, the project rarely fails loudly. It quietly produces output nobody can act on, and then it gets shelved. What PII masking was built to do PII masking comes from a world of databases and reports, not language models. The original job was to take a table full of names, account numbers, and birthdates and hand it to someone who should not see the real values: a test environment, an offshore analyst, a dashboard running aggregate queries. Black out the identifiers, swap in XXXX or a random placeholder, and the recipient can still run counts and joins. For that job, data masking works well. Nobody downstream needed the real name to compute a monthly total. A language model is a different kind of recipient. It does not run a fixed query against known columns. It reads a document the way a person would, and it leans on exactly the parts that masking throws away: who is referring to whom, in what order, and how one figure derives from another. Why masking starves a model Mask a contract and you do not get a slightly redacted contract. You get a page of holes. The clause that referenced "the counterparty" now points at [REDACTED], the payment schedule lost the figures that made it a schedule, and the model has to guess at relationships that used to be explicit. Ask it to summarize and it summarizes a document that no longer holds together. The output is safe and close to useless. The deeper problem is that plain masking treats every sensitive value as noise to be deleted. In a real workflow those values carry the structure: which party owes which amount, when each event happened, how a total was derived. Strip the values and you strip the structure with them, so the context the model needed is the first thing lost. For more on why that structure is itself the signal, see document layout preservation. What workflow enablement does instead Enablement keeps the structure and changes only what the structure is made of. Sensitive values are replaced with stand-ins, and the relationships between fields stay exactly as they were. The contract still reads as a contract: two parties, a schedule, a set of obligations that reference each other correctly. The model sees a coherent task instead of a redacted blank, does the work, and the result is reconstructed with the real values back in place inside your environment. The mechanism, step by step, is covered in substitute, execute, reconstruct. The idea is to send the model the work's structure, not the raw values, let it run in place, and reconstruct the business meaning locally when it returns. This is delivered by a Context-Preserving Data Layer for AI, in CUBIG's case LLM Capsule, which sits between your real data and the model. It keeps the structure intact on the way out and rebuilds the confidential business context on the way back. The contrast with plain masking is not a matter of degree: masking asks how much it can remove, while enablement asks how much work it can let through. Masking versus enablement, side by side The two approaches diverge on almost every dimension that matters once a language model is the consumer of the data. This table lays out where they differ. Dimension PII masking Workflow enablement Primary goal Nothing sensitive is exposed The work still runs correctly Treats values as Noise to remove Structure to preserve Relationships between fields Broken Kept intact Model output quality Degraded, often unusable Representative until you reproduce it on your own data Real values restored No Yes, locally in your environment Built for Reports, test data, analysts AI and agent workflows on confidential data How to tell which one you are actually running If you are not sure whether your current pipeline is masking or enabling, run this quick self-check on a real document from your workflow: After processing, does the document still read as a coherent whole, or is it a set of disconnected blanks? Can the model still tell which party refers to which, and how one figure derives from another? Does the finished result come back with the real values restored inside your environment, or does someone re-enter them by hand? When output quality drops, can you trace it to lost structure rather than a weaker model? Do the confidential values ever have to leave your systems for the work to complete? If the answers point to disconnected blanks, manual re-entry, and structure loss, you are masking. If the document stays coherent and the real values return locally, you are enabling the workflow. Where it fits This shift is the conceptual core of sensitive AI workflow enablement, the practice of running AI on confidential data without that data leaving your environment. Once you stop asking what to hide and start asking what to enable, the whole approach to confidential data changes. The local mapping store that maps stand-ins back to real values, described in the local mapping store, is what makes that reconstruction possible inside your own boundary. All of it runs on the CUBIG Syntitan platform. LLM Capsule is what the shift from masking to enablement looks like in practice. It substitutes instead of blanking, so the model keeps the structure it needs while your real values stay inside. --- title: "Local Mapping Store: Keeping the Mapping Inside the Boundary" url: "https://cubig.ai/articles/local-mapping-store/" source: live --- A local mapping store is the component that keeps the mapping between your real values and their structure-preserving stand-ins inside your own environment, so an AI workflow can run on substituted data while the key to restore it never crosses the enterprise boundary. When AI runs on confidential data by substitution, one detail decides whether the approach is sound or theater: where the mapping lives. Send the model stand-ins for the sensitive values, and you still hold a table that links each stand-in back to the real thing. If that table travels, the substitution was pointless. The local mapping store is where it stays put. The obligation is already written down: GDPR Article 32 requires technical and organizational measures appropriate to the risk, and names pseudonymization and encryption explicitly, and where the mapping lives decides whether the pseudonymization is real. The EU cybersecurity agency ENISA sets out the pseudonymisation techniques and best practices this relies on, where the protection is only as strong as the control kept over the mapping. Standards exist for exactly this substitution: NIST's format-preserving encryption (SP 800-38G) transforms a value while keeping its original shape, so a stand-in still fits the field it came from. What a local mapping store holds During substitution, each sensitive value is swapped for a stand-in that keeps its shape: a company name becomes a consistent company-shaped stand-in, a figure becomes a figure that holds its place in the table. Something has to remember that "Stand-in 4471" was a specific counterparty and that a placeholder amount maps to a real one. That record is the mapping, and the local mapping store is where it is stored and resolved. The vault is deliberately narrow. It is not a general data store and not a place the model reaches into. It holds the correspondence between originals and stand-ins, nothing more, and it answers exactly one kind of request: given these stand-ins in a returned result, resolve them back to the real values, here, inside the environment that owns them. Why the mapping has to stay local The whole point of substitution is that the work leaves and the values do not. The mapping is the values, in a different form. A stand-in table you can resolve is functionally the sensitive data, so if it sits in an external service, you have simply moved the exposure one hop and called it safe. Keeping the vault local closes that hop. The model receives structure it can reason over and never receives the key that would turn structure back into identities. Substitution on the way out and reconstruction on the way back both read from a vault that lives where your data already lives, whether that is a cloud tenant, an on-prem cluster, or an air-gapped network. Design question Mapping held externally Local mapping store Does the key to real values leave? Yes No Model reads a coherent task? Yes Yes Result reconstructs in your context? Depends on the service Yes, locally Works in an air-gapped network? No Yes The vault and the enterprise boundary The enterprise boundary is the line your data is not supposed to cross without a reason. Substitution lets useful work cross it while the values stay behind, and the local mapping store is what makes that split honest rather than aspirational. Structure flows outward to the model; the mapping and the originals flow nowhere. Reconstruction happens on the inside, so the boundary is not a wall the workflow bounces off but a membrane that lets the task through and keeps the identities home. This is what turns the pattern into a restorable AI data boundary. The result comes back inside and resolves against the real context, rather than leaving for good in a form someone has to translate by hand. Without a local mapping store, there is nothing to resolve against, and the returned answer stays keyed to stand-ins nobody recognizes. Does your AI workflow keep the key at home? Run your current confidential AI setup through these questions: When you substitute sensitive values, where does the mapping physically live? Could that mapping be resolved by anyone outside your environment? Does reconstruction happen inside your systems, or in a third-party service? Could the same workflow run unchanged in an air-gapped network? If the mapping leaked, would the substitution still mean anything? If the mapping lives anywhere but inside your boundary, the substitution is only as local as its weakest hop. Where the local mapping store fits The local mapping store is one piece of the mechanism behind sensitive AI workflow enablement. It holds the mapping created during substitute, execute, reconstruct, so the substitution step has somewhere safe to record its stand-ins and the reconstruction step has something to resolve against. In production these steps are delivered by a Context-Preserving Data Layer for AI, in CUBIG's case LLM Capsule, which keeps the vault inside your environment and runs on the CUBIG Syntitan platform. In LLM Capsule the vault is not a diagram, it is where your real values live. The stand-ins go to the model while the originals never leave, and the mapping restores them locally when the result returns. --- title: "Restorable AI Data Boundary: Two-Way, Not a Blocking Gate" url: "https://cubig.ai/articles/restorable-ai-data-boundary/" source: live --- A restorable AI data boundary is an execution architecture where only the structure of a task crosses out to the model and the result is rebuilt back inside your environment, so the boundary carries work in both directions instead of only stopping data on the way out. The stakes are not academic. Latanya Sweeney's foundational work found 87% of Americans are uniquely identifiable from just ZIP code, sex, and date of birth, and Narayanan and Shmatikov later de-anonymized a supposedly anonymous dataset from only a few outside facts; in regulated firms the response is mundane: the sensitive data the workflow needs is never allowed to reach the model. The workflow does not fail on model quality. It fails at the edge of the network, where a one-way control says no and the work stops there. The word that carries the weight in a restorable AI data boundary is restorable. Most boundaries an enterprise puts around its data exist to stop things. A restorable boundary is built to let work through and bring it home, which is a different shape of thing, so it is worth being precise about why the two behave so differently. The boundary most teams already have is one-way Consider how data-loss prevention behaves. It watches the edge of the network, and the moment something sensitive tries to leave, it blocks the request. The logic is a gate: sensitive value detected, request denied. That is a one-way boundary, and for its own purpose it works well enough. What it cannot do is get the work done. When an analyst wants to run a confidential contract through a model, a one-way boundary has exactly one answer, which is no. It cannot let the task out and bring a result back, because it was never designed to bring anything back; it only knows how to stop. So the AI project stalls at the perimeter, and the work either dies quietly or leaks out through a side channel nobody is watching. That gap is the whole problem. A blocking control protects the data by making the workflow impossible, and the question a restorable boundary asks runs the other way: how do you let the workflow run while the sensitive values stay inside the entire time. What crosses the boundary, and what comes back A restorable AI data boundary splits the request into two things that a blocking control treats as one. There is the confidential value, and there is the structure of the work. The value stays in. The structure goes out. On the way out, sensitive content is substituted with structure-preserving stand-ins, so what crosses the boundary is a coherent task with the identifying values swapped. The model receives a real problem to solve; it does not receive your data. The model runs, its output crosses back in, and here is the move that makes the boundary restorable: the stand-ins are resolved to the real values, in their original context, inside your environment. The output reconstructs into your actual document. So the boundary is two-way by design. Structure out, result back, and the original values never cross in either direction. Set that against the one-way gate, which owns a single move and spends it saying no. The restorable boundary owns two moves, substitute on the way out and reconstruct on the way back, and it spends them handing you finished work. Restorable versus blocking, side by side The contrast comes down to what each boundary is measured by. A blocking boundary is judged by how much it stops, so more blocks read as more wins. A restorable boundary is judged by how much previously off-limits work it lets you complete. Those are not the same metric, and they do not even point the same direction. Property Blocking boundary (plain masking, DLP-style) Restorable AI data boundary Direction One-way: outbound only Two-way: structure out, result back Answer to a sensitive workflow No Yes, here is the finished work What crosses out Nothing, or a redacted stub A coherent task with substituted values Reconstruction of the result None Rebuilt against original values, locally Success measured by Volume blocked Workflows completed This is why a restorable boundary is not a security control with extra features bolted on. It is an enablement architecture. Security and privacy reviewers do gain something real from it, because the original values stay inside and never cross out, which gives them a basis to approve the workflow. GDPR Article 32 speaks the same language, requiring technical and organizational measures appropriate to the risk, with pseudonymization named explicitly and regular testing of those measures. That approval is a by-product of the design rather than its purpose. The purpose is to make the work run. Why reconstruction has to happen inside Reconstruction is the half of the boundary that blocking controls simply do not have, and it only holds up if it happens inside your environment. The mapping between each original value and its stand-in is what turns the model's output back into a real document. If that mapping lived outside your walls, you would have relocated the exposure rather than removed it. So the boundary keeps the mapping local, and reconstruction runs where your data already lives. The result is rebuilt in place, and nothing has to travel out to be restored. That single design choice, keeping the mapping inside, is what separates a restorable boundary from a masking tool that merely hides values and hopes the workflow can proceed on the redacted version. The redacted version usually cannot proceed, which is exactly why one-way controls stall so many projects. How to tell whether your boundary is restorable Run this quick self-check against whatever sits between your teams and their models today. If you answer no to more than one, you have a blocking boundary rather than a restorable one. When a workflow needs sensitive data, does the boundary return a completed result, or only a denial? Does a coherent task cross out to the model, rather than a redacted stub the model cannot actually use? Are the original values reconstructed into the final output automatically, or does someone stitch them back by hand? Does the mapping between values and stand-ins stay inside your environment at all times? Is the boundary measured by workflows completed, not just by volume blocked? Where it fits A restorable AI data boundary is the architecture behind sensitive AI workflow enablement. The mechanism that crosses it in each direction is set out in substitute, execute, reconstruct, where reconstruction is the step that makes the boundary two-way. The mapping that makes reconstruction possible stays in a local mapping store inside the enterprise boundary, and the shift in thinking that gets teams here is covered in from PII masking to workflow enablement. This architecture is delivered by a Context-Preserving Data Layer for AI, in CUBIG's case LLM Capsule, and runs on the CUBIG Syntitan platform. LLM Capsule is how CUBIG turns this boundary into something work can cross. The structure goes out to the model, the real values stay home, and the finished result is rebuilt inside your environment. --- title: "Substitute, Execute, Reconstruct: The Three Steps to Run AI on Confidential Data" url: "https://cubig.ai/articles/substitute-execute-reconstruct/" source: live --- Substitute, execute, reconstruct is the three-step method that lets AI run on confidential enterprise data: sensitive values are substituted for structure-preserving stand-ins before the work leaves your systems, the model executes on that substituted version, and the result is reconstructed back into your original context locally. Enterprises are not short on AI ambition; they are short on data they can safely put in front of a model. When the work involves contracts, patient records, or operational documents, the usual answer is to keep the data away from the model entirely, which also keeps the value away. Substitute, execute, reconstruct is the method that breaks that trade-off: it sends the model the work's structure, not the raw values, and rebuilds the answer where your data already lives. What does substitute, execute, reconstruct mean? Most explanations stop at the headline: the model never sees your real data. That is true, but it skips what each step actually does. Every step has a distinct job, and each one decides whether the workflow comes back finished or comes back as a chore someone has to untangle by hand. The three steps map to three plain questions. What did you send instead of the real values? What did the model do with it? And how did the answer get back into your world? Read in that order, the method stops being jargon and starts being a checklist you can run against any confidential AI workflow. Step 1: Substitute the values, keep the structure This is where the approach either holds or falls apart, and the difference comes down to one thing: structure. Plain masking removes the sensitive value and leaves a hole. A name becomes [REDACTED]. A price becomes XXX. A clause naming a specific counterparty becomes a gap the model guesses around. The document is now safe and also incoherent, because the model is reading a page full of blanks and cannot reason about relationships that no longer appear on it. Structure-preserving substitution. It replaces each sensitive value with a stand-in that keeps the shape of the original. A company name becomes a different but consistent company-shaped stand-in. A figure becomes a figure that holds its place in the table and its relationship to the other figures. Contract indentation, the row-and-column logic of a spreadsheet, the ordering of steps in an operational note: all of it survives. What leaves is the identifying content. What stays is everything the model needs to do the task. So the question that matters here is what you kept, not just what you hid. A good substitution keeps the work intact and lets the values go, which is the line between a redacted blank and a coherent task the model can actually complete. Step 2: Execute on the substituted version Now the substituted version goes to the model, and this is the step that changes the least. You run the model you already chose: an external frontier LLM, an on-prem open-weight model, a retrieval pipeline, or an agent that was built around your workflow. It executes on the substituted data much as it would on the original, because from its point of view the task still reads as complete and consistent. The reasoning a model does over a substituted contract tracks the reasoning it would do over the real one, since the relationships it reads from are still in place. Meanwhile the original values are absent from this step. They never left the environment where they live. What went out was the work, not the record, and that separation is the whole point: you get the model's full capability without putting confidential values in front of it. The principle behind this is well established: differential privacy, as Penn's Aaron Roth explains, lets a system surface representative patterns while preventing anyone from revealing information about a specific individual. Step 3: Reconstruct the answer in your context The model returns a result, but that result is written against the substituted version. It refers to the stand-in names, the stand-in figures, the placeholder clauses. On its own it is not yet something your team can use. Reconstruction maps it back. The stand-ins resolve to the real values in the real context, so the output arrives as a finished business document instead of a draft keyed to stand-ins nobody recognizes. A reviewed contract reads against the actual counterparty and the actual numbers. A summarized report names the real metrics. The reviewer opens it and sees their own document, completed, rather than a puzzle to reassemble. This step is what makes the method worth running. Without it you would hold a safe result that still needs manual translation, which is most of the effort you were trying to skip. With it, you reach Workflow Closure: the work is actually done, in place, and ready to use. How the three steps run in place A fourth idea wraps the other three, and it is easy to miss because nothing visibly moves. The whole sequence runs inside the environment you already operate, whether that is a cloud tenant, an on-prem cluster, or an air-gapped network. The substitution mapping and the reconstruction both happen locally, held in a Local Mapping Store, so your infrastructure does not have to be rebuilt to accommodate the method. Regulation points the same way: GDPR Article 32 requires technical and organizational measures appropriate to the risk, and it names pseudonymization and encryption explicitly. Substituting real values with structure-preserving stand-ins is the same trade the Royal Society and Alan Turing Institute describe for synthetic data: privacy and fidelity move together, so the aim is a faithful stand-in rather than a blank. Running in place matters because the alternative, routing data through some new external service to get the benefit, reintroduces exactly the LLM Data Egress you were avoiding. Keeping the flow local is what makes the first three steps honest. The work goes out as structure; the values and the mapping stay home. Together, steps three and four describe a Restorable AI Data Boundary, where the answer comes back inside rather than leaving for good. Substitute, execute, reconstruct versus plain masking The clearest way to see the difference is to line the method up against the redaction approach most teams reach for first. Question Plain masking Substitute, execute, reconstruct Keeps document structure? No Yes Model reads a coherent task? Partial Yes Answer returns ready to use? No Yes Real values leave your systems? No No Manual cleanup afterward? High Low Masking and substitution both keep raw values off the wire. The gap opens after that: masking hands the model a blanked-out page and hands your team a translation job, while substitution hands the model a complete task and hands your team a finished artifact. A quick self-diagnostic Run your current confidential AI workflow through these questions. If you answer no to more than one, the method is worth a closer look. When you strip sensitive values, does the document still read as a coherent task, or does it become a page of blanks? Do the real values ever cross the boundary out to the model, even briefly in a log or a cache? Does the model's answer come back ready to use, or does someone re-key it against the originals by hand? Does the substitution mapping stay inside your environment, held somewhere you control? Could you run the same flow unchanged in an air-gapped network if a regulator asked? Where it fits in CUBIG's operating layer Substitute, execute, reconstruct is the engine under sensitive AI workflow enablement. The pillar describes what the practice is for; this article is how it actually runs. In production, the four steps are delivered by a Context-Preserving Data Layer for AI, in CUBIG's case LLM Capsule, which handles the substitution on the way out and the reconstruction on the way back. It runs on the CUBIG Syntitan platform, so the same boundary applies whether the workflow reads a contract, a spreadsheet, or a scanned document across a Multimodal AI Data Boundary. The mechanism is the product; the middleware is where it lives. LLM Capsule is the layer that runs this loop for you. It substitutes the sensitive values, lets your model execute on the structure, and reconstructs the real result inside your environment. --- title: "OT Data Reproducibility: Making Manufacturing AI Repeatable" url: "https://cubig.ai/articles/ai-in-manufacturing-reproducibility/" source: live --- OT data reproducibility is the ability to rebuild the exact operational-technology data state, sensor calibrations, sampling rates, line configuration, and tag mappings, that produced a manufacturing AI result, so the same process outcome can be examined and reproduced later. On the factory floor, the data a model runs on rarely sits still, and that instability is expensive. Operational technology is distinct enough that NIST maintains a dedicated guide to securing it, SP 800-82 Rev. 3, written around OT's unique performance, reliability, and safety constraints. Manufacturing lines are among the hardest cases: sensors get recalibrated, sampling windows shift, and tag maps change after a line reconfiguration, so "the same model" quietly runs on different inputs from one week to the next. Why OT data breaks reproducibility on the plant floor A manufacturer runs a model for quality prediction or process optimization on operational-technology data. It performs well for a while, then degrades. The model binary is unchanged, yet the results drift. A temperature sensor was recalibrated, a sampling window was shortened during a maintenance pass, or a tag was remapped after an equipment swap. When the team tries to reproduce last month's good result to understand the regression, they cannot. The OT data state that produced it is gone, overwritten by the live stream. This is the industrial form of the same execution-drift and schema-change patterns that undermine model reproducibility everywhere, and it is why preprocessing drift is so hard to catch on the floor: the changes are physical and easy to miss. OT volatility is dangerous precisely because it does not raise obvious errors. A recalibration shifts a sensor's baseline by a fraction. A sampling-rate change alters the signal's shape without breaking the pipeline. A tag remap points the model at a different measurement while every dashboard stays green. None of this looks like a failure until the output quietly gets worse. The pattern is endemic well beyond manufacturing: in a CHI study of high-stakes AI, 92% of practitioners interviewed had experienced data cascades, compounding downstream problems triggered by upstream data issues that nobody treated as AI work. What makes a manufacturing AI result reproducible A model-centric view answers only which model ran. For production AI on OT data the useful question is which data state and execution conditions produced the result. When you can name and rebuild that state, a regression stops being a mystery and becomes something you can investigate. Reproducibility on the floor rests on three things: capturing the conditions behind every run, scoring the stability of those conditions so drift is visible early, and being able to restore a past state to compare against a current one. Miss any one of them and you are back to guessing which physical change moved the output. How Release State and Run Binding restore an OT data state Binding each run to a Release State makes process results reproducible without exporting raw OT streams off-site. The plant keeps its data; the operating layer records the state and the conditions around it. Run Binding captures the calibration, sampling configuration, line setup, and tag mapping behind each result, so a run is tied to the exact inputs it saw. Reproducibility, one of the six readiness axes, keeps the stability of these conditions in view as a score you can watch trend, rather than a surprise you discover after quality drops. Diff pinpoints whether a recalibration or a remap explains a regression between two runs, instead of leaving the team to hunt through change logs. Reproduce restores the OT data state that produced the good run, so it can be examined next to the degraded one. For production AI the question is not only which model ran, but which data state and execution conditions produced the result. Syntitan, CUBIG's AI-Ready Data Platform, scores enterprise data on six axes: Usability, Integrity, Context, Consistency, Reproducibility, and Traceability; it rebuilds what blocks execution, and it binds every AI or agent run to a data state you can diff and reproduce. Release State compared with a plain data snapshot A snapshot copies values at a point in time. A Release State records the executable conditions around a run, which is what OT reproducibility actually needs. The difference shows up the moment you try to explain a regression. Capability Plain OT snapshot Release State + Run Binding Stores raw sensor values Yes Yes Captures calibration and sampling config No Yes Records tag mapping at run time Partial Yes Diff two runs to isolate a change No Yes Restore the state that produced a result No Yes Keeps raw OT streams inside the plant Yes Yes A quick self-diagnostic for your OT workflow Run this test against one model you already trust on the line. If you cannot answer most of these in minutes, your process results are not yet reproducible. Can you name the exact calibration and sampling configuration behind a result from six weeks ago? When a sensor is recalibrated, does that event get bound to the runs it affects? Can you diff two runs and point to the single physical change that moved the output? Can you restore last month's OT data state and rerun against it? Do you track a stability score for these conditions, or only notice drift after quality drops? Where OT data reproducibility fits in CUBIG's operating layer This is the operating layer for AI-ready data, not a security control and not a monitoring dashboard. It sits under the model and treats the data state as a first-class, reproducible object. On the floor that means a recalibration or a remap becomes a recorded, diffable event rather than an untracked physical change that silently degrades output. The payoff is concrete. Instead of "the model got worse on line 3," the team can say "the temperature sensor was recalibrated on the 12th, here is the diff against the prior run, and here is that earlier state restored for comparison." The regression turns into a process you can investigate and close, which is what separates a durable manufacturing AI program from one that erodes as the plant changes. Erosion is the default more often than teams expect: a Scientific Reports study that aged four standard model types across 32 industry datasets found temporal degradation in 91% of the combinations tested. To see how this connects to the broader foundation, start with what AI-ready data means, then read how schema changes break production AI and how AI fails after deployment. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "Public Sector: Running AI on Citizen Service Workflows" url: "https://cubig.ai/articles/public-sector-workflows/" source: live --- Running AI on citizen service workflows means sending an external model the structure of a case, not the resident's real records, so a public agency can draft decisions and clear casework backlogs while every citizen identifier stays inside its own environment. The gap between what AI could do for public agencies and what they are allowed to try is wide, and it is widening. What good looks like is already written down: GAO built its accountability framework for federal AI on four principles, governance, data, performance, and monitoring, and the data principle is where agencies get stuck. In government, the data that would make a project work is usually the data an agency is legally bound to keep in place. The result is a backlog nobody can automate and a technology nobody can point at. The appetite is already there: an Alan Turing Institute study found that 45% of public servants were aware of generative AI in their work and 22% actively used it, well ahead of the rules meant to govern it. Why citizen casework resists automation Picture a housing benefit appeal. The case file that lands on a caseworker's desk runs to forty pages: the original application, income statements, a prior decision, correspondence, and supporting documents that name the applicant's national ID, address, household members, and medical circumstances. Someone has to read all of it, weigh it against policy, and write a reasoned decision. There are thousands of these, and the backlog is measured in months. A model could draft the decision and surface the governing policy in minutes. But the file is citizen data held under a legal duty, and it does not leave the agency. This is the public-sector shape of a problem every regulated organization meets: the work that would benefit most from AI sits on exactly the records that are least free to move. The two responses that fail Agencies usually reach for one of two answers, and both cost them the outcome. The first is to anonymize the case file before any model sees it. A benefit decision, though, turns on specifics: this household, this income, this prior ruling, these circumstances. Strip them out and the model reasons about a generic case that resembles no real applicant, so its draft is useless for the actual decision in front of the caseworker. The second answer is to forbid the model outright and keep every file inside the agency, which is where most public bodies sit today. The caution is widespread: a 2023 BlackBerry survey found 75% of organizations were implementing or considering bans on generative AI apps at work. The backlog stays, the caseworkers stay overloaded, and the technology that could help waits outside the door. Both responses rest on the same mistaken assumption: that the citizen's identifying details and the substance of the case are one and the same thing. They are not, and separating them is what makes government casework automation possible. Enablement applied to a citizen case file There is a third path. Send the model the structure of the case instead of the file itself. The income figures keep their relationships, the timeline of applications and decisions stays in order, and the policy-relevant facts stay intact. What gets substituted is the identifying layer: the national ID, the name and address, the household members, all replaced with consistent stand-ins before anything reaches the model. The model reads a coherent case and drafts a reasoned decision against it. That draft is then reconstructed inside the agency's own systems, where the real applicant's identity returns and the output becomes a document a caseworker can review and issue. The casework ran, the citizen data never left the agency's environment, and the caseworker gets a draft grounded in the actual facts rather than a generic template. This substitute, execute, reconstruct pattern is the practical form of sensitive AI workflow enablement applied to resident records, and it is a different move from the plain masking most agencies already know: masking blanks a value, while this preserves the working structure the model needs. What runs and what stays put A Context-Preserving Data Layer for AI, in CUBIG's case LLM Capsule, holds the real values inside a Local Mapping Store and rebuilds the finished decision once the model has done its part. The Restorable AI Data Boundary is the line the substitution never crosses: structure goes out, identity stays home, and the reconstruction closes the loop so the workflow actually finishes rather than stopping at a redacted draft nobody can act on. Element of the case Sent to the model Restored locally National ID, name, address No Yes Household members No, substituted Yes Income figures and relationships Yes, structure preserved Kept Application and decision timeline Yes Kept Policy-relevant facts Yes Kept The reader test is simple. Run through this list against any citizen service workflow you want to open up to AI, and if you cannot answer yes to the first four, the workflow is not ready to run. Can you name which fields are true identifiers and which are the substance of the case? Does the substitution keep figures and timelines in their real relationships? Does the finished output get reconstructed inside your own environment, not the vendor's? Can a caseworker act on the draft as if it were written on the real file? Does your data protection officer have a concrete basis to sign off? Where AI on citizen service workflows fits This is one public-sector view of a pattern that recurs wherever confidential records meet a regulated mandate. In telecom, the protected file is subscriber and network topology data rather than a case file; in healthcare, it is protected health information in clinical workflows. The mechanism is the same in each. All of it runs on the CUBIG Syntitan platform, CUBIG's AI-Ready Data Platform. Because no citizen identifiers cross the boundary, the workflow lines up with the data-protection duties that govern public records, and the agency also gains a reviewable trail for how each AI-generated decision was produced. The risk being managed is a named one: NIST's Generative AI Profile counts data privacy and information security among the twelve risks it treats as unique to generative AI or amplified by it. The data protection officer's approval is the by-product. The point was always to clear the backlog without putting a single resident's file at risk, and to move public-sector AI from a slide deck into the casework queue. LLM Capsule puts this pattern to work on real case files. It substitutes the citizen identifiers before anything reaches the model and rebuilds the decision inside the agency, so the casework runs while the records stay put. --- title: "Backtest Reproducibility Under Audit in Finance" url: "https://cubig.ai/articles/financial-backtest-reproducibility-under-audit/" source: live --- Backtest reproducibility means being able to restore the exact data state a model backtest ran on, so an auditor can confirm the reported performance came from the data you claimed, not from a version that has since moved on. Representative example. Figures are illustrative until reproduced on your own model and data. The stakes here are not academic. The Federal Reserve's SR 11-7 guidance on model risk management expects model outputs to be validated, documented, and monitored on an ongoing basis, and that expectation assumes you can show what a model actually ran on. In regulated finance the fastest way to get an AI result thrown out is to fail the one question every auditor asks: show me that this run came from the data you say it did. A backtest you cannot reproduce is a backtest you cannot defend. Why a backtest fails an audit A risk team backtests a model and reports strong performance to a committee. The committee acts on it. Months later an auditor arrives with a simple request: reproduce the run that produced these numbers. The model version is sitting in the registry, so that part checks out. But the data behind the backtest has kept moving. The window shifted, corporate-actions adjustments were reapplied, reference rates were restated, and the point-in-time view that existed on the day of the run is gone. The team can rerun the model, just not on the same state. The numbers come back different, and now the original result is the thing under question, not the audit. The failure is not in the model, and it is not a lapse in anyone's diligence. It is that the data state was treated as transient. Nobody wrote down the exact conditions the result depended on, because at the time the result looked self-evident. The weakness is not unique to finance: when Nature surveyed 1,576 researchers, more than 70% had tried and failed to reproduce another scientist's experiment. Why backtests are especially exposed Backtests live or die on point-in-time correctness. Use today's adjusted prices instead of the prices as they stood on the trade date, and you get a different answer, often a flattering one, because the future has already been folded into the inputs. Restated fundamentals do the same thing more quietly. Survivorship in a reference universe does it again. Each of these is a change in data state, not a change in code. Model versioning captures the code and the weights; it says nothing about which slice of a moving market the model was scored against. So when the auditor says "reproduce it," a team that only versioned the model is left rebuilding the data from memory, and memory is exactly what an audit is designed to distrust. Supervisors have already written that distrust into standing rules: BCBS 239 requires banks to aggregate risk data accurately and completely on a largely automated basis, so as to minimize the probability of errors. This is the gap between AI-ready data and merely clean data: the data can be pristine and still be undefendable if the state it was in during the run was never captured. For a fuller treatment of that distinction, see our piece on AI-ready data versus clean data. What operating control changes When each backtest run is bound to a Release State, the audit question stops being a threat and becomes a lookup. A Release State is the recorded, restorable condition of the data at the moment of the run: the window, the adjustments applied, the reference data in force, and the access permissions that shaped what the model could see. Three operations do the work: Operation What it answers for the auditor Artifact produced RUN BINDING Which exact data state produced this reported number? A backtest run tied to one Release State DIFF What changed between that state and today, and does it explain the gap? A named, inspectable set of differences REPRODUCE Can the committee's numbers be examined under the conditions that made them? The original state restored for inspection Run Binding identifies the state. Diff explains what changed. Reproduce restores the conditions for inspection. Run Binding ties the reported result to the exact state it used. Diff shows what moved between the original state and any later one, so a discrepancy has a named cause instead of a shrug. Reproduce restores the original state so the committee's numbers can be examined under the conditions that produced them. For production AI the question is not only which model ran; it is which data state and execution conditions produced the result. Syntitan, CUBIG's AI-Ready Data Platform, scores enterprise data on six axes, rebuilds what blocks execution, and binds every AI or agent run to a data state you can diff and reproduce. The six axes are Usability, Integrity, Context, Consistency, Reproducibility, and Traceability, and a defensible backtest leans hardest on the last two. Is your backtest audit-ready? A quick self-test Run your last committee-reported backtest through these five checks. If you answer "no" or "not sure" to more than one, the result is more fragile than it looks: Can you restore the exact point-in-time data the backtest scored against, not a rebuilt approximation? Is that run bound to a single named data state rather than "the production tables as of roughly then"? Can you produce a diff between that state and today's data, with each change named? Do the corporate-actions adjustments and reference rates in the run carry their own as-of dates? If two people reproduce the run independently, do they start from the same restored state? Where it fits in CUBIG's operating layer The defensible version of a backtest is not "trust our numbers." It is "here is the state the run used, here is what has changed since, and here is that state restored for you to inspect." Operating control does not promise byte-identical output, and it should not; performance figures stay representative until you reproduce them on your own model and data. What it removes is the black box the auditor would otherwise have to take on faith. This is the operating layer for AI-ready data at work: a reproducible AI-ready state that survives contact with a regulator. The same binding that defends a backtest also lets you reproduce a production incident later, which we walk through in how to reproduce an AI incident, and it sits alongside formal governance rather than duplicating it, as we cover in operating control versus AI governance frameworks. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "Operating Control vs AI Governance Frameworks: Policy and Proof" url: "https://cubig.ai/articles/operating-control-vs-ai-governance-frameworks/" source: live --- AI governance frameworks define what an organization is allowed and required to do with AI; operating control is the run-level layer that produces the executable evidence proving a specific result met those rules. Policy and proof sit on different layers, and an audit needs both. This is the comparison most often miscast as overlap, because both a framework and an operating layer touch words like "governance," "traceability," and "trust." The policy side has never been better codified: ISO/IEC 42001, the first management system standard written for AI, spans lifecycle management, risk and impact assessment, and oversight of suppliers. A framework can name the requirement. It cannot, by itself, satisfy it on the data that a given run actually used. That gap is where most governance programs quietly stall. What AI governance frameworks do well Frameworks set the rules. They include internal AI policies, risk taxonomies, and external standards such as the NIST AI Risk Management Framework or the obligations written into the EU AI Act. Forty-seven governments have adhered to the OECD AI Principles, which name transparency, accountability, and robustness, security and safety among their commitments. A good framework tells the organization which use cases are permitted, what documentation each system must carry, and what reproducibility, oversight, and traceability have to exist before a model goes to production. That alignment is genuinely valuable. It gives legal, risk, and engineering a shared definition of responsible AI, and it turns "be careful with AI" into named obligations that a review board can check against. The NIST framework, for instance, calls for AI systems to be documented, traceable, and auditable. Nobody serious argues the policy layer is optional. Where the gap opens The trouble starts at the word "traceable." Article 12 of the EU AI Act goes as far as text can, obliging high-risk systems to log events automatically across their lifetime so outcomes remain traceable. A framework can restate that duty, yet neither the statute nor the framework holds a mechanism to make any single result reproducible. Between the written policy and an actual audit sits a concrete, unglamorous question: for this specific output, can you show which data state produced it, what changed since the last known-good run, and can you restore that state to inspect it? If the honest answer is "we'd reconstruct it from scattered logs, notebooks, and a snapshot someone hopefully kept," the policy is aspirational. The requirement exists on paper; the evidence does not exist at run time. Auditors and regulators increasingly ask for the second thing, not the first, and a binder full of principles does not survive that request. Operating control: policy turned into evidence Operating control is the layer that answers the run-level question. It binds each AI or agent run to a Verifiable Data State, keeps an operating record of which data produced which result, and lets you diff two states and reproduce either one. Where a framework says "results must be reproducible," operating control is the mechanism that makes a named result reproducible on demand. Concretely, that means three owned capabilities. A Release State captures the exact data conditions a run depended on. Run Binding ties the run to that state so the link is not a guess later. Diff and Reproduce let a reviewer replay the state, compare it against another, and see precisely what moved. The framework's Traceability requirement stops being a promise and becomes something you can hand over. Are AI governance frameworks and operating control competing? No, and treating them as rivals is the mistake. They are complementary layers of the same program. The framework decides the rule; operating control supplies the proof for each run that the rule was met. One without the other fails in a predictable way: policy without evidence is unenforceable, and evidence without policy has nothing to prove against. The clearest way to see the split is to line up who owns each responsibility. Responsibility AI GovernanceFrameworks Operating Control Define permitted use cases and required controls YES NO State that results must be reproducible and traceable YES PARTIAL Enforces the stated rule Make a specific result reproducible on demand NO YES Bind a run to the exact data state it used NO YES Show a diff of what changed since the last good run NO YES Restore a past state to inspect it during an audit NO YES Assign accountability and approval workflows YES NO Governance defines the rules. Operating control makes a specific AI result reproducible against the exact data state it used. A quick self-diagnostic Run your own program through this short test. If you cannot answer several of these with a live artifact rather than a policy sentence, your framework is ahead of your evidence. Pick a production AI output from last quarter. Can you name the exact data state that produced it? Can you diff that state against the current one and see what changed, without manual archaeology? Can you restore the earlier state and re-run to inspect the result? When your framework says "traceable," is there a run-level record that satisfies it, or only the requirement? If an auditor asked "show me which data produced this decision," would you answer with a reproduced state or a promise? Where it fits in CUBIG's operating layer For production AI, the question is not only which model ran; it is which data state and execution conditions produced the result. Syntitan, CUBIG's AI-Ready Data Platform, scores enterprise data on six readiness axes (Usability, Integrity, Context, Consistency, Reproducibility, and Traceability), rebuilds what blocks execution, and binds every AI or agent run to a data state you can diff and reproduce. That is the operating layer beneath your governance framework. The framework names Reproducibility and Traceability as goals; Run Binding and the Traceability axis are how they become real at the moment a run happens. You keep your policy where it belongs, in risk and legal, and you gain the evidence that lets the policy stand up under review. For a deeper treatment of the run-level mechanics, see what Run Binding is and how to reproduce an AI incident. This article compares the concepts; for the same distinction drawn at the product level, with CUBIG's operating layer set beside AI governance platforms, see CUBIG's operating layer vs. AI governance platforms. How to adopt both together Start from the framework you already have, because it names your obligations. Then, for each requirement that uses the words "reproducible," "traceable," or "auditable," ask what run-level artifact would satisfy an outside reviewer. Wire operating control to produce exactly that artifact, and the requirement moves from aspiration to something you can demonstrate on a Tuesday afternoon when the request arrives. Teams that do this in the wrong order tend to buy more policy and more dashboards, then discover at audit time that neither can restore a past state. The order that works is: keep the policy, add the proof. Performance and coverage here are representative until you reproduce them on your own data, which is exactly the point, because the reproduction is the evidence. For a worked example of what that proof looks like when a decision is later challenged, see how an audit trail for AI decisions in the public sector lets a reviewer restore the exact data state and execution conditions behind a past decision. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "Reproducible Cohort Analysis in Healthcare: A Proof Example" url: "https://cubig.ai/articles/healthcare-reproducible-cohort-analysis/" source: live --- Reproducible cohort analysis means any patient population you define for research or clinical work can be rebuilt later, exactly as it was, from a recorded data state, so the same definition returns the same group and the same finding survives a second look. The stakes are concrete. The FDA's approach to AI in software as a medical device is total-product-lifecycle oversight, backed by Good Machine Learning Practice guiding principles, and oversight across a lifecycle assumes results that can be revisited. In a hospital or a research group, that expectation shows up as a specific failure mode: a cohort that cannot be rebuilt, and a result no reviewer can stand behind. What reproducible cohort analysis actually requires A cohort is not just a query. It sits on top of a stack of moving parts: diagnosis and procedure code sets, terminology mappings, eligibility rules, a data window, and the reference tables that resolve all of them. When a team reports a finding, that finding belongs to one specific combination of those inputs at one moment in time. The problem is that the query text stays the same while the stack underneath it shifts. A code set gets updated to a new revision. An inclusion criterion is refined after a chart review. A mapping table absorbs a correction. Nobody edits the SQL, yet the population it returns next quarter is a different group of people. This is the same quiet failure we describe in the hidden cost of stale reference data: nothing errors, so nothing warns you. Reproducibility, then, is not a property of the code. It is a property of the whole data state that produced the result. To make cohort analysis reproducible you have to capture that state and be able to return to it on demand. Where cohort analysis breaks in practice Consider a research team studying outcomes for patients with a chronic condition. They define the cohort by a set of diagnosis codes, an age range, and a two-year lookback on encounters. They run the analysis, publish an internal finding, and move on. Six months later a regulator, an internal review board, or a journal reviewer asks them to reproduce the result. The team reruns the "same" cohort and the N is off by several hundred patients. The lookback window was recalculated against a refreshed encounter table; two of the diagnosis codes were remapped when the terminology set was updated; one eligibility rule was tightened. Each change was reasonable on its own, but together they mean the original cohort no longer exists anywhere, and neither does the ability to explain the finding it supported. This is not a privacy incident and it is not a modeling error. It is a data-state problem, and it is why so much clinical and research AI work becomes hard to defend the moment someone asks a second time. The wider research world knows the feeling: when Nature surveyed 1,576 researchers, more than 70% had tried and failed to reproduce another scientist's experiment. How a bound data state makes cohort analysis reproducible The fix is to bind each cohort analysis to a recorded Release State: a captured version of the cohort definition, the code sets, the mappings, and the data window that produced a given result. Three operations follow from that binding. Run Binding ties a specific result to the exact state that produced it, so "the cohort behind this finding" is an address you can return to, not a description you have to reconstruct from memory. Diff shows why a re-run returns a different population: it names the updated code set or the shifted criterion instead of leaving the team to guess. Reproduce rebuilds the original state so a reviewer can inspect the cohort as it stood when the finding was made. You can read more about the comparison mechanics in how diff works on data state. Crucially, none of this requires moving raw patient records. The system records the state and the conditions of execution, so reproducibility and patient-data handling stay separate concerns rather than trading off against each other. The handling side is tightly specified: HIPAA recognizes exactly two de-identification paths, Expert Determination and Safe Harbor, which removes 18 specified identifiers. Rebuilding a cohort is a different problem, and neither path solves it. Teams that run analytics directly on protected health information can pair this with the patterns in healthcare workflows on PHI. What changes when a cohort is bound versus not Question a reviewer asks Unbound cohort Bound to a Release State Can you rebuild the exact population from the original run? No Yes Why does the re-run return a different N? Guesswork Named by Diff Which code set and data window applied? Often unknown Recorded in the state Does answering require moving raw records? Usually No How long to produce an audit answer? Days of forensics Direct replay The point of the table is not that reproducibility guarantees identical downstream statistics. It is that the cohort stops being a moving target. The numbers still belong to the analysis; what changes is that the conditions behind them are recoverable, so a result is representative until you reproduce it on your own data and confirm it. A quick self-diagnostic for your cohort work Run this test against a cohort your team reported in the last year. If you cannot answer yes to most of these, your cohort work is not yet reproducible: Can you point to the exact code sets and terminology mappings that were in effect when the cohort was defined? If you rerun the definition today, can you explain every difference in the resulting N? Can a reviewer inspect the original population without you rebuilding it by hand? Is the data window pinned to a recorded state, not to "whatever the table holds now"? Can you produce all of the above without exporting raw patient records? Where it fits in CUBIG's operating layer For production AI, the question is not only which model ran; it is which data state and execution conditions produced the result. Syntitan, CUBIG's AI-Ready Data Platform, scores enterprise data on six axes, Usability, Integrity, Context, Consistency, Reproducibility, and Traceability, rebuilds what blocks execution, and binds every AI or agent run to a data state you can diff and reproduce. For a cohort, that means the definition, code sets, and window travel with the result as a verifiable data state rather than living in someone's memory. To see how this connects to the broader category, start with what AI-ready data means and the mechanics of Run Binding. The healthcare payoff is direct: a reviewer can say "show me this exact cohort as it was defined then" and get an answer instead of an apology. Reproducible cohort analysis turns a fragile finding into one your team can defend twice. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "Operating Control vs CI/CD for ML: What Each Reproduces" url: "https://cubig.ai/articles/operating-control-vs-cicd-for-ml/" source: live --- Operating control vs CI/CD for ML is the difference between a pipeline that ships reliably and a production result you can reproduce: CI/CD reproduces the build, while operating control binds each run to the data state that actually produced the answer. The gap matters because even a mature pipeline stops at the build. Google Cloud's MLOps guide grades the practice into maturity levels 0 through 2 and singles out continuous training as the one piece that only ML systems need on top of standard CI/CD. Automate all of it and you can still ship a model flawlessly yet fail to explain the number it produced last Tuesday. What CI/CD for ML actually reproduces Continuous integration and continuous delivery for machine learning automates the path from commit to production. It versions code, runs automated checks, packages models, and controls rollouts and rollback. When the pipeline works, the same commit yields the same deployment every time. That discipline is real and worth keeping. CI/CD gives teams model versioning, reproducible builds, and a clean audit of what shipped and when. If your concern is "did this exact pipeline definition make it to production," CI/CD answers it well, and I would never suggest replacing it. The limit is narrow and specific. CI/CD pins the code and the pipeline definition; it does not pin the live data that pipeline reads once it is running. The build is frozen, but the world the build operates on keeps moving. Why a reproducible build still gives a non-reproducible result A pipeline can be byte-for-byte reproducible and still produce an answer nobody can recreate three months later. The reason sits outside the build entirely: the production data state changed after the build shipped. Operating control vs CI/CD for ML: what each one binds The cleanest way to see the split is to ask what each layer binds to a run. CI/CD binds code and artifacts. Operating control binds the data state and the execution conditions. You want both, because a production result depends on both. Question CI/CD for ML Operating control Reproduces the build and deployment Yes No Versions code, models, pipeline definition Yes Partial Pins the live production data state per run No Yes Diff two runs by data state, not just code No Yes Rebuild the exact answer an old run gave Partial Yes Ships and rolls back deployments Yes No Read the table as a division of labor, not a contest. CI/CD owns shipping. Operating control owns the run. The Partial marks are where the two overlap: CI/CD versions the pipeline that transforms data, and operating control records enough to rebuild an answer, but neither covers the other's core job. How the two fit together in one flow Operating control is not a replacement for your ML CI/CD; it runs inside it. Keep CI/CD for build and deploy, then add a step that binds each production run to a Release State: a captured, addressable record of the data state and execution conditions behind that run. With both in place you can make two statements instead of one. CI/CD lets you say "this commit produced this deployment." Operating control lets you add "this deployment, on this data state, produced this result," and then Diff that state against another run or Reproduce it on demand. Together they cover both halves of reproducibility, the code half and the data half. The code half already has a working standard: the Reproducible Builds project counts a build as trustworthy only when a deterministic process and a documented environment let an independent party rebuild the artifact and match the output. A bound Release State brings that same bar of independent verification to the data half. This is where the six readiness axes come in. Before a run is worth binding, the data behind it has to be operable, so Syntitan scores enterprise data on Usability, Integrity, Context, Consistency, Reproducibility, and Traceability, then rebuilds what blocks execution. Once the state is sound, every run binds to it. A quick self-diagnostic Run this test against your own stack. Google's ML Test Score takes the long route, 28 tests for production readiness that weigh data and infrastructure as heavily as the model itself; the short list below covers the data-state half. If you answer "no" to more than one of these, your CI/CD is doing its job and your data state is still unbound. Can you name the exact data window a production run read, not just the commit it ran? If a source schema changed last month, would your pipeline logs show it, or would the build still read green? Can you Diff two runs by their data state and see what moved? Given a result from six months ago, can you Reproduce it, or only rebuild the deployment? When an auditor asks "why this output," do you have the state, or only the code? Where it fits For production AI the question is not only which model ran; it is which data state and execution conditions produced the result. That is the layer operating control lives in, and it sits beneath your CI/CD rather than beside it. The Release State it binds is the same unit that diffing a data state works on, and model versioning alone leaves the same gap CI/CD does. Syntitan, CUBIG's AI-Ready Data Platform, is where this runs. It scores enterprise data on the six axes above, rebuilds what blocks execution, and binds every AI or agent run to a data state you can diff and reproduce, so a green pipeline and a reproducible answer finally mean the same thing. For the broader picture, see what AI-ready data means. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "Operating Control vs Feature Store: What Each One Records" url: "https://cubig.ai/articles/operating-control-vs-feature-store/" source: live --- Operating control vs feature store is a division of labor, not a rivalry: a feature store serves consistent features to training and inference, while operating control records the AI-ready data state those features belonged to and binds each run to a state you can diff and reproduce. Feature stores earned their place by killing a specific bug, train/serve skew, so the line between them and operating control deserves a careful hand. The pattern has real pedigree: Uber's Michelangelo platform introduced the feature store, with roughly 10,000 canonical, shared features that teams consume online and offline by name. Serving the right numbers, though, is not the same as being able to rebuild the exact conditions a past run executed on. A feature store closes part of that gap. It does not close all of it. What a feature store actually solves A feature store centralizes feature definitions, serves the same features online and offline, and enforces point-in-time correctness so a model trained on last quarter's data sees the values that were true then, not the values as they look today. For consistent, reusable features shared across many models, it is the right infrastructure and there is no reason to replace it. The scope is precise, though. A feature store guarantees that a named feature is computed the same way wherever it is read. That is a statement about the columns it owns. It is not a statement about everything else the run touched. Why operating control vs feature store is the wrong fight Two runs can pull byte-identical features from the same store and still diverge, because something outside the feature definitions moved: the reference table a feature joined against, the schema of an upstream source, the permission scope that decided which rows were visible, the wider data window the model conditioned on. The store kept its own inputs honest. It never claimed to capture the full state around them. The pattern is common enough to have a name: in a CHI study of high-stakes AI, 92% of practitioners interviewed had experienced data cascades, compounding downstream problems triggered by upstream data issues that nobody treated as AI work. Operating control fills that space. It treats the entire data state a run executed against as one addressable object, a Release State, and it binds the run to that object through Run Binding. When a result later needs explaining, you do not reconstruct the world from memory. You open the state, Diff it against another, and Reproduce the run on the state it was bound to. Serving features and recording state are simply different jobs, and asking one tool to do both is where teams get stuck. Feature store and operating control, side by side The cleanest way to see the relationship is to lay the two capabilities against the questions a production team asks after a result looks wrong. Question after a result Feature store Operating control Were the named features computed the same online and offline? Yes Relies on the store Was train/serve skew prevented for those features? Yes Relies on the store Is the full data state around the run captured as one object? No Yes, as a Release State Can I diff this run's state against a passing run? No Yes Can I reproduce the exact run months later under audit? Partial Yes, through Run Binding Does it record schema, reference data, and permission scope? No Yes Read the table as a handoff rather than a scorecard. Everything the store owns, operating control leans on. Everything the store leaves open, operating control records. How the two work together in practice The pattern is straightforward once the boundary is clear. You keep serving features from the store, and you wrap each run in a Release State. The feature store guarantees consistency of the inputs it defines; Run Binding, Diff, and Reproduce cover the state of the run as a whole. When something breaks, you have both halves in front of you: the feature values, and the data state they sat inside. Google Cloud's MLOps guide grades this march toward automation from manual level 0 to fully automated level 2, and the further up that scale a pipeline sits, the more the recorded state matters when a retrain needs explaining. Picture a fraud model that flagged fewer cases this week than last. Your feature store confirms the transaction features were computed identically, so the usual suspect is ruled out fast. With operating control you go further and Diff this week's Release State against last week's. The features match, but a merchant-category reference table was refreshed midweek, and a permission change narrowed which regions the run could see. Neither shift lived in the feature definitions, yet both moved the output. Because each run was bound to its state, you can Reproduce last week's run on this week's data and confirm the cause in minutes instead of arguing about it for a day. For production AI the question is rarely only which model ran. It is which data state and execution conditions produced the result. Syntitan, CUBIG's AI-Ready Data Platform, scores enterprise data on six axes, Usability, Integrity, Context, Consistency, Reproducibility, and Traceability, then rebuilds what blocks execution and binds every AI or agent run to a data state you can diff and reproduce. The feature store makes the features trustworthy; the platform makes the whole run rebuildable. A quick test: do features alone explain your run? Run this check against your own pipeline before you decide the feature store has you covered. If you answer no to two or more, the gap operating control fills is already costing you. Can you name every reference table a run joined against, at the version it used that day? If an upstream schema changed last month, can you tell which runs saw the old shape? Can you reproduce a six-month-old run's inputs without asking three people what changed? When two runs disagree, can you diff their full data state, not just their feature values? Is the permission scope that filtered visible rows recorded alongside the run? Where it fits in the CUBIG operating layer A feature store lives at the input layer, keeping specific features consistent. Operating control lives one level up, at the operating layer for AI-ready data, where the concern is the Verifiable Data State a run was bound to and whether you can return to it. The two do not overlap so much as stack. Serve from the store; record with the platform; and when a regulator, a customer, or your own postmortem asks what produced a result, you answer with a reproducible AI-ready state instead of a best guess. Performance you see today is representative until you reproduce it on your own data, which is exactly what Run Binding lets you do. Try it on your data for free. Run a sample proof and see it on your own workflow. For adjacent comparisons, see how operating control differs from a passive inventory in Operating Control vs Data Catalog and from experiment tracking in Operating Control vs MLflow, and read why a saved Release State beats a dataset snapshot when you need to rebuild a run. The pillar concept sits in what AI-ready data means. --- title: "Operating Control vs Data Catalog: Discovery vs Reproducibility" url: "https://cubig.ai/articles/operating-control-vs-data-catalog/" source: live --- Operating control vs data catalog is not a choice between two versions of the same tool: a data catalog helps people find, understand, and govern the data an organization holds, while operating control binds each AI run to the exact data state it executed on so you can diff it and reproduce it.Getting data ready for AI starts with knowing what you hold. AWS defines a data catalog as exactly that: an inventory of an organization's data, with metadata organized to support governance and discovery, a system that enforces policies rather than sets them. A catalog answers part of that readiness problem. It does not answer the part that decides whether a given AI result can be defended later. Teams conflate the two because they overlap on metadata. Both touch schemas, definitions, and lineage. But they act at different moments: a catalog works before the run, describing data at rest, and operating control works at the run, recording which state actually produced a result. What a data catalog does well A data catalog inventories the assets an organization holds. It captures lineage, records column and table definitions, manages ownership and access policy, and makes data discoverable across teams that would otherwise duplicate work or query the wrong table. For the questions "what data do we have, what does each field mean, and who owns it," a catalog is the right system. Enterprises invest in cataloging precisely because bad data costs the US economy about $3.1 trillion a year, and much of that waste comes from people acting on data they misunderstood or could not find. Good discovery and clear governance cut directly into that loss. None of what follows is an argument against catalogs. Where the gap opens: operating control vs data catalog at run time A catalog describes data at rest. It tells you what a dataset is in general. It does not record that, on one specific run, a model consumed this window of that dataset, under this schema version, after this preprocessing step, within these permissions. So when a production output needs explaining, whether a regulator asks, an internal review flags it, or an agent takes an action someone disputes, the catalog gives you the dataset's general shape and lineage. It cannot give you the run-time evidence: which exact state produced the particular result you are investigating, and whether you can rebuild that state to check. Discovery metadata sits one layer above the thing you actually need to reproduce a decision. The two jobs side by side The clearest way to see it is a direct comparison of what each system owns. Question Data catalog Operating control What data do we have? Yes No What does each field mean and who owns it? Yes No Which exact data state did this run use? No Yes Can I diff two runs' data states? No Yes Can I reproduce the state behind an output? No Yes Best moment Before the run At the run Read down the two columns and the division of labor is obvious. The catalog governs data as an asset. Operating control governs a run as an event with a fixed, replayable input. Why lineage alone does not close it Data lineage, the catalog feature teams reach for when they want reproducibility, tracks how a dataset was produced and where it flows. That is genuinely useful for tracing the origin of a field, and it is usually the first thing an investigation pulls up. Lineage still describes the pipeline, not the run. It tells you the path a dataset traveled to reach a table. It does not pin the specific version of that table, plus the preprocessing and permissions in force, at the instant a model read it. Two runs a week apart can share identical lineage and still execute on different states, because a backfill landed, a schema migrated, or a reference file went stale between them. Lineage will look the same in both; the outputs will not. Operating control captures what lineage leaves out: the Verifiable Data State bound to each run. How the two work together These systems are complementary, not competing. Use the catalog to find and govern the data going in; use operating control to bind each run to the Release State it executed on. The catalog's definitions and lineage feed the readiness score, especially the Context axis, since knowing what a field means is part of judging whether data is ready. Run Binding, Diff, and Reproduce then supply the run-time half a catalog was never designed to cover. The regulatory bar is written in the same terms: EU AI Act Article 10 requires the training, validation, and testing data of high-risk AI systems to meet quality criteria under documented data-governance practices, relevant, sufficiently representative, and as error-free and complete as possible. A production result has two parents, the model that ran and the data state it ran on, and an audit asks about both. Syntitan, CUBIG's AI-Ready Data Platform, covers the data parent: it scores enterprise data on six axes, Usability, Integrity, Context, Consistency, Reproducibility, and Traceability, rebuilds what blocks execution, and binds every AI or agent run to a data state you can diff and reproduce. The catalog tells you what you have; operating control tells you what actually ran. A quick test: which layer are you missing? Run this on your own stack. If you answer "no" to more than one of these, you have a catalog but no operating control: Can you name the exact data state, not the dataset, that produced a specific AI output from last quarter? Can you diff that state against the one a similar run used, and see what changed? Can you rebuild that state today and re-run the model on it? When a schema migrated or a backfill landed, did anything record that runs before and after used different states? If a regulator asked you to reproduce one decision, would you reach for lineage or for a bound, replayable state? Where it fits Operating control is the run-time layer of AI-ready execution. It assumes you already have discovery and governance, from a catalog or elsewhere, and adds the piece those tools were never meant to hold: a reproducible AI-ready state tied to each run. In CUBIG's terms, that is Release State, Run Binding, Diff, and Reproduce working on top of whatever catalog you already trust. To see how the underlying readiness question is framed, start with what AI-ready data means, then compare the adjacent tools in operating control vs feature store. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "Operating Control vs AI Monitoring Tools: Detect vs Reproduce" url: "https://cubig.ai/articles/operating-control-vs-ai-monitoring/" source: live --- AI monitoring tells you a model's output has drifted; operating control tells you which data state caused the drift and lets you reproduce the run, which is the difference between detecting a problem and being able to fix it.AI monitoring has become standard practice, and rightly so. Teams watch accuracy, latency, and input distributions, and they get alerted when something moves. The trouble starts one step later. An alert tells you the output changed; it does not hand you the data state that produced the change, so the investigation begins exactly where the monitoring tool stops. Drift is not an edge case either: a Scientific Reports study that aged four standard model types across 32 industry datasets found temporal degradation in 91% of the combinations tested. What AI monitoring does well Monitoring is the smoke detector of a production model. It tracks prediction distributions, flags data drift and concept drift, watches for anomalies, and pages someone when a metric crosses a threshold. For catching problems early, this is essential, and a team without it is flying blind. What a monitor gives you is a signal in the present tense: something is off, right now, relative to a baseline. That signal is valuable and incomplete. Knowing that an output drifted is not the same as knowing why, and it is a long way from being able to rebuild the run and prove the cause. Where AI monitoring stops The alert fires, and the questions start. Which data did this run actually read? What did the schema look like that day? Did a reference table get refreshed, or did a preprocessing job change a default? A monitoring tool usually cannot answer these, because it observes outputs and aggregate inputs rather than capturing the resolved data state behind each run. It measures the symptom, not the cause. So the team reconstructs. Someone tries to remember what shipped, digs through migration logs, and guesses at the reference data that was live. This is slow, and worse, it is not evidence. When the drift touches a regulated decision, "we think the reference table changed" does not hold up. Capability AI monitoring Operating control Detect that output drifted Yes Partial Alert in real time Yes No Identify the data state behind a run No Yes Diff a good run against a bad one No Yes Reproduce the run to confirm the cause No Yes Detection without reproduction is half a loop Think of an incident as a loop: detect, diagnose, fix, verify. Monitoring owns detection and does it well. Diagnosis and verification both need the data state, and that is precisely what a monitor does not keep. Without it, diagnosis becomes archaeology and verification becomes hope, because you cannot replay the failing run to confirm your fix actually addressed the cause. Google Cloud's MLOps guide pushes the automation side of this loop, singling out continuous training as the practice CI/CD never needed before ML, yet an automated retrain still inherits the diagnosis problem when nobody can say what the failing run read. The two tools are answering different questions. Monitoring asks, is something wrong now. Operating control asks, what produced this specific result, and can I rebuild it. You want both, in that order. The other half of the loop Operating control keeps what the monitor discards: the data state behind each run, versioned. Every run records a Release State, the resolved data, schema, and configuration it executed on, linked to the run through Run Binding. When a monitor flags drift, you pull the Release State behind the suspect run, Diff it against the last known-good run to see exactly what moved, and Reproduce the run to confirm the cause before you change anything. The readiness axes preserved along the way are Usability, Integrity, Context, Consistency, Reproducibility, and Traceability. This does not replace your monitoring stack. The monitor still raises the alarm; operating control turns the alarm into a fix you can defend. Keep the smoke detector and add the ability to walk back into the room and see what burned. Can you resolve a drift alert, not just receive it? Run this quick check the next time your monitoring dashboard turns red. Google's ML Test Score, a 28-test rubric for ML production readiness, scores data and infrastructure checks with the same weight as model checks, and the questions below apply that weighting to drift: Can you name the exact data state the flagged run read? Can you diff it against the last good run to see what data changed? Can you replay the failing run to confirm the cause before you retrain? Could you show an auditor the same evidence twice and get the same answer? A no to two or more means your monitoring catches problems it cannot help you close. Where this sits in your stack For production AI, the question is not only that an output moved but which data state and execution conditions produced it. That gap is CUBIG's layer. Syntitan, CUBIG's AI-Ready Data Platform, scores enterprise data on the six readiness axes, rebuilds what blocks execution, then binds each AI or agent run to a data state you can Diff and Reproduce. It sits downstream of your monitoring, turning an alert into a reproducible root cause. For more, see what makes data AI-ready, why AI fails after deployment, and how operating control compares to generic MLOps. Any performance figure you see is representative until you reproduce it on your own data. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "Operating Control vs Generic MLOps: Versioning the Data State" url: "https://cubig.ai/articles/operating-control-vs-generic-mlops/" source: live --- Operating control is the layer that records which data state each AI run executed on and lets you reproduce that run exactly, a job generic MLOps leaves undone because it versions pipelines and models while treating the data behind each run as a moving target.Generic MLOps solved the engineering half of machine learning: build a pipeline, deploy a model, monitor it, repeat. It did not solve the question a regulated team gets asked after something goes wrong, which is what exactly this run read and whether you can rebuild it. The discipline is thoroughly mapped, with the MLOps overview by Kreuzberger et al. cataloging the principles, components, and roles it takes to move ML from development into production operation, and the distance between that map and a reproducible run is a large part of why regulated teams get stuck. What generic MLOps covers well A generic MLOps stack does real work, and none of it is in dispute. It runs continuous integration for models, tracks experiments and hyperparameters, packages and deploys artifacts, and watches latency and error rates in production. That machinery is how teams ship models at all, and pulling it out would be a step backward. Google Cloud's MLOps guide gives that machinery a maturity scale, defining levels 0 to 2 and adding continuous training to CI/CD as the practice unique to ML systems. The scope is the point. Generic MLOps versions the things that engineers change on purpose: the code, the container, the model artifact. It assumes the data is an input that flows through, not an object that needs to be pinned. On a static benchmark that assumption holds. Production breaks it. What generic MLOps leaves out The same pipeline, unchanged, produces different results when the data underneath it moves. Four things sit outside what most MLOps tools capture per run: The data window. Which rows and values the run actually read, not the live table that has since changed. The schema. Column types and constraints as they stood at run time, before a migration reshaped them. The preprocessing version. The exact feature logic that turned raw inputs into what the model saw. Reference data. The lookup tables, thresholds, and enrichment sources the run depended on. A pipeline definition tells you these steps ran. It does not freeze what passed through them on a given Tuesday. When an output drifts, the pipeline log says nothing changed while the result clearly did, and that gap is where days of debugging disappear. Question at incident time Generic MLOps Operating control Which model and pipeline ran? Yes Yes Which data state did the run read? No Yes Can you rebuild that exact data state? No Yes Diff two runs to see what data moved? No Yes Reproduce a six-month-old result? Partial Yes Why the gap costs you at the worst moment The failure rarely shows up in a demo. It shows up when a customer disputes a decision, or an auditor asks you to justify one from last quarter. A generic MLOps stack can show which model version served the request and which pipeline processed it, then goes quiet on the reference tables, thresholds, and enriched attributes the model actually read, all of which have since been updated. You are left reconstructing the past from memory, which is not evidence. The scale of the downside is documented: RAND's interview study of engineers and data scientists puts the failure rate of AI projects above 80%, twice the rate of IT projects that involve no AI, and ranks data quality and availability among the five leading root causes. This is not a monitoring problem. Monitoring tells you an output looked off. It does not let you rebuild the conditions that produced it. Reproducibility is a different capability, and generic MLOps was never designed to provide it. What operating control adds Operating control treats the data state as a first-class, versioned object. Each run captures a Release State, the resolved data, schema, and configuration it executed on, and every run records the state it consumed through Run Binding. When two runs disagree you compare their states with Diff, and you restore a prior one with Reproduce. The readiness axes this preserves are Usability, Integrity, Context, Consistency, Reproducibility, and Traceability, the last two of which a pipeline version does not touch. None of this replaces your MLOps stack. A pipeline tool answers which code ran; operating control answers which data state produced the result. Production reproducibility needs both, so you keep the pipeline and add the state. Is your MLOps stack missing the data half? Run this quick check against a real production model you own: For an output from last quarter, can you name the exact data state it read? When a metric moves on an unchanged pipeline, can you diff the data before you touch the model? Can you rebuild a past run without the original engineer or a lucky backup? Would your reconstruction survive an auditor asking for the same evidence twice? A no to two or more means the reproducibility your MLOps stack promises stops where most production failures begin. Where operating control fits At incident time, knowing which model ran is the easy half; the hard half is knowing which data state and execution conditions produced the result. That harder half is the layer CUBIG operates in. Syntitan, CUBIG's AI-Ready Data Platform, scores enterprise data on the six axes, rebuilds what blocks execution, and binds every AI or agent run to a data state you can Diff and Reproduce. It sits alongside your MLOps tooling rather than competing with it, filling the half a pipeline tool leaves out. For the fuller picture, see what separates AI-ready data from clean data, why model versioning is not enough on its own, and how operating control compares to a model registry. Any performance figure you see is representative until you reproduce it on your own data. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "Operating Control vs Model Registry: Which Reproducibility Gap Are You Missing?" url: "https://cubig.ai/articles/operating-control-vs-model-registry/" source: live --- Operating control vs model registry is not a choice between two tools: a model registry versions the model artifact and tells you which model ran, while operating control versions the data state and tells you which data produced the result. The gap is expensive because teams assume the registry covers reproducibility, and it does, for exactly one half of the picture. Consider what even best-in-class lifecycle tooling captures: MLflow Tracking records parameters, metrics, code version, and artifacts for each run, a record of what the model and code did, not of the state of the data the run consumed. A registry that logs model versions cleanly but never records the data state behind a result leaves that half unmanaged, and it is usually the half that moves. What a model registry does well A model registry catalogs model versions, promotes them through dev, staging, and production, and gives you clean rollback and audit of the model artifact. Popular registries, whether inside MLflow, a cloud ML platform, or a standalone tool, all solve the same real problem: they stop teams from emailing model files around and losing track of which one is live. For governing which model is in production, moving a version through approval, and rolling back a bad model, the registry is the right tool. It answers "which model version is serving traffic" with precision, and that answer matters. The trouble starts when people expect the same record to explain a result that changed while the model did not. Where the data-state gap sits The registry's record ends at the model. The same registered version can return different outputs when the data window moves, a schema shifts under the pipeline, preprocessing logic changes, or a permission boundary narrows what the model can read at run time. In each case the registry faithfully reports that the model version never changed, which is true and unhelpful at the same time, because the output clearly did. Google researchers documented this failure mode as underspecification: models that score identically on held-out tests can behave very differently once deployed. "Nothing changed" is the most expensive sentence in a drift investigation, and the registry is often where it comes from. An engineer opens the version log, sees a stable model, and rules out the one place they can inspect, so the search moves to retraining or infrastructure while the actual cause, a moved data state, stays invisible. Operating control closes that gap by treating the data state as a first-class versioned object through a reproducible Release State, Run Binding, Diff, and Reproduce. Consider a fraud model that scored a transaction as safe in March and flags an identical pattern as risky in July. The registry shows the same version both times. What moved was the reference table the model joins against, refreshed on a schedule nobody logged as a release. Without a versioned data state, the team cannot say which table the March run read, so they cannot reproduce the March decision an auditor now wants to see. The model version, the one thing the registry captured perfectly, is the one thing that did not change. Operating control vs model registry: what each one versions Question Model registry Operating control What it versions The model artifact The data state Primary answer Which model is live? Which data produced the result? Rollback target A model version A data state, via Reproduce When output drifts “Version unchanged” “Here is the diff between states” Audit scope Model lineage Data state at execution time Read across any row and the pattern holds: the two tools version different objects, so they answer different questions. Neither row makes the other redundant. A registry that also tags a dataset name does not close the gap, because a name is not a state; the same named dataset can hold different rows, a different schema, and different preprocessing on two different days. Operating control records the state itself, not a label pointing at whatever the store happens to contain now. Registry for the model, operating control for the state Keep the registry for what it governs, the model, and add operating control for the data state, so every run is bound to the Release State that produced it. When two runs of the same registered model disagree, you diff their states and reproduce the prior one instead of concluding "nothing changed" because the version log says so. The axes operating control adds are Consistency, Reproducibility, and Traceability for the data, the half a registry cannot see. Production AI has to answer two questions at once: which model ran, and which data state produced the result. Syntitan, CUBIG's AI-Ready Data Platform, covers the second. It scores enterprise data on six axes, rebuilds what blocks execution, and binds every AI or agent run to a data state you can diff and reproduce. Paired with a registry, you can say both "this model version was live" and "this data state produced this result, and here it is, restored." Four questions for a recent production result Run this quick test against a recent production result. If you cannot answer cleanly on the data side, the registry is doing its job and something else is missing. The framing is not ours alone: Google's ML Test Score, a 28-test rubric for ML production readiness, scores data and infrastructure checks with the same weight as model checks. For a past result, can you name both the model version and the data state it ran on? When a registered model drifts, do you start by diffing the data state, or by retraining? Could you reproduce a six-month-old result for an auditor with the data exactly as it was? When someone says "nothing changed," can you prove it for the data, not just the model? Where it fits in an AI-ready data operating layer A model registry lives in the model layer; operating control lives in the layer between data management and AI execution. That operating layer is where a result becomes reproducible: it records the six readiness axes of Usability, Integrity, Context, Consistency, Reproducibility, and Traceability, and it binds each run to a Release State you can restore. The arc is simple to state and hard to fake: make the data ready, then keep it reproducible. A registry gives you the model side of that arc for free. Operating control gives you the data side, and for regulated teams the data side is usually where the audit lands. Any performance figure you see is representative until you reproduce it on your own model and data. Try it on your data for free. Run a sample proof and see it on your own workflow. Related reading: Operating Control vs MLflow, Model Versioning Is Not Enough, and What is Run Binding? --- title: "Operating Control vs MLflow: Where Reproducibility Breaks" url: "https://cubig.ai/articles/operating-control-vs-mlflow/" source: live --- Operating control vs MLflow is not a contest between two tools: MLflow tracks the model lifecycle, while operating control records the data state behind each production run, and reproducible AI needs both. MLflow's own documentation is precise about what tracking captures: parameters, metrics, code version, and artifacts for each run. That is a record of what the model and the code did, not of the state of the data the run consumed, and the difference shows up after deployment, when a model that reproduced perfectly in an experiment tracker no longer reproduces in production. MLflow is usually where that reproducibility story starts, and, for many teams, where it quietly stops. This piece maps where one layer ends and the other begins, so you stop asking either to do the other's job. "Operating control" here means the AI-ready data operating layer: a reproducible data state built on Release State, Run Binding, Diff, and Reproduce. What MLflow does well MLflow tracks experiments, parameters, metrics, and model artifacts across the training lifecycle. It is genuinely strong at comparing runs during development, packaging a model, and managing the model side of a deployment. If your question is "which experiment produced this model, with which parameters and which training-time metrics," MLflow answers it cleanly and is hard to beat. Notice what that answer is anchored to: the model and the inputs it saw at training time. That anchor is exactly right for the development loop. It is also where a production result can slip away from you, because the thing that moved was never the model. Where the data-state gap opens MLflow pins the model and its training-time inputs. What it does not pin is the execution-time data state a live run actually saw: the data window in production, a schema that shifted after training, a preprocessing step that changed version, or a permission boundary that moved and quietly narrowed what the run could read. When a deployed result drifts and the model version is unchanged, those four are the usual suspects, and they sit outside a model-centric view. The experiment that looked perfectly reproducible in your tracker is not the same object as a production result you can rebuild on demand. Model reproducibility answers half the question; the data state answers the other half. None of this is a new observation: the NeurIPS paper on hidden technical debt in machine learning argued a decade ago that data dependencies cost more than code dependencies in production ML, precisely because nothing tracks them. Consider a concrete case. A risk model ships in March, logged cleanly in MLflow with its parameters and validation metrics. In June, an upstream team renames a column and backfills a default value, and a scheduled job starts feeding the model a slightly different window. The model version never changes, so the tracker still shows the same artifact it always did. The output, however, has moved, and nothing in the model's own history explains why. What you need is a record of the data state each run was bound to, and a way to diff June's state against March's. That record is the operating layer's job, not the tracker's. Operating control vs MLflow, side by side The division of labor is clean once you name it. One layer owns the model and its history. The other owns the data state each run was bound to. Dimension MLflow Operating control Centered on The model lifecycle The data state behind each run Captures Experiments, params, metrics, artifacts Window, schema, preprocessing, permissions Reproduces A training experiment A production result's data state Answers Which experiment made this model? Which data state produced this result? Time of record Training time Execution time, per run Complementary, not competing Keep tracking experiments and models in MLflow. Then bind each production run to a Release State you can diff and reproduce. Together they let you say, for any result, both which model ran and which data state produced it, which is what reproducibility actually requires in an audited environment. The axes operating control adds are Consistency, Reproducibility, and Traceability at execution time, not only at training time. MLflow gives you the model's story; the operating layer gives you the run's story. The MLOps overview by Kreuzberger et al. maps the principles, components, and roles a team needs to move ML from development into production operation, and no single component on that map covers both stories. Neither is a substitute for the other, and treating one as if it covered both is how teams end up with a green experiment log next to an incident they cannot rebuild. How to tell where your gap is You can reproduce a training experiment from MLflow, but can you reproduce a production result's exact data state? When a deployed model drifts on an unchanged version, does your tracker tell you what moved, or only what the model was? Is the live data window and schema each run saw recorded anywhere, or only the training data? If an auditor asked you to rebuild last quarter's decision, could you diff the data state, not just re-list the model version? Do a preprocessing change and a schema change leave a trace you can point to after the fact? If you answered "the model, not the data" to more than one of these, the gap is on the execution side, and no amount of model tracking closes it. Where Syntitan fits Syntitan, CUBIG's AI-Ready Data Platform, supplies the execution-time half that MLflow leaves out. For production AI the question is not only which model ran but which data state and execution conditions produced the result. Syntitan scores enterprise data on six axes: Usability, Integrity, Context, Consistency, Reproducibility, and Traceability. It rebuilds what blocks execution, and it binds every AI or agent run to a data state you can diff and reproduce. In practice: use MLflow for the model, and use the operating layer for the state. That arc, make the data ready and keep the run reproducible, is the missing layer between data management and AI execution. For a fuller definition of that layer, see what AI-ready data means; for the adjacent tools this sits next to, see operating control vs a model registry and operating control vs generic MLOps. Any performance figure you see is representative until you reproduce it on your own model and data. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "The Hidden Cost of Stale Reference Data in Production AI" url: "https://cubig.ai/articles/the-hidden-cost-of-stale-reference-data/" source: live --- Stale reference data is the silent failure mode where the lookups, mappings, and rules an AI model relies on for context have fallen behind reality, so the system keeps answering confidently against an outdated map while nothing on the surface looks broken. Most drift conversations stay fixed on the primary dataset, the rows that feed the model directly. The reference layer sits underneath it: product catalogs, account hierarchies, region codes, policy rules, pricing tables. It changes slowly, rarely trips an alarm, and gets far less scrutiny than the model itself. That is precisely why it is so easy to ship and so expensive to find. More than 80% of AI projects fail by RAND's estimate, double the rate of IT projects that involve no AI, and when its researchers interviewed the engineers behind those projects, data problems kept surfacing among the leading root causes. A reference layer no one is tracking is one of the ways otherwise healthy projects slip into that column. Why stale reference data stays silent When the primary data is wrong, results usually look obviously off, and someone flags them. When the reference data is stale, the model still returns a clean, plausible answer. It has simply resolved a code, a mapping, or a rule against last quarter's version. A discontinued product still maps to its old category. A reorganized account still rolls up to the wrong parent. A region code points at a boundary that changed months ago. The output passes review because it looks completely normal; only the context behind it is wrong. This is the same mechanism that erodes trust across AI systems generally. The answer that is almost right, correct in form and wrong in fact, is harder to catch than the answer that is visibly broken, and a confident response built on a stale lookup is exactly that kind of near-miss. Reviewers approve it, downstream systems act on it, and the error compounds without ever announcing itself. Where the cost of stale reference data actually lands The damage shows up downstream, well away from the model, which is why it is so hard to attribute back to its source. Three patterns recur: Wrong-but-confident answers that pass human review because nothing looks anomalous, so they flow straight into decisions. Slow erosion of trust as occasional errors accumulate without a clear cause, until people quietly stop relying on the system and route around it. Expensive forensics when someone finally traces a bad decision back to a lookup table nobody thought to check, weeks after the fact. This compounding has a name. In a CHI study of high-stakes AI, 92% of practitioners interviewed had experienced data cascades, downstream problems triggered by upstream data issues that nobody treated as AI work, and a stale lookup is a canonical trigger. Because the model and the primary data both look fine, teams often burn the entire investigation in the wrong place, retraining, re-tuning, and auditing the pipeline, while a three-month-old mapping table sits untouched at the bottom of the stack. The table below shows how ordinary reference changes turn into silent errors. Reference data What changed The silent error Product catalog Item discontinued or recategorized Maps to an old category Account hierarchy Reorganization Rolls up to the wrong parent Region or code table Boundary or code update Resolves to a stale region Policy or pricing rules Rule revision Applies last quarter's logic Why monitoring does not catch it Standard monitoring watches the model output and the primary inputs. Both can look unchanged for months while the reference layer underneath goes stale, because a reference table is not part of the request the monitor sees. Accuracy metrics stay flat because the model is doing exactly what it was told; the instructions themselves are just out of date. This is one route into a well-measured outcome: when a Scientific Reports study aged four standard model types across 32 industry datasets, 91% of the combinations lost quality over time. You can add data-quality checks on the reference tables, but freshness is not a quality problem you can validate in isolation. A three-month-old region map is not malformed or incomplete. It is internally consistent and simply describes a world that no longer exists. The only way to know it matters is to connect that table's version to the runs that depend on it. Treat reference data as part of the data state Reference data deserves the same treatment as any other input that shapes a result: a known freshness, captured in the data state a run is bound to. When the reference layer is part of a versioned Release State, staleness becomes measurable rather than assumed. The Context and Consistency axes reflect it, and a Diff between two runs can surface a plain statement like "the region mapping is three months older here" instead of leaving that fact invisible. Reproduce then lets you rebuild a past answer and confirm it was built on the reference data it should have used. A deployed answer inherits its context from the data state behind it, which is why the reference layer belongs inside that state rather than beside it. Syntitan, CUBIG's AI-Ready Data Platform, scores enterprise data on six axes, rebuilds what blocks execution, and binds every AI or agent run to a data state you can diff and reproduce. The reference layer stops being a blind spot and becomes one more thing you can see and compare. A self-diagnostic for stale reference data For any AI workflow, ask what reference data it depends on and how you would know if that data were stale. Run through this quick test: Is each reference table's freshness recorded somewhere, or is it simply assumed to be current? When an answer turns out to be wrong, can you check which version of a lookup it used? Does a reference update register against the runs it affects, or does it change silently in the background? Could you reproduce a decision from last quarter with the exact reference data that was live at the time? If the honest answer to any of these is "we wouldn't know," that gap is the hidden cost: an invisible risk sitting under workflows that otherwise look perfectly healthy. Where it fits in an AI-ready data operating layer Catching stale reference data is one instance of a larger job: making data ready for AI and keeping it reproducible run after run. Syntitan turns reference freshness from an assumption into a number, then holds that number as part of the state each run is bound to, so staleness shows up as a low Context or Consistency score and as a visible Diff when it moves. This is the work of an AI-ready data operating layer, the missing layer between data management and AI execution. Any performance figure you see stays representative until you reproduce it on your own model and data. If you want to go deeper on the failure modes around it, see how preprocessing drift and schema changes break production AI, why these show up as part of the broader pattern in why AI fails after deployment, and how the same discipline supports agent-ready data. The foundation for all of it is AI-ready data. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "Preprocessing Drift in Production AI" url: "https://cubig.ai/articles/preprocessing-drift/" source: live --- Preprocessing drift is what happens when the steps between raw data and model input change, so the model receives different inputs while its version never moves and the raw data looks untouched. It is the quiet sibling of the schema change. The columns look the same, the model artifact is untouched, but what actually reaches the model has shifted, and the output shifts with it. Aging of this kind is the norm: in a Scientific Reports study that followed four standard model types across 32 industry datasets, 91% of the combinations degraded over time. What ages is rarely the artifact itself; it is the untracked layer that turns raw data into model input. This article maps where the drift hides, why teams misdiagnose it, and how to bring preprocessing under the same reproducibility discipline you already apply to the model. Where preprocessing drift actually hides Preprocessing is a pipeline of small decisions, and each one changes the numbers the model sees. None of them is stored inside the model version, and many of them live in code or third-party libraries that are maintained on a completely different schedule than the model. Imputation: how missing values get filled. Switch from "drop the row" to "fill with the column mean" and the distribution moves, so the model learns a signal that was never in the raw data. Scaling and normalization: refit a scaler on a fresh window and the same raw value maps to a different scaled value, which means identical inputs land in different places. Encoding: category handling that adds, drops, or reorders levels quietly changes what each encoded position means, and the model reads those positions literally. Tokenization and feature extraction: a library upgrade that changes how text or signals are parsed, with effects that ripple through every downstream feature at once. Any single one of these can move the output enough to matter. The trouble is that the change leaves no mark on anything a team normally inspects when results go sideways. Why teams misdiagnose it When the output drifts and the model version has not changed, most teams reach straight for a retrain or a model rollback. The reflex runs deep, and mature stacks even automate it: Google Cloud's MLOps guide treats continuous training as the pipeline practice unique to ML, so retraining is always the nearest lever. If the real cause was a refit scaler or an upgraded tokenizer, those two remedies do nothing: the drift returns on the very next run, because the layer that changed is not the layer anyone is watching. The investigation stalls in the wrong place, and the cost keeps compounding while the true cause sits in plain sight. This is also why preprocessing drift and data drift get confused. Data drift is a change in the incoming data itself; preprocessing drift is a change in how unchanged data gets prepared. The symptom, worse predictions with no obvious trigger, looks identical from the dashboard, which is exactly why you need to separate the two before you spend a sprint retraining against a cause that was never the model. Preprocessing step A small change Effect on the model Imputation Drop rows changed to mean-fill Distribution shifts; a false signal is learned Scaling Scaler refit on a new window Same raw value, different scaled input Encoding New or reordered categories Encoded positions change meaning Tokenization Library version upgrade Features parsed differently across the board Bind preprocessing to the data state The reliable fix is to stop treating preprocessing as invisible glue and start treating it as part of the data state. A production AI result is explained as much by the data state and execution conditions behind it as by the model version, and preprocessing sits squarely inside those conditions. When the preprocessing version is captured in the Release State a run is bound to, a change stops being a mystery and becomes a visible difference between two states. From there the workflow is concrete. Run Binding ties each run to the exact data state, preprocessing included, that produced it. Diff two runs and the imputation rule or the refit scaler surfaces as a ranked candidate cause instead of something you hunt for by hand. Reproduce restores the earlier preprocessing so you can confirm the cause rather than guess at it. The six readiness axes this preserves most directly are Consistency and Reproducibility: the promise that the same inputs yield the same model inputs, run after run. A quick self-diagnostic for preprocessing drift The sharpest test is simple. Take a past run and try to reproduce it exactly. If you cannot restore the preprocessing that produced it, the imputation rule, the fitted scaler, the tokenizer version, then you cannot rule it out as the cause of your next incident. Run through this checklist and count how many you can answer with a confident yes: Is your preprocessing versioned alongside the model, or does it live in code that changes with no release note? When a scaler or encoder is refit, is that recorded against the runs it affects? Could you reproduce last quarter's result with last quarter's preprocessing, not today's? When output drifts, can you tell within minutes whether the model, the data, or the preprocessing moved? If two or more of those are a no, preprocessing drift is already a live risk in your stack, and the bill is not small: by the estimate Harvard Business Review cites, bad data costs the US about $3.1 trillion a year. Where it fits in an AI-ready data layer Syntitan, CUBIG's AI-Ready Data Platform, pulls preprocessing into the reproducible state instead of leaving it as untracked glue. It scores enterprise data on six axes, Usability, Integrity, Context, Consistency, Reproducibility, and Traceability, rebuilds what blocks execution, and binds every AI or agent run to a data state you can diff and reproduce. Because the preprocessing behind a run lives inside that bound state, a change shows up in a Diff rather than as mystery drift, and a Reproduce gives you the earlier behavior back to confirm the cause. That arc, make the data ready and then keep it reproducible, is the job of an operating layer for AI-ready data, the missing layer between traditional data management and AI execution. Any performance figure you see for it is representative until you reproduce it on your own model and data. Try it on your data for free. Run a sample proof and see it on your own workflow. For the failure it sits closest to, read Schema Changes Break Production AI. For the broader pattern, see Why AI Fails After Deployment and The Hidden Cost of Stale Reference Data. --- title: "Schema Changes Break Production AI: The Silent Release" url: "https://cubig.ai/articles/schema-drift-production-ai/" source: live --- Schema changes break production AI because a renamed column, a new category, or a type shift alters what actually reaches the model, so its behavior changes while the model version stays exactly the same. The people who make a schema change and the people who feel it sit on opposite sides of the system, which is why this is one of the most common and hardest-to-diagnose ways production AI fails. In Sambasivan and colleagues' CHI study of high-stakes AI, 92% of practitioners had lived through a data cascade, an upstream data change compounding into downstream model failures, and a quiet schema edit is exactly the kind of event that starts one. This article explains why a schema change is effectively an AI release, why it slips past everyone, and how to make schema evolution visible instead of dangerous. Why a schema change is really an AI release A model learns the shape of its inputs. When that shape moves, behavior moves with it, even though no one deployed a new model. The change never appears in the model registry, so from a versioning standpoint nothing happened, yet the output tells a different story. Four edits account for most of the damage: A renamed field stops mapping to the feature the model expects, so a real signal silently drops to zero. A new category arrives that the model never saw in training, and it has no learned response for that value. A type or unit change, such as a number stored as text or a shifted date format, quietly alters how a value is read. A dropped or merged field removes signal the model relied on, or blends two distinct meanings into one column. None of these touch the model artifact. From the output's point of view, a release just shipped, and it shipped untested, because nothing in the workflow treated it as a change to the AI system. Why schema changes are so hard to catch Schema changes are routine and usually well intentioned: a cleanup, a migration, a fresh source, a column split for clarity. They pass through the data team's process, not the model team's, so no one evaluates them as a change to the AI system. By the time the output drifts, the edit is days old and buried in a migration log that nobody connects to the model. This is hidden technical debt in its purest form; the NeurIPS paper that named the concept priced data dependencies above code dependencies for exactly this reason. So two teams read two correct records that disagree. The model team sees an unchanged version and starts inspecting the model, while the data team sees a successful migration and considers the work finished. Neither record says "this is why the AI changed," and that gap is where days of debugging disappear. What each change looks like from both sides The same edit reads as harmless housekeeping to one team and as a behavior change to the other. Laying the two views next to each other is the fastest way to see why these events stay invisible until production moves. Schema change To the data team To the model Rename a column Cosmetic cleanup A feature disappears Add a category More complete data An unseen input Change a type or unit Standardization Values read differently Drop or merge a field Simplification Lost or blended signal The fix: treat the data state as versioned If each run binds to a fixed data state, a schema change becomes visible as a difference between states instead of an invisible event. You compare the Release State behind a good result with the one behind a bad result, and the renamed column or new category surfaces as a ranked likely cause through Diff, rather than a needle you hunt for by hand in migration logs. Reproduce then rebuilds the earlier state so you can confirm the cause before you ship a fix. The readiness axes this preserves are Consistency, which asks whether the input shape held, and Traceability, which asks whether you can point to exactly what moved. For production AI the useful question is not only which model ran but which data state and execution conditions produced the result. How to tell if schema changes can bite you Run this quick self-check against your own stack. If you answer "no" or "I'm not sure" to more than one, an ordinary migration can quietly reshape your model's behavior: When a model drifts, can you diff the schema between the good and bad runs, or do you read migration logs by hand? Do upstream schema changes notify the AI system, or only the data pipeline? Could you reproduce a result from before a recent migration, with the old schema intact? Do you know which schema version each production run was bound to? Does a "successful migration" ever get reviewed as a change to the model, not just to the warehouse? Where it fits in an AI-ready data layer Syntitan, CUBIG's AI-Ready Data Platform, makes schema evolution safe instead of silent. It scores enterprise data on six axes, Usability, Integrity, Context, Consistency, Reproducibility, and Traceability, rebuilds what blocks execution, and binds every AI or agent run to a data state you can diff and reproduce. A schema change is then captured in the Release State and surfaced by Diff the moment behavior moves, so you get a reproducible AI-ready state rather than a mystery. The goal is never to freeze schemas, because data has to evolve. The goal is to make a schema change with production impact as visible as any other release, which is the job of the operating layer between data management and AI execution. Any performance figure you see is representative until you reproduce it on your own model and data. For the broader pattern, see why AI fails after deployment, the related failure mode in preprocessing drift in production, and the foundation in what AI-ready data means. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "How to Reproduce an AI Incident: A 6-Step Playbook" url: "https://cubig.ai/articles/how-to-reproduce-an-ai-incident/" source: live --- To reproduce an AI incident, you rebuild the exact data state and execution conditions that produced the wrong output, then rerun the model against that bound state until the same result appears on demand.Most teams cannot do this. When a production model returns a bad answer, the code is still in git and the model weights are still in the registry, but the data the model actually read has already moved on. The AI Incident Database has cataloged more than 1,500 real-world AI failures, and behind nearly every entry sits the same uncomfortable question: could the operator reconstruct what went wrong? If you cannot reproduce an AI incident, you cannot debug it, defend it to an auditor, or prove the fix actually worked. Why you cannot reproduce an AI incident by default A production AI run reads from live tables, feature pipelines, retrieval indexes, and reference lookups. All of those keep changing. By the time an incident lands in your queue, the upstream table has new rows, the feature job has rerun, and the vector index has been reindexed. You still have the model. You no longer have the input. Software engineers solved this for code decades ago with version control and reproducible builds. AI teams inherited versioned models and versioned code but skipped the hardest layer: the data state at the moment of the run. Model versioning tells you which weights ran; it does not tell you which rows, which schema, and which preprocessing produced the output you are now trying to explain. The practical test is simple. Pick any prediction your system made last Tuesday at 2:14 p.m. Can you feed the model the same bytes it saw then and watch the same answer come back? For most enterprise teams the honest answer is no, and that gap is where incidents become unsolvable. The data state, not the model, is usually the culprit When an AI system that worked in testing fails in production, people reach for the model first. In our experience the input is far more often to blame. A schema migration renamed a column and the feature now arrives as null. An upstream job started truncating decimals. A reference table went stale and the model is scoring against last quarter's thresholds. The weights never changed; the data underneath them did. This is why reproducibility has to bind the run to a verifiable data state, not just a model tag. A Release State captures the exact condition of the data a run consumed, and Run Binding links that specific run to that specific state. With both in place, reproduction stops being an archaeology project and becomes a replay you can trust. A playbook to reproduce an AI incident Here is the sequence we use with regulated enterprises. It assumes you have some way to pin and replay a data state; if you do not yet, the closing section covers where that fits. 1. Freeze the evidence. The moment an incident is confirmed, stop treating the involved tables and indexes as mutable. Snapshot or bind the current data state so it stops drifting while you investigate, because you cannot debug a moving target. 2. Identify the exact run. Find the specific inference or agent run that produced the wrong output: its timestamp, its request, and the Run Binding that ties it to a data state. If runs are not bound to states, you are already reconstructing from logs and guesswork. 3. Rebuild the input state. Reconstruct the data as it existed at run time, meaning the rows, the schema version, the feature values, and the retrieved context. This is the step most teams skip, and it is the step that decides whether reproduction succeeds. 4. Replay against the bound state. Run the same model version against the rebuilt state. If the wrong output reappears, you have reproduced the incident and isolated it to data plus model rather than to environment noise. 5. Diff to localize the cause. Compare the incident-time state against a known-good state. The Diff shows you which fields, rows, or schema elements changed, and that difference is almost always your root cause. 6. Fix, then prove it. Apply the fix, then rerun against both the incident state and current data. Reproduction is not finished when the bug is found; it is finished when you can show the same input no longer yields the same failure. What a reproducible incident record contains Auditors and incident reviewers do not want a narrative. They want the artifacts. Google's ML Test Score rubric made the same point about production readiness, scoring data and infrastructure checks with the same weight as model checks. A record that lets someone else reproduce an AI incident without your help should carry a few concrete things, and the table below separates what most teams log from what reproduction actually requires. Element Commonly logged Needed to reproduce Model version Yes Yes Code commit Yes Yes Input data state Rarely Yes Schema version at run time No Yes Retrieved context or features Partial Yes Run-to-state binding No Yes The pattern is clear. The rows everyone already fills in are the model and the code; the rows that decide reproducibility are the data state and its binding to the run. NIST makes the same point in its AI Risk Management Framework, which calls for AI systems to be documented, traceable, and auditable, not merely monitored. Can you reproduce an AI incident today? A self-check Run this quick diagnostic against your own stack. If you answer no to two or more, incident reproduction is currently a matter of luck. Can you name the exact data state a given production run consumed? Can you rebuild that state weeks later, after upstream tables have changed? Can you diff the incident-time state against a known-good baseline? Can you rerun the model against the rebuilt state and get the same output? Can you hand all of this to an auditor without narrating it yourself? Where this fits in CUBIG's operating layer Reproducing an incident is only tractable when the run was bound to a data state in the first place. For production AI, the question is not only which model ran but which data state and which execution conditions produced the result. Syntitan, CUBIG's AI-Ready Data Platform, scores enterprise data on six axes, Usability, Integrity, Context, Consistency, Reproducibility, and Traceability, rebuilds what blocks execution, and binds every AI or agent run to a data state you can diff and reproduce. In practice that means the Release State and Run Binding this playbook depends on are captured continuously, not scrambled together after the fact. When something breaks, you already hold the bound state, so reproduction is a replay rather than a forensic reconstruction. For the foundational concept behind this, see what AI-ready data means; for the failure modes that trigger these incidents, see why AI fails after deployment and how schema changes break production AI. For how the same reproduction discipline extends to accountable public-sector decisions, see an audit trail for AI decisions in the public sector. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "Why AI Fails After Deployment: It's the Data State" url: "https://cubig.ai/articles/why-ai-fails-after-deployment/" source: live --- Why AI fails after deployment is rarely a model problem: the model that passed the pilot is usually unchanged, and what moved is the data state and execution conditions behind each production run. The pattern has become familiar to anyone who ships models. A system clears every evaluation, earns sign-off, goes live, and then quietly degrades the first time it meets real operating conditions. RAND's interview study of engineers and data scientists puts the failure rate of AI projects above 80%, twice the rate of IT projects that involve no AI, and ranks data problems among the leading root causes. Those failures cluster after deployment, and they cluster around data rather than algorithms. Why AI fails after deployment: the pilot proved the model, not the conditions A pilot runs on a curated dataset that sits still while you evaluate it. That is the correct way to measure a model, and it is exactly why a strong pilot score predicts production behavior so poorly. Google researchers gave the effect a name, underspecification: models that score identically on held-out tests can behave very differently once deployed. The pilot holds constant everything production refuses to hold constant: access is open, the schema is frozen, the data window does not move, definitions are agreed, and one team owns the whole thing. Production introduces all of that at once. Access is restricted and partial, the schema drifts as upstream teams do routine work, the data window slides forward every day, and definitions differ between the team that produced the data and the team consuming it. The output looks like a model regression, but it is really an environment the model was never shown. What actually changes on the way to production When a deployed result drifts and the model version is unchanged, the cause is almost always in the state around the model. This is the norm, not an edge case: a Scientific Reports study that aged four standard model types across 32 industry datasets found temporal degradation in 91% of the combinations tested. The usual suspects are mundane, which is part of why they get missed. This is the data drift that observability dashboards rarely name, because they are watching the model instead of the data state feeding it. The data window moved. The run saw a different slice than the pilot: a fresh month, a new season, a population that shifted. The same query returns different rows. A schema changed. A column was renamed, retyped, or split during a routine migration. For the data team it was one line; for the model it was a behavior change no one filed as a release. Preprocessing changed version. An imputation rule shifted, a scaler was refit on new data, or a tokenizer was upgraded, so the same raw value reaches the model as a different number. A permission boundary moved. What was accessible and trusted at run time changed, so the model quietly ran on a narrower or different view of the data. None of these touches the model version. From the registry's point of view, nothing happened. From the output's point of view, an untested release just shipped. That gap between what the registry records and what the model actually ran on is the reason so much AI in production drifts without a clear cause. Condition Pilot Production Dataset Fixed, curated, sits still Live, moving window every day Schema Frozen for the evaluation Drifts with upstream changes Access Open to the pilot team Restricted, partial, role-bound Preprocessing One agreed version Refit and upgraded over time What is tested The model The state around the model Why teams debug the model instead of the state The tooling is the tell. Almost all production AI observability is built around the model: version, latency, accuracy on a held-out set, sometimes input-distribution alerts. Very little is built around the data state at execution time. So when something breaks, the only instruments pointing anywhere point at the model, and the team retrains, re-tunes, or rolls back a model that was never the problem. The drift returns on the next run, because the thing that changed was never the thing being inspected. The deeper reason is organizational, the pattern a CHI study of high-stakes AI called data cascades: 92% of the practitioners interviewed had watched upstream data issues cascade into downstream model failures. The change that broke the result was usually made by a different team, through a different process, and logged as routine data work rather than as a change to the AI system. By the time output drifts, the change is days old and buried in a migration log that no one connects to the model. The model team and the data team end up looking at two different records of the same event, and neither record says "this is why the AI changed." How to tell whether your own deployment is exposed A short diagnostic answers this faster than a post-mortem. Run your last incident, or your next one, against these questions: When a result drifts, can you diff the data state between the run that worked and the run that did not, or are you correlating logs by hand? Can you take any past result and reproduce the exact data it ran on: the window, the schema, the preprocessing, and the permissions? Do you know which data state each production run was bound to, or only which model version was live? Do you have a number for readiness across all six axes, or only a "the pipeline is green" feeling? If the answers run to "no," the constraint is not model quality. It is the state of the data reaching the model, which is a different problem with a different fix. Naming that difference is the first step, and it is also why teams that keep chasing model reproducibility alone stay stuck: the model was reproducible; the data state was not. How readiness gets measured "Ready" only means something if you can measure it. Syntitan scores enterprise data on six readiness axes, which turns AI-ready from a claim into a number: Usability, Integrity, Context, Consistency, Reproducibility, and Traceability. Deployments that fail the way this article describes tend to score acceptably on Integrity and Usability while scoring low on Consistency, Reproducibility, and Traceability. That is the profile of data that passes review and then drifts in production, with no clean way to prove what moved. If you want the upstream view of readiness, start with what AI-ready data actually means. Where it fits For production AI the question is not only which model ran, but which data state and execution conditions produced the result. Syntitan, CUBIG's AI-Ready Data Platform, answers both halves. It scores enterprise data on the six axes, rebuilds what blocks execution while preserving the context ordinary pipelines drop, and binds every AI or agent run to a Release State you can diff and reproduce. So when a deployed result is questioned, the answer is not "the model degraded" but "here is the data state behind the good run, here is the diff against the bad one, and here is that state restored." That arc, make it ready and keep it reproducible, is the job of an AI-ready data operating layer, the missing layer between data management and AI execution. The mechanics of that binding are covered in how Run Binding ties a run to its data state, and the failure modes in how schema changes break production AI. Any performance figure you see is representative until you reproduce it on your own data. Validated lift comes from your own model, in your own environment, not from a benchmark slide. If your AI worked in the pilot and broke in production, the data state is the first place to look, not the model. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "Release State vs Dataset Snapshot: What Reproducible AI Needs" url: "https://cubig.ai/articles/release-state-vs-dataset-snapshot/" source: live --- Release State vs dataset snapshot comes down to one difference: a dataset snapshot is a passive copy of data at a point in time, while a Release State is a scored, run-bound reference point that lets you prove which data produced which AI result.Both a snapshot and a Release State claim to "save the data," so teams treat them as interchangeable until a result gets challenged. That assumption gets expensive. When Nature surveyed 1,576 researchers, more than 70% had tried and failed to reproduce another scientist's experiment, and that was in labs built around controlled conditions. Enterprise AI teams hit the same wall with worse odds: when someone asks how a production result was produced, nobody can rebuild the conditions behind it. Storing the data and being able to tie an outcome to it are two different jobs. What a dataset snapshot actually is A dataset snapshot captures data as it was at a moment in time. It is good for backup, rollback, and cheap storage, and most data platforms make snapshots trivial to take. Data versioning tools sit in the same family: they preserve bytes so you can retrieve an earlier copy. The limit is that a snapshot is passive. It sits in a bucket, disconnected from the runs that consumed it. Nothing inside a snapshot tells you which AI run used it, under what permissions, or whether the result you are now investigating came from this snapshot or another one taken the same week. When output drifts, a pile of copies leaves you guessing, because a pile of copies is not an audit trail. This is an old warning: the NeurIPS paper on hidden technical debt in machine learning argued a decade ago that data dependencies cost more than code dependencies, precisely because nothing tracks them. "We have snapshots" is not the same claim as "we can reproduce this result." Teams discover this the hard way during an incident. A model output looks wrong, someone pulls the snapshots from that period, and then the questions start: which of these did the run read, was preprocessing applied before or after this copy, did the feature window include the last week of records or not? The snapshots hold the raw material, yet the one fact you need, the link from an output back to its inputs, was never captured. That missing link is why a snapshot answers "what did the data look like" but never "what produced this result." What a Release State adds A Release State is the data state used as the operational reference point for AI runs, agent executions, and review. It differs from a snapshot in four ways, and each one matters precisely when something has gone wrong. Scored. A Release State is evaluated for readiness on six axes, so you know whether the data was fit to run on, not merely that it existed. A snapshot of unready data is still unready data. Bound. Through Run Binding, every run connects to the exact Release State that produced its output. That link is the entire point, and it is what a snapshot lacks. Diffable. Two Release States can be compared with Diff to surface what changed between a working result and a broken one, whether that is schema, time window, or preprocessing, returned as a ranked set of candidate causes. Restorable. You can Reproduce a prior state for inspection instead of hoping a snapshot still matches the run's conditions. Release State vs dataset snapshot: side by side Dataset snapshot Release State Nature Passive copy at a point in time Operational reference point Bound to a run? No Yes, via Run Binding Scored for readiness? No Yes, on six axes Diff between two? Manual, if at all Built in Restorable to a run's conditions? Partial Yes Best for Storage and backup Production AI reproducibility Why the distinction shows up under pressure The gap is invisible on a good day and decisive on a bad one. In a regulated setting the stakes get sharper: if you cannot reproduce the exact data a result ran on, you may not be permitted to rely on that result at all. The EU AI Act's Article 12 requires high-risk AI systems to automatically record events over their lifetime so that results stay traceable, and a passive copy sitting in cold storage does not clear that bar. A bound, restorable state does, because it answers the question an auditor actually asks: which data produced this, and can you show it to me now? This is also where data versioning and a plain data snapshot quietly diverge in value. Versioning proves the data existed in some form; a Release State proves the run happened on a specific, ready form. When a customer disputes a credit decision or a regulator reviews a model output, the second claim is the one that holds. Banking supervisors put this expectation in writing back in 2011: the Federal Reserve's SR 11-7 guidance on model risk management expects model outputs to be validated and monitored on an ongoing basis, which assumes you can show what a model actually ran on. When to use which Keep snapshots for what they do well, which is storage and backup. Reach for a Release State when the goal is production reproducibility, when you need to answer "which data produced this result, and can we restore and inspect it?" A snapshot can sit behind a Release State as raw material; the binding and the readiness score are what make the result explainable. The short version: a snapshot remembers the data, and a Release State remembers the run. In practice the two coexist. Your storage layer keeps taking cheap snapshots on its own schedule, and that is fine. The Release State does not replace them; it sits above them, adding the score and the binding at the moment a run happens, so the copies you already keep become traceable rather than merely present. If you are choosing between the two, you are usually asking the wrong question. The real question is whether the runs on top of your data are bound to anything at all. A quick check on your own setup For a past production result, can you name the exact data state it ran on, or only the day a snapshot was taken? Can you Diff two points in time to see what changed, or would that be a manual reconciliation across buckets? If an auditor asked you to reproduce a result, would you restore a bound state, or hunt through snapshots hoping one matches? Do you know whether the data behind a given run was ready to run on, or only that a copy of it exists somewhere? Where it fits in an AI-ready data layer For production AI the question is not only which model ran but which data state and execution conditions produced the result. Syntitan, CUBIG's AI-Ready Data Platform, makes the data state operational rather than passive. It scores enterprise data on six axes, Usability, Integrity, Context, Consistency, Reproducibility, and Traceability, rebuilds what blocks execution, and binds every AI or agent run to a Verifiable Data State you can Diff and Reproduce. That arc, make the data ready and keep it reproducible, is the job of the operating layer for AI-ready data, the layer that sits between data management and AI execution. Any performance figure you see is representative until you reproduce it on your own model and data. Try it on your data for free. Run a sample proof and see it on your own workflow. Related reading: What is Release State?, Model Versioning Is Not Enough, and How Diff Works on Data State. --- title: "How Diff Works on Data State: See What Changed in an AI Run" url: "https://cubig.ai/articles/ai-data-diff/" source: live --- A diff on data state is a field-level, structural comparison between two Verifiable Data States that shows exactly what changed in the data an AI run consumed: which columns appeared or disappeared, which distributions shifted, which categories were remapped, and which preprocessing rules moved. Most production AI teams can compare code and model artifacts, but they still struggle to compare the data state behind two runs. TensorFlow Data Validation shows why schema changes, skew, and drift need systematic checks, and data versioning tools such as DVC make the dataset itself part of the execution record. A diff on data state brings those ideas together for AI operations: it turns “the data drifted” into a specific list of field, distribution, encoding, preprocessing, and lineage changes. Reproducible builds set the same bar for software, an independently verifiable path from source to binary, and it matters here for exactly the same reason: without a reproducible data state, teams cannot explain why one run worked and the next one failed. What a diff on data state actually compares Code diffs work on text. A diff on data state works on the meaning and shape of data, so it compares things a line-by-line text tool would miss. When two states go in, the diff reports differences across a small set of concrete dimensions. - Schema: columns added, dropped, renamed, or retyped; nullability and key changes.- Distribution: shifts in mean, variance, and category frequency for the fields a model depends on.- Encoding and preprocessing: changes to normalization ranges, category-to-index maps, tokenizer settings, and fill rules.- Lineage: the upstream source or join that produced each field, so a change traces back to where it entered. The point is not to flag every byte that moved. Reference tables update constantly, and that is normal. The point is to surface the changes that alter what a model or agent learns from or acts on, and to leave the rest quiet so reviewers are not buried in noise. Why a data diff matters more than a model diff When an AI system regresses, most teams reach for the model first. They compare weights, check the prompt, and re-read the config. That instinct is understandable, and it is often the wrong starting point, because the model artifact frequently did not change at all. The pipeline feeding it did. Consider a churn model that held steady for months and then started missing obvious cases. The weights were byte-identical to the last good release. Upstream, a billing system had split one status field into two and backfilled the old column with a default, so a feature the model leaned on went flat. A model comparison shows nothing. A diff on data state shows the split, the backfill, and the collapsed distribution in one view. This is why data versioning matters as much as model versioning, and why comparing models alone leaves you looking in the wrong place. Data diff versus code diff versus model comparison These three comparisons answer different questions, and teams get into trouble when they use one to answer another. A quick side by side makes the boundaries clear. Comparison Operates on Answers Catches data-state change? Code diff Source text, config files What did an engineer edit? No Model comparison Weights, hyperparameters Did the trained artifact change? No Data diff Schema, distributions, encodings, lineage What changed in the data a run used? Yes Code and model comparisons are useful, and you still want them. But when the code is unchanged and the weights match, the difference that explains a behavior change lives in the data state, and only a data diff surfaces it. How the diff runs against a Verifiable Data State A diff is only as trustworthy as the two states it compares. If either side is a loose "the data as of roughly last week," the comparison is guesswork. That is why the diff operates against a Verifiable Data State: a captured, addressable record of the data an AI run consumed, including schema, distribution summaries, and preprocessing rules. Because each AI or agent run is bound to the exact state it used, you can name two runs and ask for the difference between the data behind them. The diff then walks both states, aligns fields by lineage rather than by position, and reports the meaningful changes. The output is not a wall of raw rows; it is a scoped, reviewable summary an engineer or auditor can act on, which is the kind of documented, traceable record the NIST AI Risk Management Framework asks AI systems to keep. Naming the two states you compare is a form of run binding, which is what makes the whole comparison reproducible instead of anecdotal. Reading a diff without drowning in noise A good diff ranks changes by how much they can move a model's behavior, not by how many cells moved. This quick diagnostic tells you whether your pipeline is set up to answer the question at all.- Can you name the exact data state behind any past AI run, not just the model version?- When two runs disagree, can you produce a field-level diff of the data between them in minutes?- Does the diff separate high-impact changes, such as a dropped feature, from routine reference updates- Can a non-author, such as an auditor, read the diff and understand what changed without reading pipeline code?- Does each changed field trace back through lineage to the source that produced it? If you answered no to two or more, a regression today would send your team to the model when the real change is in the data, and the search would take days instead of minutes. Where a diff on data state fits For production AI the question is not only which model ran but which data state and execution conditions produced the result. Syntitan, CUBIG's AI-Ready Data Platform, scores enterprise data on six axes, Usability, Integrity, Context, Consistency, Reproducibility, and Traceability, rebuilds what blocks execution, and binds every AI or agent run to a data state you can diff and reproduce. The diff is the readable surface of that binding: it is how a team moves from "the model got worse, we think the data changed" to a specific list of what moved, ranked by impact, traceable to source. That capability sits inside a broader picture of AI-ready data, and it pairs closely with how a Release State captures a data state and how Run Binding ties each run to one. If you are still comparing only artifacts, start with why model versioning is not enough, then see how a diff feeds directly into reproducing an AI incident. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "Data Blockers: Restricted, Unusable, Unstable in Enterprise AI" url: "https://cubig.ai/articles/three-data-blockers/" source: live --- The three data blockers are the three specific reasons enterprise AI stalls on its data: it is restricted, so you cannot use it as it stands; unusable, so a model cannot learn from it; or unstable, so a result that worked once cannot be reproduced.Naming which of these data blockers you are hitting is the first practical move, because each one has a different fix. By RAND's estimate, more than 80% of AI projects fail, twice the failure rate of IT projects that involve no AI. When RAND's researchers interviewed the engineers behind those projects, data quality and availability kept appearing among the leading root causes. Many of those failed projects are stuck on exactly one of these three. Teams that skip the naming step spend months on the wrong remedy: cleaning data that was never the problem, or retraining a model when the real issue was access or reproducibility. Blocker 1: restricted data you cannot put in front of a model The data exists and would be useful, but you cannot use it as it stands. Regulation, contracts, or internal policy keep it from leaving its environment or reaching a model. This is not a quality problem, because the values are fine. It is a usability-under-constraint problem, and teams most often mistake it for a dead end. The wrong response is to copy the raw data somewhere less restricted, which trades a blocker for a liability. The better response is to rebuild the data into a form that preserves its structure, its statistics, and the rare patterns that matter, while removing the exposure that blocked it, so a model can learn from it within the limits that apply. Restricted is what the Usability axis is really measuring under constraint: can this be used safely here, not merely does it exist. Blocker 2: unusable data a model cannot learn from The data is accessible but not learnable. Missing values cluster in the wrong places, the rare patterns that carry the signal barely appear, definitions differ across teams, and the context a model needs got stripped somewhere in transformation. The dataset passes basic checks and still produces weak results. In a CHI study of high-stakes AI, 92% of practitioners interviewed had experienced data cascades: upstream data issues compounding into failures far downstream, long after anyone treated them as AI work. This is the blocker closest to traditional data quality, but the bar sits higher. Quality asks whether you can trust the values. Readiness asks whether a model can actually learn the task from what is left. A dataset can be spotless and still unusable: a blank that meant "test not ordered" gets imputed to a column mean, or a bimodal column that signalled two distinct populations gets normalized into a smooth ramp. Fixing it means restoring context, correcting skew, and strengthening the patterns the target metric depends on, which is what the Context and Integrity axes capture. Blocker 3: unstable data whose state you cannot reproduce The data works today and breaks next month with no visible change. Model, prompt, and code are identical, but the data window shifted, a schema changed, a preprocessing library moved a version, or a permission boundary moved under the run. The result is no longer reproducible, and no one can point to what moved. This is an execution-state problem, and it is the one teams most often misdiagnose as a model problem, because everything about the model says nothing changed. It also survives a successful pilot and only surfaces in production. A decade ago, the NeurIPS paper on hidden technical debt in machine learning already ranked data dependencies as costlier than code dependencies in production ML. The fix is to treat the data state behind each run as something you can fix, version, diff, and reproduce, rather than something that drifts in the dark. That is the territory of the Consistency, Reproducibility, and Traceability axes. Three data blockers, three different fixes Naming matters because the responses do not overlap. Apply the wrong fix and the project stays stuck while looking busy: you clean a dataset that was actually restricted, or you retrain against an unstable state that will drift again on the next run. Blocker What is wrong The fix Axes most affected Restricted Cannot be used as-is for compliance reasons Rebuild within constraints, preserving structure Usability, Integrity Unusable Present but not learnable Restore context, fix skew, strengthen rare patterns Context, Integrity Unstable State changes, so the result cannot be reproduced Bind each run to a fixed, diffable, restorable state Consistency, Reproducibility, Traceability The blocker teams underestimate most is the unstable one, because the data looks fine until a result that worked last month cannot be reproduced this month. By then the pilot has shipped and the failure lands in production, where it is expensive to trace. How to tell which data blocker you are on Run this quick self-check before you commit engineering time to a fix: If the useful data cannot be put in front of a model at all for compliance reasons, that is Restricted. If the data is available but the model underperforms and you cannot say why, suspect Unusable: context or signal was lost upstream. If a result that worked before now cannot be reproduced and the model is unchanged, that is Unstable. If all you have is a "the pipeline is green" feeling and no score across the six axes, you cannot yet tell which, and that gap is itself the first thing to fix. Most stalled projects hit more than one blocker at once, which is why a single measurement across all six axes beats arguing about which one it is. Where it fits: readiness as the state past all three blockers "AI-ready" is simply the state on the other side of these three data blockers: data you can use within its constraints, that a model can learn from, and that stays reproducible once it reaches production. For a fuller definition, see what AI-ready data is and how it differs from clean data. Syntitan, CUBIG's AI-Ready Data Platform, turns that state into a number by scoring enterprise data on six axes: Usability, Integrity, Context, Consistency, Reproducibility, and Traceability. A blocker then shows up as a low score on a specific axis instead of a vague sense that the data is not ready. For production AI the question is not only which model ran, but which data state and execution conditions produced the result; Syntitan scores data on those six axes, rebuilds what blocks execution, and binds every AI or agent run to a data state you can diff and reproduce. That arc, making data ready and keeping it in a reproducible AI-ready state, is the job of the operating layer between data management and AI execution. When this state slips in production, the failure often looks like a model problem; why AI fails after deployment traces how that happens. Any performance figure you see is representative until you reproduce it on your own model and data. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "Why an AI Gateway Alone Isn't Enough for Regulated AI" url: "https://cubig.ai/articles/why-ai-gateway-alone-isnt-enough/" source: live --- An AI gateway routes, meters, and logs the traffic between your applications and large language models, but it does not make the confidential data inside that traffic safe to send, nor does it make the run you sent reproducible later. Most teams reach for an AI gateway first because the pain it solves is loud: shadow usage, runaway spend, no central log of who called which model. That work matters. It also stops well short of the two problems that actually keep regulated data off of production AI. IBM's Cost of a Data Breach 2025 prices that shadow usage: breaches involving shadow AI cost $4.63 million on average, $670,000 above the global average, and 97% of organizations with AI-related incidents lacked proper AI access controls. A gateway sees the request go by; it cannot fix the data state the request depended on, and it cannot let the model touch data your policy says it may never see. What an AI gateway actually does, and where it stops An AI gateway (some vendors call it an LLM gateway) is a control point in front of one or many model providers. It handles routing across models, rate limiting, key management, cost attribution per team, caching, and a request log for audit. If your problem is that forty engineers are each holding a personal API key and finance cannot tell what any of it costs, a gateway is the right tool and you should deploy one. The trouble starts when people assume the gateway also makes it safe to send sensitive data, and that whatever the model returned can be defended in a review six months on. It does neither. A gateway inspects and forwards; it can redact obvious patterns with plain masking, but masking a field is not the same as letting the model do the work that field was part of. And a request log records that a call happened, not the exact data state that produced the answer. Those two gaps sit on different axes, so no single gateway feature closes both. The two problems a gateway leaves open Break the work into what a gateway can and cannot own. The first axis is data access: can the model do useful work on records your policy forbids you to expose, such as patient identifiers, account balances, or a live contract. Samsung showed how that axis fails in 2023, temporarily banning generative AI on company devices after employees pushed sensitive internal data into ChatGPT. The second axis is execution state: if the output is ever questioned, can you rebuild the exact conditions that produced it, a bar EU AI Act Article 12 now writes into law by requiring high-risk systems to record events automatically over their lifetime. A gateway touches the wire between these two axes without resolving either. Capability AI gateway What closes the gap Route and load-balance across models Yes Gateway Meter cost and rate-limit by team Yes Gateway Central request log for audit Partial Gateway logs the call, not the data state Run AI on data policy forbids you to expose No A context-preserving data layer Rebuild the exact run months later No A bound, reproducible data state Make scarce or unusable data trainable No A rebuild engine for the data itself Data access: sending structure instead of raw values When a compliance rule says a record cannot leave your boundary, redaction inside the gateway does not unlock the workflow. Strip the account numbers and the model can no longer reconcile the ledger; leave them in and you have shipped regulated values to a third party. Plain masking treats the field as noise to hide, which is why work that depends on that field stalls. A different approach sends the model the structure of the work rather than the confidential values, runs the reasoning, and reconstructs the real answer locally against an internal mapping that never left your environment. In CUBIG's case this is a context-preserving data layer, LLM Capsule, which substitutes sensitive values before egress, lets the model execute on the substituted structure, and reconstructs the result inside your boundary so the workflow actually closes. It runs on the CUBIG Syntitan platform. The point a gateway misses: the goal is to complete the work on confidential data, not merely to hide the data and accept that the work stops. Execution state: what a request log cannot rebuild Say an auditor asks in November why your model flagged a March transaction. Your gateway log shows the call, the prompt, the model version, the response. It does not show the reference table that had been half-populated that week, the preprocessing script that changed in April, or the schema migration that quietly dropped a column. Same model, different data state, and the number you defended in the board deck no longer reproduces. Performance you cannot rebuild is performance you cannot keep in a budget review. This is the second axis, and it belongs to the data, not the wire. Syntitan, CUBIG's AI-Ready Data Platform, scores enterprise data on six axes, Usability, Integrity, Context, Consistency, Reproducibility, and Traceability, rebuilds what blocks execution, and binds every AI or agent run to a data state you can diff and reproduce. For production AI the question is not only which model ran but which data state and which execution conditions produced the result. A gateway records the first half of that sentence and none of the second. Where the third gap sits: data that was never usable There is a case that neither a gateway nor a context-preserving data layer addresses: the data you need does not exist in a form a model can learn from. Fraud labels are too rare, a clinical cohort is too small, or the only records you hold are restricted. Here the fix is to rebuild the data itself. DTS, CUBIG's AI-ready data transformation engine, diagnoses what is missing, transforms it with structure-preserving synthesis, and augments rare patterns without copying restricted originals, using a non-access architecture. It also runs on the Syntitan platform. A gateway forwards whatever you have; it cannot manufacture the signal you never captured. A quick self-diagnostic Run this test before you decide a gateway is enough: Can your team run AI on records that policy forbids you to send to a model provider, and still complete the task? If an output is challenged next quarter, can you rebuild the exact data state that produced it, not just the log line? Do you know which data state each agent run was bound to, or only that the run happened? When a schema or preprocessing step changes, does anything tell you the run is no longer comparable? Is the data you most need to model actually usable, or too rare, too small, or too restricted to train on? If a gateway answers the first two questions, keep it and move on. In most regulated shops it answers neither, which means the gateway is one layer and the confidential-execution and reproducibility layers are separate work. Where it fits with CUBIG's operating layer Read the three tools by the question each one answers. An AI gateway governs the traffic. LLM Capsule lets AI run on confidential data by sending work structure and reconstructing results in place. Syntitan and its Release State bind each run to a data state you can diff and reproduce, so an output stays defensible. DTS rebuilds data that was never usable into data a model can learn on. The gateway and the operating layer are complementary: put the gateway in front for routing and cost, and use the operating layer underneath so the data going through it is one you can expose safely and reproduce later. For the broader picture, see our primer on what AI-ready data means, the deeper treatment of sensitive AI workflow enablement, and the mechanics of LLM data egress. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "LLM Data Egress: What Actually Crosses Into the Model and What Should" url: "https://cubig.ai/articles/llm-data-egress/" source: live --- LLM data egress is the moment confidential data leaves your control by entering a language model's input, which turns the prompt itself into a compliance perimeter: the AI-era counterpart of network egress, where the boundary that matters is no longer the firewall but the field you type into.Network teams have watched egress for decades, because data leaving the network is the event you log and govern. That same logic now applies one layer up, and most enterprises have not moved their controls to match. Samsung made the gap concrete in 2023, when it temporarily banned generative AI tools on company devices after employees leaked sensitive internal data to ChatGPT, exactly the kind of crossing that happens with nobody watching. Why the prompt became a perimeter For most of enterprise history the perimeter was physical, then virtual: the building, then the network edge. AI moved it again. The models that create the most value run on infrastructure you do not own, so getting value from them means handing over the data they need to work on. That handoff is the egress event, and it happens dozens of times a day in places no one instrumented. The hard part is that the egress is usually invisible. A reviewer pastes a renewal agreement into a chat window to summarize the obligations, or an analyst drops a quarter of internal sales figures into a model to find the outliers. Nothing alerts, no firewall rule fires, and the data is simply gone, sitting in a context window governed by someone else's terms. The compliance event already happened, quietly, before anyone could approve it. The volume is measurable: Netskope Threat Labs found that more than a third of the sensitive data employees feed into generative AI apps is regulated data they are legally bound to protect, with source code among the most common categories exposed. What counts as crossing the line The trigger is not storage. It is transmission. The instant confidential values appear in a prompt, whether typed, pasted, retrieved by a RAG step, or returned by an agent's tool call, they have crossed the perimeter. This is why the problem resists the usual controls: you can encrypt data at rest and in transit and still hand it to the model in plain language inside the request. Encryption protects the pipe, and it does nothing about what you decided to put through it. The data that creates the most value tends to be the data you least want to send. Not only personal records, but the confidential business context that makes a task worth doing: pricing logic, contract terms, the operational numbers that describe how the company actually runs. Strip those out and the model has nothing to reason about; leave them in and you have an egress event you cannot defend. The exposure is not hypothetical: Harmonic Security measured that nearly 22% of files and 4.37% of prompts employees sent to generative AI tools in early 2025 carried sensitive content, so the prompt itself is where the crossing happens. How to handle LLM data egress without the values crossing it The way through is to separate two things that look like one thing. The model does not need your raw values; it needs the shape of the work, the structure and relationships and the task itself. So you send the structure across and keep the values at home. In CUBIG's terms this is the Restorable AI Data Boundary: a checkpoint work can pass through, not a wall that blocks it. In practice the prompt carries a structure-preserving substitution of the real document. Tables stay tables. A renewal contract stays a renewal contract, with its clauses and cross-references intact, while the counterparty name, the negotiated price, and the account numbers are stand-ins held in a protected mapping layer inside your environment. The model does real work on a faithful copy. The result comes back and the original values are restored locally, so what your team receives is a finished document rather than a placeholder. That three-part motion, Substitute, Execute, Reconstruct, is what lets the egress happen while the confidential values never join it. How the boundary changes what you measure Once the perimeter is a checkpoint rather than a wall, the metric changes. The old question was how much you managed to block. The better question is how many previously impossible workflows you can now run while the values that triggered the compliance event stay inside the building. That is Workflow Closure: the work actually finishes, end to end, and the finished result is an operational artifact your team uses directly. Use this quick self-check to tell whether a workflow has an egress problem you cannot currently defend: Does the prompt, at any point, contain raw customer names, account numbers, or pricing that would fail an audit if logged? Can a reviewer paste a confidential document into a model without any approval step firing? Do retrieved passages from a RAG index arrive at the model with confidential values still in them? When an agent calls a tool, does the tool's output flow back into the model unfiltered? If a regulator asked what left your environment last quarter, could you answer at the level of the prompt? Answering yes to any of these means the boundary is real and you are already crossing it. The comparison below shows why the usual controls stop short. Control Stops raw values entering the prompt? Lets the workflow still finish? Encryption in transit and at rest No Yes Plain masking of the input Partial No, the model loses the structure it needs Blocking access to the model Yes No Restorable AI Data Boundary Yes Yes, the result is reconstructed locally Plain masking is worth singling out. It removes values, which reads as safe, but it also removes the structure the model reasons over, so the answer that comes back is weaker or wrong. Substituting instead of stripping keeps the document usable while the real values stay home. Where it fits Treating LLM data egress as a perimeter is the practical core of sensitive AI workflow enablement. What crosses the boundary is structure; what comes back is a reconstructed operational artifact your team can use as-is. The substitution and restoration are handled by a context-preserving data layer, in CUBIG's case LLM Capsule, which sits between your real data and the model and runs on the CUBIG Syntitan platform. The same boundary logic extends past text: images, PDFs, and diagrams travel through a multimodal AI data boundary on the same terms. Once egress is no longer the thing blocking the project, the underlying data can become the next focus rather than the excuse. LLM Capsule is how CUBIG keeps your raw values out of the egress event. It substitutes the confidential fields before the prompt leaves and restores them inside your environment, so what crosses the boundary is structure, not data. --- title: "Document Layout Preservation: Why Visual Structure Matters to AI" url: "https://cubig.ai/articles/document-layout-preservation/" source: live --- Document layout preservation is the practice of keeping a document's visual structure, its tables, headings, indentation, and the row-and-column relationships between fields, intact when sensitive values are substituted, so an AI model reads the same meaning it would read from the original. The stakes here are not academic. Samsung temporarily banned generative AI tools on company devices in 2023 after employees leaked sensitive internal data to ChatGPT, and in an enterprise that sensitive data arrives as documents: contracts, statements, forms, scanned pages. When a preprocessing step flattens those documents to strip sensitive values, it quietly destroys the thing the model was supposed to read. Most people picture a document as the words inside it. A model does not read it that way. When it processes a contract or a financial statement, a large part of what it understands comes from where things sit: which cell lines up under which header, which clause is nested under which section, what the indentation says about precedence. Strip that arrangement away and you have not just hidden the sensitive parts, you have changed what the document means. Why visual structure is meaning, not decoration Take a payment schedule. A table with three columns, milestone, due date, and amount, tells the model that each row is one obligation, and that the number on the right belongs to the date in the middle and the event on the left. The relationship lives in the layout. Flatten that table into a paragraph of comma-separated values and the model has to guess which number pairs with which date. Sometimes it guesses wrong. Benchmarks bear this out: on DocVQA, models still fall short precisely on the questions where understanding a document's structure is crucial. The same thing happens with hierarchy. A clause indented under "Termination" is governed by that heading. Move it, or drop the indentation, and a model can read it as a standalone term. Indentation is not styling here; it is the document telling you what depends on what. This is why a naive redaction step so often breaks the work. It treats the page as text with some words blacked out. But the moment you remove a value and leave a gap, you disturb the grid the value was sitting in. The model now reads a table with a hole in it, and a table with a hole in it is a different table. Recovering that grid is a research problem in itself: TableBank exists because identifying the row-and-column structure of tables, especially in scanned images, is hard enough to need dedicated models. What breaks when document layout is lost Consider a review task on a renewal contract where the counterparty's pricing has been blacked out. The reviewer wants the model to flag whether the new terms are worse than last year's. If the redaction collapsed the pricing table into loose text, the model can no longer line up this year's figures against last year's. It will still produce an answer, and the answer will be confident and wrong, because the spatial cue that told it "these two numbers are comparable" is gone. Financial statements fail the same way. A balance sheet means what it means because assets sit above liabilities and the subtotals roll up in a fixed order. Lose the order and a model can misattribute a figure to the wrong line. The output looks plausible, and it is the kind of error nobody catches until an auditor does. EU AI Act Article 10 sets the expectation in law, requiring data for high-risk AI systems to be relevant, sufficiently representative, and as error-free and complete as possible under documented data-governance practices. How structure-preserving substitution keeps document layout intact The alternative separates two things that redaction lumps together: the sensitive value, and the slot it occupies. You can replace the value and keep the slot. A real dollar amount becomes a substitute amount, but it stays in the same cell, under the same header, in the same row. The model sees a complete, coherent table where every relationship it relies on is still present. Only the underlying numbers have changed, and those get restored once the work comes back. This is the core idea behind moving from PII masking to workflow enablement. Plain masking asks "what can we remove and still feel safe." Layout preservation asks a better question: what does the model actually need to read, and how do we keep that readable while the real values stay home. Send the model the structure of the work, not the raw values, then reconstruct the result inside your own systems. The privacy side of that balance is now formally cataloged: NIST's Generative AI Profile lists data privacy and information security among twelve risks it treats as unique to or amplified by generative AI. Approach Sensitive value handled Table and hierarchy preserved Model reads original meaning Plain redaction (black out) Removed, leaves a gap No No Text extraction only Sometimes stripped Partial Partial Structure-preserving substitution Substituted in place, restored later Yes Yes A quick test for your own documents Before you send a redacted document to a model, run it through this short check. If you answer "no" to any item, the layout is likely doing work that your preprocessing step is about to break. Do the tables still have every row and column they started with, including the ones with substituted values? Does each figure still sit under its correct header and in its correct row? Is the section and clause nesting (what is indented under what) unchanged? Can a person read the document and tell it apart from the original only by the swapped values, not by a broken shape? When the result comes back, can you restore the real values into the exact slots they came from? Where it fits Layout preservation is one piece of a larger boundary. The mechanics follow a simple loop: substitute the sensitive values, let the model execute on the structure, then reconstruct the real result locally. In CUBIG's stack that work is done by a context-preserving data layer, in CUBIG's case LLM Capsule, which performs the substitution so the document's shape survives the trip to the model and rebuilds the original values when the result returns. When a document is not plain text but a mix of tables, scanned pages, and images, the same principle has to hold across all of them, which is the subject of the multimodal AI data boundary. And it is a building block of sensitive AI workflow enablement, the practice of running AI on confidential data without that data leaving your environment. All of this runs on the CUBIG Syntitan platform. This is what LLM Capsule does in practice. It substitutes the sensitive values while the tables, headers, and nesting stay exactly where they were, so the model reads the real structure and the originals never leave your environment. --- title: "Model Versioning Is Not Enough for Reproducible AI" url: "https://cubig.ai/articles/model-versioning-is-not-enough/" source: live --- Model versioning answers one question well, which model ran, but it does not answer the question that actually breaks production AI: which data state produced the result.Versioning the model is necessary, and most teams already do it well: MLflow Tracking alone records parameters, metrics, code version, and artifacts for every run. The trouble is that a pinned model artifact creates a false sense of reproducibility, because the thing that usually moves between a working result and a broken one is not the model, it is the data state around it. Google researchers call this effect underspecification: models that score identically on held-out tests can behave very differently once deployed, and much of that gap traces back to results no one can explain or rebuild. This article sets out what model versioning captures, what it silently misses, and why production reproducibility needs both halves. What model versioning captures A model version records the trained artifact: the weights, the architecture, and often the training configuration. Redeploy that exact version and you get the same function back. That is genuinely useful for rollback, for comparing model iterations, and for auditing the model itself. None of that is in dispute. It is also where most reproducibility setups stop, and that stopping point is the problem. The model is the one part of a production AI system that usually does not change between a result that worked and one that did not. What model versioning misses The same model version can produce different outputs whenever anything underneath it moves. Four shifts do most of the damage: The data window. Which slice of data the run actually saw. A model scored on last quarter behaves differently on this one, even with identical weights. The schema. A renamed or retyped column that the data team treated as routine, which the model reads as a different feature entirely. Preprocessing. A changed imputation rule, a scaler refit on newer data, or an upgraded tokenizer, so the same raw value reaches the model as a different number. Permissions and context. What was accessible and trusted at run time, which can quietly narrow or shift the view the model ran on. None of these moves the model version. So when output drifts, the version log states that nothing changed while the result clearly did. The distance between those two statements is where days of debugging disappear, and where teams retrain a model that was never the cause. Model version Data state (Release State) Answers Which model ran? Which data produced the result? Captures Weights, architecture, train config Window, schema, preprocessing, permissions Changes when You retrain or redeploy Upstream data work happens, often invisibly Enough to reproduce a result? No, on its own Yes, with the model Why the reproducibility gap stays invisible Reproducibility is widely assumed to be solved once models are versioned, which is exactly why the gap survives. The change that breaks a result is usually made by a different team, through a different process, and logged as routine data work rather than a change to the AI system. The model registry faithfully reports an unchanged version; the data change sits in a migration log no one connects to the output. You end up with two accurate records, neither of which explains why the AI behaved differently. The registry says the model held steady, the migration log says the pipeline ran as scheduled, and the result still moved. Most of these failures are unversioned state, not unversioned models, so the fix has to sit on the data side. Versioning the state, not just the model The fix is to make the data state a first-class, versioned object: a fixed reference point captured at run time, against which any result can be explained. In Syntitan this is a Release State, and every run connects to the state that produced it through Run Binding. When two runs disagree, you compare their states with Diff and restore a prior one with Reproduce. For production AI the question is not only which model ran but which data state and execution conditions produced the result. Syntitan, CUBIG's AI-Ready Data Platform, scores enterprise data on six axes, rebuilds what blocks execution, and binds every AI or agent run to a data state you can diff and reproduce. The readiness axes this really preserves are Consistency, Reproducibility, and Traceability, the three a model version never touches. Model registry vs Release State: complementary, not competing This is not an argument against model registries or experiment trackers. A registry answers which model; a Release State answers which data state; and production reproducibility needs both. Keep the registry and add the state. The teams that can actually reproduce a six-month-old result, under audit, after a drift, or on request, are the ones that versioned both halves. Google's ML Test Score points the same way: the 28-test production-readiness rubric scores data and infrastructure checks with the same weight as model checks, a bar you cannot clear from a model version alone. A quick test on your own setup Run this diagnostic against your current stack. If the honest answers point at the model only, your reproducibility stops where most production failures begin. Take a result from last quarter. Can you reproduce the exact data it ran on, or only the model that produced it? When output drifts on an unchanged model version, can you diff the state to find what moved, or do you start by retraining? For any run, do you know which data state it was bound to, or only its model version and a timestamp? Can a different team reconstruct a past result without asking the person who first built the pipeline? Where it fits in CUBIG's operating layer Syntitan supplies the half a registry leaves out. It scores data on the six axes, rebuilds what blocks execution, and binds every AI or agent run to a Release State you can diff and reproduce, so alongside "this model version was live," you can also say "this data state produced this result, and here it is, restored." That arc, make the data ready and keep it reproducible, is the job of an AI-ready data operating layer, the layer that sits between data management and AI execution. Any performance figure you see is representative until you reproduce it on your own model and data. Try it on your data for free. Run a sample proof and see it on your own workflow. Related reading: What is Release State?, Release State vs Model Versioning: Why Registries Aren’t Enough, Release State vs Dataset Snapshot, and Operating Control vs Model Registry. --- title: "What Is Run Binding? Tying AI Runs to Data State" url: "https://cubig.ai/articles/what-is-run-binding/" source: live --- Run binding is the practice of tying every AI or agent run to the exact data state that produced it, so the result can be diffed, audited, and reproduced later. Most teams can tell you which model version served a prediction. Far fewer can tell you which data state the model actually read at that moment. That gap is where production AI quietly breaks. Even teams with excellent tracking stop at the model's edge: MLflow Tracking records parameters, metrics, code version, and artifacts for each run, a full record of what the model and code did that still says nothing about the data the run consumed. Run binding closes the part of the gap nobody else records: what the data looked like when a run happened. What run binding actually means When an AI system produces an output, several things come together at once: the model, the prompt or feature vector, the runtime configuration, and the data the system reads from, meaning tables, embeddings, reference lists, retrieved documents. Model versioning captures the first of those. Run binding captures the rest, and it captures them as one linked record. A bound run answers a precise question months later: given this specific output, which data state was live, and can I rebuild that state exactly. Not an approximation, not "the table around that date," but the resolved contents the system saw. When you can answer that, an incident review stops being archaeology and becomes a lookup. Why model versioning is not enough Model versioning tools do their job well. They record weights, hyperparameters, and training lineage. The problem is scope. In production, the model is often the most stable part of the system, while the data underneath it shifts daily: a reference table gets a new row, an upstream job rewrites a column, a retrieval index gets re-embedded. The model version stays "v3.1" through all of it. So when an output looks wrong, the model log tells you nothing changed, because on the model's side, nothing did. The change was in the data state, and the data state was never bound to the run. You end up guessing. Run binding removes the guessing by recording the data side of the run with the same rigor teams already apply to code and model artifacts. Here is the distinction in one view. Question at incident time Model versioning answers Run binding answers Which model served this? Yes Yes Which data state did it read? No Yes Can I rebuild that exact state? No Yes What changed since the last good run? Partial Yes, via diff Is the whole run reproducible? Partial Yes Run binding and the Release State Run binding depends on there being something stable to bind to. That anchor is the Release State: a named, resolved snapshot of the data as it stood at a moment, with its contents fixed rather than left as a live pointer. A Release State is to data roughly what a tagged commit is to code. Once a Release State exists, binding is straightforward: each run records the Release State identifier it consumed, alongside the model and configuration. The run is now a Verifiable Data State, not a loose event. You can Reproduce it by resolving the same Release State and replaying the run, and you can Diff two Release States to see exactly which rows, columns, or documents moved between them. How run binding improves AI reproducibility AI reproducibility has a low bar and a high bar. The low bar is rerunning the same code and getting the same number on a fixed file. The high bar, the one regulated teams actually need, is reconstructing a past production result from records alone, without the original engineer, the original notebook, or a lucky backup. Run binding is what raises you to the high bar. Because the data state is part of the run record, model reproducibility and data reproducibility stop being separate problems. This matters well beyond debugging. EU AI Act Article 12 requires high-risk AI systems to automatically record events over their lifetime; you cannot claim that traceability if the data half of every run is missing. A bound run is auditable by construction, the same property reproducible builds give software: an independently verifiable path from inputs to output. Picture a lending model that declined an applicant in January. In March, a regulator asks the bank to justify that specific decision. With model versioning alone, the team can show which model scored the file but not the reference tables, risk thresholds, or enriched attributes the model read that day, all of which have since been updated. With run binding, the January run carries its Release State, so the team resolves that state, replays the score, and reproduces the exact decision the regulator is questioning. The difference is not cosmetic: one path is a defensible reconstruction, the other is a best-effort narrative that tends to fall apart under a second question. Binding also shortens ordinary incident reviews, because the first thing an engineer checks, "did the data move," is answered before anyone opens the model code. Run this quick self-test on any production AI system you own. If you answer "no" to two or more, your runs are not bound. For a random output from three months ago, can you name the exact data state it read? Can you rebuild that data state today without asking the original engineer? Can you show, as a diff, what changed in the data between two runs? When a metric moves, can you rule out silent data change before you touch the model? Would your answer survive an auditor asking for the same evidence twice? What run binding is not Run binding is not logging. Logs tell you an output happened and roughly when; they rarely let you rebuild the inputs. It is not a data catalog either: a catalog describes where data lives and who owns it, not the resolved contents of a specific run. And it is not backup or snapshotting for its own sake. A snapshot you cannot tie to a run, and cannot diff against another, gives you storage without answers. Run binding is the link, the diff, and the replay together, treated as one operating discipline. Where run binding fits For production AI, the question is not only which model ran but which data state and execution conditions produced the result. That is the layer CUBIG operates in. Syntitan, CUBIG's AI-Ready Data Platform, scores enterprise data on six axes, Usability, Integrity, Context, Consistency, Reproducibility, and Traceability, rebuilds what blocks execution, and binds every AI or agent run to a data state you can Diff and Reproduce. Run binding is the mechanism behind the last two axes: it turns Reproducibility and Traceability from aspirations into something you can point at. If your team already versions models but still argues about "what changed" during incidents, run binding is the missing half. It fits alongside your existing MLOps stack rather than replacing it, sitting where the data state meets the run. For the broader picture, see how this connects to AI-ready data, why model versioning is not enough on its own, what a Release State is, and how Release State compares to a plain dataset snapshot. Try it on your data for free. Run a sample proof and see it on your own workflow. --- title: "What Is Release State? Reproducing the Exact Data Your AI Ran On" url: "https://cubig.ai/articles/what-is-release-state/" source: live --- Release state is the exact, versioned snapshot of the data, schema, and configuration an AI system ran on at a given moment, captured so that any past run can be reproduced or rolled back long after it happened. Teams rarely lose an AI system to a bad model. They lose it to a data state they can no longer rebuild. In a third-quarter 2024 Gartner survey of 248 data-management leaders, 63% either did not have or were unsure they had the data-management practices their AI work required, and Gartner expects organizations to abandon 60% of AI projects through 2026 for lack of AI-ready data. This is not only an enterprise story. A 2023 study in Patterns traced reproducibility failures caused by data leakage across 17 scientific fields and 294 papers, with the root cause sitting in how the data was handled rather than in the model. The break usually surfaces after launch. A model that cleared every test starts returning wrong answers, and no one can rebuild the exact inputs, schema, and settings that produced them. What release state captures A model artifact on its own explains very little about a single prediction. Release state widens the lens to everything that fed the run, then freezes it as one addressable version. Five things travel together: Data snapshot. The precise rows and values the run read, not a live table that has since changed. Schema. Column names, data types, nullability, and constraints as they stood at run time, before any migration reshaped them. Configuration. Model parameters, prompt templates, thresholds, and feature flags active for that run. Transformation code. The version of the preprocessing and feature logic that turned raw inputs into what the model actually saw. Run binding. An identifier that ties all of the above to one production execution, so a prediction points back to the state that produced it. Miss any one of these and reproduction turns into guesswork. A team that saved the model but not the schema will discover that last quarter's migration quietly changed a column type, and the earlier prediction can no longer be recreated. Release state vs model versioning Model versioning and a model registry answer one question well: which artifact shipped. That matters, and it falls short of reproduction. Versioning tracks the model; release state tracks the run. A registry can tell you that version 4.2 went live in March. It cannot tell you which data version 4.2 scored on a given Tuesday, what the schema looked like that day, or which preprocessing build sat in front of it. Data versioning and data lineage tools close part of the gap, though they usually stop at the dataset and leave the configuration and the run link out. Release state binds all of it to a single reproducible point. What it answers Model registry / versioning Release State Which model artifact shipped ✓ Yes ✓ Yes Training & prompt configuration ◐ Partial ✓ Yes Exact input data the run read — No ✓ Yes Schema as it stood at run time — No ✓ Yes Preprocessing / feature-logic version — No ✓ Yes Reference & lookup data used — No ✓ Yes Link to one specific production run — No ✓ Yes Roll back to reproduce a past prediction — No ✓ Yes How release state makes an AI incident reproducible Reproducibility is the difference between fixing an incident and arguing about it. When a complaint arrives about a prediction from three weeks ago, a release-stateful system lets an engineer pull the exact state behind that run, replay it, and compare the result against current behavior. Data lineage shows where each input came from; the release state shows what those inputs actually were at execution time. Diff the two states and the cause stops hiding, whether it was a schema migration or a stale reference table feeding the model. This kind of traceability is also what the NIST AI Risk Management Framework asks for when it calls on teams to keep AI systems documented and auditable. Without it, root-cause work becomes reconstruction from memory, and the answer often lands after the customer has already left. When regulators ask, reproducibility is not optional In regulated industries, reproducing a result is a requirement rather than a convenience. In US banking, the Federal Reserve and OCC guidance on model risk management, SR 11-7, requires a model's development, validation, and deployment to be documented well enough to reproduce its results. In Europe, Article 12 of the EU AI Act obliges high-risk AI systems to log events automatically and to retain those logs for at least six months. Both point at the same bar: any decision has to be traceable back to the conditions that produced it. A model version and a timestamp do not clear it. You need the data state to trace back to. Where release state sits in an AI-ready data stack Release state is one of the properties that separates AI-ready data from data that is merely clean. Clean data is correct today; AI-ready data stays correct and reconstructable across every run, a different and harder guarantee that we draw out in AI-ready data vs clean data. For agents that act on live systems the same property underwrites trust, since an agent's decision is only as auditable as the state you can replay behind it, a point developed in agent-ready data needs semantic context. This is the layer CUBIG operates on. Syntitan, CUBIG's AI-Ready Data Platform, scores enterprise data on six readiness axes, rebuilds what blocks execution, and binds every AI or agent run to a release state you can diff and reproduce, so a result from any point in the past can be rebuilt and defended. Is your data release-stateful? A quick self-check Pick one prediction your system made last Tuesday, then ask: Can you retrieve the exact rows and values that prediction read? Do you know the schema as it stood that day, before any later migration? Can you name the preprocessing and configuration version the run used? Does that prediction link back to a single, addressable run? Could you replay it today and get the same output? A "no" to any one of these marks a spot where an incident would outrun your ability to explain it. Try it on your data for free. Run a sample proof and reproduce a past run from end to end. --- title: "AI Readiness Assessment: The Six Readiness Axes" url: "https://cubig.ai/articles/ai-readiness-assessment/" source: live --- An AI readiness assessment scores your data on six readiness axes: Usability, Integrity, Context, Consistency, Reproducibility, and Traceability. It turns "AI-ready data" from a claim into a measurable score. Each axis is scored from 0 to 100%, and each maps one way data breaks a model in production to one fix. CUBIG runs this AI readiness assessment as a single diagnostic: a dataset gets a percentage on each axis, an overall readiness number, and a ranked list of what to repair first. By Bae Ho, Founder & CEO, CUBIG Corp. · Last updated: June 2026. What is an AI readiness assessment? Most AI readiness assessments grade an organization: its people, budget, and tooling. This one assesses the data, the part a model actually runs on. You score the data on six axes instead of grading it with one number. Each axis checks something different a model needs: whether it can read the data, trust the values, understand the meaning, rely on it across runs, rebuild a past result, and trace every output to its source. A low score on any one axis can block a release. Almost everyone in enterprise AI agrees the data has to be "AI-ready." Far fewer can say what that means, and fewer still can put a number on it. So readiness becomes a feeling: a team looks at a dataset, decides it seems clean enough, ships it, and learns in production whether it was ready. The cost of guessing is old and well documented. Bad data was estimated to cost the US economy about $3.1 trillion a year back in 2016. Gartner now expects organizations to abandon 60% of AI projects through 2026 for want of AI-ready data, and finds that 63% of organizations either lack the data-management practices AI needs or are unsure they have them, which leaves only about 37% confident they have what AI needs. The gap between the teams that ship and the teams that stall is now mostly a data gap. I have watched this play out the same way across regulated deployments. A model passes every test the data team runs, ships, and then drifts the week a data window rolls forward or an upstream schema changes by one column. The model was never the problem. The data state that reached it had moved, and nobody could say when or by how much. Why are there six readiness axes instead of one data-quality score? A model can fail six different ways, and they do not share a remedy. The data arrives unusable, or wrong, or stripped of the meaning the model needed. Or it stays readable but drifts between runs, can't be rebuilt, can't be traced. A single "data quality" number folds all six into one digit and hides which failure is about to break your deployment. The axes pull the failures apart so the score points you at the cause. One low axis is usually enough to block a release, which is why a reading across all six is more honest than one composite metric. When a deployment stalls, a six-axis read shows which axis is dragging and what raising it unblocks, instead of leaving the team to argue about whether "the data is good." What are the six readiness axes? Each axis answers one question a model asks of the data, and each maps a specific production failure to a specific fix. The six readiness axes: each answers one question a model asks of the data. Axis What it asks, and what it maps to UsabilityCan a model consume the data in the form it is in? A locked permission, an unloadable format, or a sensitive column that cannot leave its source system stops the run before quality is even in question. This is where a sensitive field becomes usable without exposing it: a regulated column gets a privacy-preserving stand-in so a model can run on it, which is the masking-versus-synthetic-data decision in practice. Raising Usability gets the run to start on real data. IntegrityAre the values correct and the relationships between fields unbroken? Broken joins, silent corruption, and contradictory records teach a model something false. Integrity here is scored against what the model will actually consume, not just against a schema: a join can be technically valid and still produce duplicated rows the model over-weights, which a schema check passes and a readiness check flags. ContextDoes the meaning a model needs travel with the data? A field named seg_cat_3 tells an agent nothing, and a null might mean zero, unknown, or never-measured; this is the axis an agent leans on hardest. ConsistencyDoes the data behave the same across runs, time, and environments? A stable schema, stable preprocessing, and a defined data window keep "it worked yesterday" from becoming "it broke today." It is the axis most teams never measure, which is why it produces the quietest failures. ReproducibilityCan a past result be rebuilt from the exact data state it ran on? Without it a wrong answer is a dead end; with it, you can rebuild the exact data behind that answer and start a real investigation. TraceabilityCan each value's origin and transformations be verified, and does every run trace back to a fixed state? In a regulated setting this axis often decides whether you are allowed to rely on a result at all. What is an AI readiness score? An AI readiness score is the composite number these six axes produce: a percentage per axis from 0 to 100%, an overall readiness figure, and a ranked list of fixes. The composite is not an average you can game. One low axis can block a release on its own, because a model that cannot read a locked column does not care that the rest of the data is pristine. A Context score drops when fields carry no description, when nulls have no defined meaning, and when the relationships between tables are undocumented; the score reflects how much a model would have to guess. That is what makes the number tell you what to do next: it names the axis, and the axis names the work. Each axis has its own inputs. The table below names what drives each percentage down, so the score is something you can audit rather than take on faith. What lowers each readiness-axis score: the concrete signals behind the percentage. Axis What drags the score down UsabilityShare of columns blocked, locked, or in an unloadable format; sensitive fields the model cannot reach IntegrityFailed-join rate, out-of-range and contradictory values, duplication a model would over-weight ContextFields with no description, nulls with no defined meaning, undocumented relationships between tables ConsistencySchema, preprocessing, or data-window drift measured between runs and environments ReproducibilityShare of runs with no sealed data state behind them TraceabilityShare of values with no recorded origin or transformation history A real readout looks like a row of numbers, not one. A dataset might come back Usability 92, Integrity 88, Context 41, Consistency 60, Reproducibility 30, Traceability 35. The overall figure is gated by the lowest blocking axis, so this dataset is not "63% ready" in any usable sense: Context at 41 is what a model trips on first. The ranked fix list puts adding field descriptions and null semantics at the top, because that single move raises Context the most points for the least work. Is an AI readiness score the same as a data quality score? No. Quality tools check nulls, types, and duplicates, and they should; that work feeds the Integrity axis. But a dataset can pass every quality check and still score low on Context or Consistency, and still break the model the day it runs. McKinsey reports that 51% of organizations using AI have already hit at least one negative consequence, with inaccuracy the most common, and much of that traces to data that was clean by the dashboard's standard but never measured against execution. Trust runs along the same line: in Stack Overflow's 2025 survey, more developers distrust the accuracy of AI output (46%) than trust it (33%). The common objection is fair: a data team that already runs dbt tests, Great Expectations, or a governance catalog will ask whether this is the same thing. Those tools score the Integrity axis well and touch Context, but none of them score against a running model or hold a reproducible data state, so they cannot tell you where you stand on Consistency, Reproducibility, or Traceability. That is the line a catalog-versus-platform comparison draws out in detail. AI readiness versus data quality: what each one actually checks. Data quality score AI readiness score What it checksIs the data tidy: nulls, types, duplicates, rangesCan a model run on it, learn from it, and let you explain the result What it measures againstA schema or a rules dashboardExecution: a real model and metric Can it block a releaseCatches bad values, not missing context or driftYes; one low axis can hold the release Example failure it misses or catchesMisses a clean column whose meaning was droppedCatches the dropped meaning as a low Context score So AI-readiness is the broader claim. A quality check confirms the data is tidy; a readiness read goes further and asks whether a model can actually run on the data and whether you can explain what it produced. A dataset can be clean and unready at the same time, which most teams discover only after a deployment has shipped. We unpack that distinction in What Is AI-Ready Data? and the way clean falls short of ready in AI-Ready Data vs Clean Data. Every platform now claims AI readiness; the six axes are how you ask ready for what. Each readiness axis maps to one production failure and one fix. Axis The failure it catches What raising it unblocks UsabilityThe run never starts: blocked format, locked field, sensitive columnExecution begins on real data, not a curated sample IntegrityThe model learns something false from broken or corrupt valuesThe model trains on data that is internally true ContextThe model misreads a field because its meaning was left behindThe model reads fields as intended, nulls included ConsistencyOutput drifts when schema, preprocessing, or window shiftsThe same input behaves the same way next month ReproducibilityA past result cannot be rebuilt, so it cannot be investigatedAny result can be restored and re-examined TraceabilityNo one can prove which data produced which outputEvery run links back to the exact state behind it How do you act on a readiness score? A readiness score earns its keep only if it tells you what to do. The axes give you two moves. First, the gaps become a plan. Syntitan does not return "your data is 68%." It returns which axes are dragging the score and, once you connect a target model and metric, ranks the fixes by the points each one adds to that model's performance. Raising Usability on a blocked column might be worth six points; trimming low-signal fields, two. The preparation gets ordered by impact instead of by a generic cleanup checklist. Second, the score has to keep its meaning after the data moves. A dataset that scores well today can drift tomorrow when a schema shifts or a window rolls forward, and that is where Consistency, Reproducibility, and Traceability earn their keep, held in place by a fixed reference point. Four operations carry that weight: 1Release State Seals the exact data a run used, so the state behind a result is fixed rather than assumed. 2Run Binding Ties each AI or agent run to the Release State it ran on. 3Diff Compares two states to narrow what actually changed between them. 4Reproduce Returns to the state behind any past result and rebuilds it. The regulated case makes this concrete. When an auditor questions a credit decision or a clinical model's output from eighteen months ago, you reproduce the exact Release State the model ran on, diff it against today's data to show what moved, and trace each value to its source. Government guidance points the same way: the NIST AI Risk Management Framework frames trustworthy AI around documented, reproducible provenance, which is exactly what a fixed data state gives you. Data you can govern and reproduce is what keeps a project in production for years instead of quietly dying. A readiness number that no fixed state stands behind tells you almost nothing the day output drifts, which is why a model nobody changed can still start drifting: the data state moved underneath it. The axes and the state are two halves of one idea. Scoring tells you whether the data is ready to run; the bound state keeps that judgment true after the data changes underneath you. A score on its own captures how confident you are today. That confidence is exactly what erodes between the pipeline that worked last quarter and the same pipeline now. How does Syntitan measure the six readiness axes? Syntitan is the AI-Ready Data Platform built around the six axes. It profiles a dataset's signals per axis to produce each score, and where you connect a target model and metric, it weights the recommended fixes by the lift each one measures against that model. Syntitan then rebuilds what blocks execution without dropping the structure or context a model needs, and binds every AI or agent run to a Release State you can diff and reproduce. The pattern is simple: make the data ready, then keep it reproducible. That is the job of an AI-ready data operating layer, the missing layer between data management and AI execution, and it gives you a reproducible AI-ready state rather than a one-time cleanup. Any readiness or performance figure you see is representative until you reproduce it on your own model and data. How do you run an AI readiness assessment on your own data? (a 5-point check) Five yes-or-no questions to run before your next planning meeting: Do you have a number for readiness across all six axes, or only a "the pipeline is green" feeling? In your gold tables, can you still say why a given field is blank, or has that context already been dropped? When output drifts, can you diff the data state to see what moved, or do you start by retraining the model? Can you take any past result and reproduce the exact data it ran on? In a regulated setting, if you cannot, you may not be permitted to use that result at all. If one axis came back low, would you know which fix raises it, and how many points it buys? If the honest answers run to "no," the thing holding your AI back is the data state that reaches the model. The way to find out is to measure it, on all six axes, every release. None of this is the same as orchestrating tools across a stack; it is making one dataset readable and reproducible before a model ever touches it. --- title: "Databricks vs Syntitan: governing the estate, reproducing the run" url: "https://cubig.ai/articles/databricks-vs-syntitan/" source: live --- Databricks built one of the strongest governance layers in the market with Unity Catalog. It gives an enterprise a single way to govern data and AI assets across the lakehouse: access control, lineage, classification, and policy that reach over tables, models, and now the operational data that Lakebase brings onto the platform. If the job is to govern a whole data estate from one place, Databricks does that job well. Governing the estate is not the same as reproducing a run; the latter, as NeurIPS-led work notes, requires obtaining the same result with the same code and data. Teams line Databricks up against Syntitan because both speak the language of AI readiness, and a few Databricks features look adjacent to what Syntitan does. Put the two side by side and they answer different questions, on different layers. The difference in one lineUnity Catalog governs the whole data estate. Syntitan binds and reproduces the state behind a single AI run. Those are claims on different layers of the stack. Unity Catalog operates over the estate: the catalog of tables and models, the policies that apply to them, the lineage of where data and assets came from. Syntitan operates over a single AI run: it captures and versions the exact data state a model executed on, so that state can be compared against a later one, replayed, and the result reproduced. Reproducing ML results is a known systemic problem: one review found errors affecting 329 papers across 17 fields, often traced to data leakage. A word on Lakebase Lakebase is worth naming, because at a glance it looks like Databricks moving onto Syntitan's ground. It is not. Lakebase is a serverless operational database, a place for applications and agents to read and write transactional data on the same platform as the analytics. It stores live data and serves it fast. That is a real and useful addition. It is an operational store, not a readiness layer, and it does not capture, version, or reproduce the data state behind a specific AI run. Lakebase widens what Databricks governs. It does not change the layer the comparison turns on. Where the line falls The gap shows up the moment a result moves. A model ran last week and gave one answer. It runs this week on a refreshed table and gives another. Unity Catalog can show the lineage of the assets involved and confirm that policy held. It does not, on its own, tell you which data state the working run depended on, what changed between then and now, or how to reproduce the earlier result. Databricks has pieces that sit near this, and they are worth being precise about. Delta tables support time travel, and MLflow tracks the parameters, metrics, and artifacts of a training run. Both are real, and both solve their own problems. Time travel rolls a table back to a past version. Experiment tracking records what a run was configured with. Unity Catalog is not built primarily around capturing the released data state a specific AI run used, binding it to that run, diffing it against a later one, and replaying it to reproduce the result. Storage-level versioning and experiment-level tracking sit on a different layer from run-level reproduction. What reproduction takes Reproduction is the outcome. It rests on a set of mechanisms that an estate governance layer is not built around. Syntitan captures and versions the exact data state behind an AI run, so a team can compare, replay, and reproduce results when conditions change. In practice that means: Snapshot. The exact released state of the data a run executed on, captured at the moment it ran. Versioning. That state held as a versioned release, not overwritten by the next refresh. Diff. A clear comparison of what changed in the data between one run and the next. Replay. The earlier state re-run on demand, so the prior result can be reproduced. The shorter version Unity Catalog governs a data estate from one place, across tables, models, and operational data, and it is strong at that. Syntitan captures and versions the data state behind an AI run, so the result can be compared, replayed, and reproduced when the data moves. Both are forms of AI readiness, for different questions. Most teams running models in production will want their estate governed and their runs reproducible, which is why these sit on top of each other rather than against each other. Read nextEvery platform now claims AI readiness. Ready for what?Snowflake Horizon vs Syntitan: semantic meaning and reproducible AI runsCollibra vs Syntitan: two answers to two different questions About this piece. CUBIG builds the AI-ready data layer between enterprise data and the models and agents that run on it. Syntitan is the product. Capability descriptions reflect each platform's published and shipping focus as of 2026 and are meant to map categories, not to rank quality. --- title: "Snowflake Horizon vs Syntitan: semantic meaning and reproducible AI runs" url: "https://cubig.ai/articles/snowflake-horizon-vs-syntitan/" source: live --- Snowflake has moved Horizon well past a simple catalog. Horizon now provides semantic context, governance, and policy across the data estate: a semantic layer that lets an agent read business meaning instead of raw column names, classification and access policy, and end to end lineage. Snowflake and Databricks compete over this ground, each arguing that its platform is the better place for agents to find, understand, and be governed on enterprise data. For the question they answer, both have strong claims. That question is what the data means, across the estate. A different question sits one layer over and goes unanswered by either platform. An agent produced a result; you need to know which data state it ran on, and whether you can get that result again. Reproducibility is a recognized requirement in its own right: NeurIPS-led guidance calls obtaining the same result with the same code and data a necessary step to trust a finding. The difference in one lineHorizon tells an agent what the data means. Syntitan proves the data state an AI run can reproduce. Meaning and reproducibility are different layers, and a semantic layer does not imply the second. Horizon can tell an agent that a column is revenue, that the customer dimension joins here, that a field came from a given source. That is semantic context across the estate, and it is useful. Shared metadata standards aim at exactly this portability of meaning: Croissant defines a common representation so datasets are discoverable and interoperable across ML tools. It describes what the data means in general. It does not capture and version the specific data state a particular run executed on, hold that state so it can be compared against a later one, or replay it to reproduce the earlier result. Knowing what the data means and reproducing a run are two distinct guarantees. Where the line falls The gap shows up the moment a result moves. A model ran last week and gave one answer. It runs this week on a refreshed table and gives another. Horizon can explain what each field means and trace its lineage. It does not, on its own, tell you which data state the working run depended on, what changed between then and now, or how to reproduce the earlier result. Both platforms have pieces that look adjacent to this, and it is worth being precise about them. Snowflake offers Time Travel and zero-copy clone. Databricks has Delta time travel and, through MLflow, experiment tracking. Those are real, and they solve their own problems. Time travel rolls a table back to a past point. Experiment tracking records the parameters, metrics, and artifacts of a training run. Neither platform is built primarily around capturing, versioning, binding, diffing, and replaying the exact data state behind a specific AI run. Storage-level time travel and experiment-level tracking sit on different layers from run-level reproduction. Syntitan captures that state, versions it, diffs it against the current one, and replays it. Each row is how the two layers answer the same question, not an inventory of features. Horizon's full semantic, governance, and lineage coverage is wider than any single row shows. Capability reflects each product's published focus as of 2026, not a quality judgment.Snowflake HorizonSyntitan The questionWhat does this data mean?Can this AI result be reproduced? For an agentReads business meaning across the estateRuns on a bound, releasable data state Lineage is forUnderstanding and auditReproducing a specific run When a result changesExplains field meaning and lineageDiffs the state and re-runs the prior one ScopeThe data estate and its semanticsA single AI run's data state Sitting on top, not against Syntitan does not replace a data cloud, and it does not ask you to choose between Snowflake and Databricks. It reads a versioned, fixed data state and makes it something a model or an agent can execute on, trace, and reproduce. That works above whatever stores and serves the data underneath. A team can run Snowflake Horizon for semantic meaning across its estate and add Syntitan for the reproducibility of a specific run. The two cover different ground. The second matters once a model is in production and a result has to hold. The shorter version Horizon makes the meaning of enterprise data legible to agents across the estate, and it is strong at that. Syntitan captures and versions the data state behind an AI run, so the result can be compared, replayed, and reproduced when the data moves. Both describe themselves with the language of AI readiness, for different readings of the word. The reading that decides whether a model holds once it ships is reproduction, and most teams running models in production will want their data both governed and reproducible. Read nextEvery platform now claims AI readiness. Ready for what?Collibra vs Syntitan: two answers to two different questionsDatabricks vs Syntitan: governing the estate, reproducing the run About this piece. CUBIG builds the AI-ready data layer between enterprise data and the models and agents that run on it. Syntitan is the product. Capability descriptions reflect each platform's published and shipping focus as of 2026 and are meant to map categories, not to rank quality. --- title: "Collibra vs Syntitan: two answers to two different questions" url: "https://cubig.ai/articles/collibra-vs-syntitan/" source: live --- Collibra is one of the strongest names in data governance, and it earns the position. It gives an enterprise a business glossary, data stewardship workflows, lineage for audit, and policy controls over who may touch what. More recently it has extended into AI governance, with a command center that gives an organization visibility and policy over the AI systems running across it. If the job is to understand, organize, and control a data estate, Collibra does that job well. That job has a formal name: DAMA International defines data management as the practices that deliver, control, protect, and enhance the value of data assets across their lifecycle. Teams sometimes line Collibra up against Syntitan because both now use the language of AI readiness. Put side by side, though, they are answering two different questions, and seeing the difference clearly is more useful than ranking them. The difference in one lineGovernance answers can we use this data? Syntitan answers can this AI result be reproduced? Those are not competing claims on the same territory. They are claims on different layers of the stack. Governance operates over the data estate: the catalog of assets, the rules that apply to them, the lineage of where data came from. A data catalog, in the standard definition, is an inventory of an organization's data assets that helps people find the data they need. Syntitan operates over a single AI run: it captures and versions the exact data state a model executed on, so that state can be compared against a later one, replayed, and the result reproduced. Where the line falls The clearest way to see it is an audit question. Suppose a model produced a decision, and someone asks why. A governance system answers part of that. Collibra can trace governed assets and lineage into AI systems, showing which model, which dataset, and which pipeline were involved, and that policy was followed. That is real, and it matters for the control question. It is a different guarantee from reproduction. Lineage shows the path data took. It is not built to capture, version, replay, and compare the exact data state behind a specific run. Lineage is not reproducibility, and that is the part Syntitan holds. Two layers, two jobs. Capability reflects each product's focus as of 2026, not a quality judgment.CollibraSyntitan The questionCan we use this data?Can this AI result be reproduced? ScopeThe whole data estateA single AI run's data state Lineage is forAudit and complianceReproducing a specific run On a changed resultShows who accessed which assetsDiffs the data state and re-runs the prior one On AIGoverns AI systems across the orgBinds and reproduces the state behind a run When to reach for each If the problem is that the organization cannot agree on what its data means, who owns it, or whether policy is being followed, that is a governance problem, and Collibra is built for it. If the problem is that a model worked in the proof of concept and drifts once it is live, and nobody can reconstruct the data state that worked, governance will not close that gap on its own. That is the layer Syntitan adds, and it runs alongside a governance program rather than replacing it. A team can govern its estate with Collibra and still have no way to reproduce a given run. The two cover different ground. What reproduction takes Reproduction is the outcome. It rests on a set of mechanisms that a governance program is not built to provide. Syntitan captures and versions the exact data state behind an AI run, so a team can compare, replay, and reproduce results when conditions change. In practice that means: Snapshot. The exact released state of the data a run executed on, captured at the moment it ran.Versioning. That state held as a versioned release, not overwritten by the next refresh.Diff. A clear comparison of what changed in the data between one run and the next.Replay. The earlier state re-run on demand, so the prior result can be reproduced.Optimize. The data tuned for the specific model the run uses, then released in that state. Those are the moving parts behind the one-line difference. A governance suite can tell you a run happened and trace its lineage. These mechanisms are what let you reproduce it. The shorter version Collibra makes a data estate understood and controlled, and it is strong at that. Syntitan captures and versions the data state behind an AI run, so the result can be compared, replayed, and reproduced when the data moves. Both are forms of AI readiness, for different questions. Most teams running models in production need their data both governed and reproducible, which is why these sit on top of each other rather than against each other. Read nextEvery platform now claims AI readiness. Ready for what?Snowflake Horizon vs Syntitan: semantic meaning and reproducible AI runsDatabricks vs Syntitan: governing the estate, reproducing the run About this piece. CUBIG builds the AI-ready data layer between enterprise data and the models and agents that run on it. Syntitan is the product. Capability descriptions reflect each platform's published and shipping focus as of 2026 and are meant to map categories, not to rank quality. --- title: "AI-Ready Data: Every Platform Claims AI Readiness — Ready for What?" url: "https://cubig.ai/articles/ai-ready-data/" source: live --- Open the website of almost any data platform and you will find the phrase. AI-ready. Trusted data for AI. Data your agents can rely on. The language has converged, and the convergence makes the market look like one crowded category where every vendor competes on the same promise of AI-ready data. It is four categories wearing one phrase. The vendors saying "AI-ready" answer different questions, and the differences decide which one you need. Even documenting what a dataset contains lacks a standard: the authors of Datasheets for Datasets note the field has no standardized process for it, with severe consequences in high-stakes domains. Strip the marketing and four distinct readings of the word "ready" are in play. A buyer evaluating these tools is usually answering one of them without realizing the other three exist. The idea that readiness comes in levels is not new: an academic framework of data readiness levels argues teams need a common language for how ready a dataset is before a model can rely on it. The four readings of AI-ready data. The readingReady to…The question it answers finddiscoverCan people and agents find the right data and understand what it means? governcontrolCan the organization decide who may use this data, and prove it stayed compliant? queryserveCan the data be stored, joined, and queried fast enough at production scale? reproducereproduce an AI runA model produced a result; can you say which data state it ran on, and get that result again? The first three readings are well served. Mature products own each of them. The fourth decides whether AI holds up in production, and few products are built around it. Google researchers found the same imbalance in practice: in their study of high-stakes AI, data is the most undervalued and de-glamorised aspect of AI, even though it decides whether systems work. 01 The market looks crowded because the words overlap Four families of product have moved toward the same vocabulary from four different starting points. Their origin tells you which question each one solves. Catalogs and data intelligence grew out of discovery. Informatica, Collibra, Alation, and Microsoft Purview help an enterprise see what data it has through a data catalog, what each field means, and where it came from. Several now add a semantic layer so that an agent can read business meaning rather than raw column names. Their answer to "ready" is: the data is described, classified, and understood. Governance suites grew out of control. The same names appear here, because data catalog and governance have merged. The job is policy: access rules, data lineage for audit, masking of sensitive fields, a record that the right people touched the right data. Their answer to "ready" is: the data is permitted and accountable. Cloud data platforms grew out of storage and compute. Snowflake and Databricks store enormous volumes and run analytics and model training over them, with their own catalogs layered on top. Databricks has since added an operational database, Lakebase, so that applications and agents can read and write transactional data on the same platform. That is a real expansion, and it is worth being precise about what it is. Lakebase is an operational store for AI apps. It is not a readiness assessment, and it does not claim to be. Their answer to "ready" is: the data is stored, governed, and fast to query. Synthetic data tools grew out of access. The real data cannot always be shared, or there is not enough of it, so these tools generate a usable stand-in. Their answer to "ready" is: you have data you are allowed to work with. Each of these is a genuine solution to a genuine problem. Discovery, control, scale, and access are all real. None of them is the thing that breaks when a model moves from a proof of concept into production. 02 The reading of AI-ready data that breaks in production The familiar failure looks like this. A model works in the demo. The team runs it on a clean slice of data, the results are strong, the proof of concept is approved. Then it goes live, the production data has shifted, a column was renamed, a join changed, a source was refreshed, and the results stop reproducing. Nobody can say why, because nobody captured the exact state the working run depended on. The instinct is to blame the model. The model is rarely the problem. The data state moved underneath it, and no one had captured the state that worked. This is the reading of "ready" that the other four families are not built around. Discovery tells you the data exists. Governance tells you that you were allowed to use it, with data lineage to prove it. A platform stores and serves it. Some of these platforms have adjacent pieces, storage-level time travel that rolls a table back, or experiment tracking that records a training run's parameters and metrics. Useful as they are, none is built around capturing the released data state a specific AI run used, binding it to that run, and replaying it to reproduce the result. The definition we work fromAI-ready data is a released state of enterprise data that AI can use, trace, and reproduce. That is the missing layer. It sits between the data estate below and the models and agents above, and it is where CUBIG builds. Our AI readiness assessment scores a dataset across six readiness axes: usability, completeness, context, consistency, traceability, and reproducibility. The first four describe whether the data can be used at all. The last two, traceability and reproducibility, are where production breaks, and they are the two that most of the market is not built around. 03 Same word, different jobs Laid against the six readings, the landscape stops looking crowded. Each family is strong in the columns it grew out of and lighter in the columns where the run lives. The three on the right, run binding, diff and reproduce, and a proof run a customer can re-execute, are the part that turns a result into something you can stand behind. Where each platform concentrates, as of 2026. Filled circle: core strength. Open circle: present or developing. Small circle: not a focus. This maps focus, not quality. PlatformDiscovery & metadataGovernance & policySemantic / AI contextRun bindingDiff & reproduceCustomer proof run Informaticadata managementcore strengthcore strengthdevelopingnot a focusnot a focusnot a focus Collibragovernancecore strengthcore strengthdevelopingnot a focusnot a focusnot a focus Alationcatalogcore strengthcore strengthcore strengthnot a focusnot a focusnot a focus Microsoft Purviewgovernancecore strengthcore strengthdevelopingnot a focusnot a focusnot a focus Ataccamadata qualitycore strengthcore strengthdevelopingnot a focusnot a focusnot a focus Snowflake Horizondata cloudcore strengthcore strengthcore strengthdevelopingnot a focusnot a focus Databricks Unity Cataloglakehousecore strengthcore strengthcore strengthdevelopingnot a focusnot a focus Google Knowledge Catalogdata cloudcore strengthcore strengthcore strengthnot a focusnot a focusnot a focus SyntitanAI-ready data layerdevelopingdevelopingdevelopingcore strengthcore strengthcore strength core strength present / developing not a focus Read the table left to right and the established platforms are strong through context, then the marks thin out. Syntitan runs lighter on the left on purpose, because catalogs and governance suites already do that work well, and concentrates on the three columns that turn a run into something reproducible. The two layers are not rivals across the whole table. They meet at one edge and otherwise sit on top of each other. 04 What each one is for Because the categories solve different problems, the honest comparison is not "better" or "worse." It is which question you are trying to answer. Three contrasts make the boundary concrete. vs governance · Collibra, Informatica, PurviewGovernance answers can we use this data? Syntitan answers can this AI result be reproduced? vs the data cloud · Snowflake HorizonHorizon tells an agent what the data means. Syntitan proves the data state an AI run can reproduce. vs the lakehouse · Databricks Unity CatalogUnity Catalog governs the whole data estate. Syntitan binds and reproduces the state behind a single AI run. Read together, they describe a stack rather than a fight. A catalog makes data discoverable. A governance suite makes it permitted. A platform makes it stored and fast. Each of those should stay where it is strong. The released state that a model ran on, captured so the run can be repeated, is the layer to add on top when the goal is a result that survives production. 05 When to reach for each If the problem is that nobody can find the right table or agree on what a field means, the answer is a data catalog. If the problem is access rules and audit, the answer is a governance suite. If the problem is storage, scale, or operational reads and writes for an application, the answer is a data platform such as Snowflake or Databricks, with Lakebase where an operational store is needed. If the problem is that the real data cannot be shared, a synthetic data tool fills the gap. If the problem is that a model worked in the proof of concept and then drifts in production, the failure usually called model drift, and nobody can reconstruct the data state that worked, none of the above closes it on its own. That is the layer Syntitan is built for, and it runs alongside the rest of the stack rather than replacing any of it. 06 The point of the word Many platforms now claim AI readiness, and most of them have earned the claim for the reading they answer. The question a team should ask is which reading they need. Ready to find, ready to govern, ready to query, or ready to reproduce an AI run. The first three are well covered. The fourth decides whether the work holds, and it is worth knowing where your data stands on it before the next model goes live. Read next Collibra vs Syntitan: data governance and reproducible AI runs Snowflake Horizon vs Syntitan: semantic meaning and reproducible AI runs Databricks vs Syntitan: governing the estate, reproducing the run About this piece. CUBIG builds the AI-ready data layer between enterprise data and the models and agents that run on it. Syntitan is the product. Capability descriptions reflect each platform's published and shipping focus as of 2026 and are meant to map categories, not to rank quality. --- title: "What is DTS? The AI-Ready Data Transformation Engine" url: "https://cubig.ai/articles/what-is-dts/" source: live --- DTS is CUBIG's AI-ready data transformation engine: it rebuilds restricted, scarce, or structurally unusable enterprise data into data a model can learn on, while preserving the original's structure, statistics, and rare patterns and never exposing the original itself. It works in three moves: diagnose, transform, rebuild. The word that matters in that sentence is transformation. DTS does not invent data, and it is not a synthetic-data generator. It starts from data you already own but cannot put in front of a model, and it rebuilds that data into a form the model can actually learn from. That need is well documented: a review of machine learning for synthetic data generation notes that synthetic data becomes necessary precisely when real data is unavailable or must be kept private. The source stays where it is. What comes back is new, usable, and faithful to the original's shape. The balance it manages is measurable: NIST built its SDNist tool to score a synthetic dataset on both utility and privacy. The data you have, but can't use Most enterprises are not short on data. They are short on data a model is allowed to touch. Industry estimates put roughly 80% of enterprise data in the unstructured, rarely-used category, and much of that never reaches a model at all. The records exist. They sit in a system the training team is not permitted to draw from. Three things usually stand between the data you have and the data a model can learn on. The first is movement. The original cannot leave its environment for legal, contractual, or security reasons, so the team that needs it for training never receives a usable copy. The second is sensitivity. The records carry context you cannot expose, so they go untouched even by people cleared to see them, because routing them anywhere new carries a risk nobody wants to own. The third is scarcity, and teams tend to notice it last. The events a model most needs to recognize, such as fraud, equipment failure, or the rare clinical case, are exactly the ones that barely appear in the data. Train on what you have and the model learns the common case well and misses the case that matters. None of this is a data-quality problem in the ordinary sense. The data is not dirty or wrong. It is simply not in a state a model can learn from, and cleaning it does not change that. That gap between raw records and a trainable dataset is what the data blockers describe. What blocks the dataWhy a model can't use itWhat DTS doesIt can't moveThe original can't leave its environmentReads the structure indirectly and rebuilds without moving the raw recordsIt's too sensitiveExposing the records crosses a legal or security lineBounds re-identification with differential privacy; the original stays putRare patterns are missingThe events that matter barely appearAmplifies those patterns from the real distribution Transformation, not generation This is the line that separates DTS from a generator. A generator invents records from scratch. Generic or random rows can pass a human glance and still teach a model nothing, because they carry none of the relationships that made the original predictive. DTS keeps those relationships. It preserves the statistical structure, the correlations between fields, and the rare patterns that hold the most signal, so what you train on still behaves like the real thing. The output is rebuilt data. It is not fabricated, and it is not the original record either. That distinction earns its keep in practice. Because the result is derived from your data rather than copied out of it, you can move it across teams or to partners within limits the source records could never clear. And because it tracks the original's structure, a model trained on the rebuild performs close to one trained on the source. That principle is well established. In MIT's Synthetic Data Vault study, models built on data that preserved the correlations between fields matched or outperformed models built on the original in more than 70% of cases. DTS applies the same structure-preserving idea to the data you already hold, instead of generating new records from nothing. Diagnose, transform, rebuild DTS runs as a loop, not a one-time export. The three moves are what give the engine its name. Diagnose. Before changing anything, the engine reads what makes the data unusable for your specific target model: what is missing, what is skewed, which patterns are too thin to learn from, and where the sensitivity sits. A rebuild aimed at the wrong problem produces data that looks fine and trains badly, so the diagnosis comes first. Transform. Next it works on the parts that block learning. It fills the gaps, corrects the skew, and strengthens the rare patterns a model needs more examples of. Scarcity gets solved here: the rare event is amplified from the real distribution, not conjured out of nothing. Rebuild. The output is an AI-ready dataset you can check against the original's structure and statistics before you trust it. You are not asked to take the result on faith. You compare distributions, confirm the correlations held, and only then train. Is your data AI-ready? A quick diagnosis Before any rebuild, it helps to know which blocker you are dealing with. Five questions usually settle it: Movement. Can a copy of this data legally leave the environment it lives in? Sensitivity. Could the records go to an external model without crossing a legal or contractual line? Scarcity. Do the rare events you need the model to catch appear often enough to learn from? Quality. Are gaps, skew, or imbalance likely to bias what the model learns? Shareability. Can the result be shared across teams or partners under your regulations? A "no" to any of these means the data is not yet in a s How AI-ready data transformation works Three mechanisms make the rebuild possible. Each one answers a specific objection a security or compliance reviewer will raise, which is why they are worth stating plainly. The first is a non-access architecture. The original stays on its server. DTS reads only statistical information, and reads it indirectly, so the rebuild happens without the raw records moving anywhere. That is what lets the work begin at all in environments where copying data out is the one thing nobody will approve. The same principle, applied to running a model rather than rebuilding data, is the restorable AI data boundary. The second is differential privacy. Re-identification of any individual record is held to a measurable mathematical bound, rather than controlled by deleting fields until the data is hollow. Masking makes data safe by making it poorer. Differential privacy keeps the data useful while putting a provable limit on what can be traced back. The approach is proven at scale: the U.S. Census Bureau adopted differential privacy to protect respondents in the 2020 Census. The third is distribution calibration. The rebuilt data is tuned to follow the original closely enough that a model trained on it performs in line with one trained on the source. Any lift figure you see is representative until you reproduce it on your own model and your own data, and that reproduced number is the only one that should decide anything. Where DTS fits Once data is rebuilt, the work that used to be off-limits becomes routine. You train and fine-tune on data that was locked. You evaluate and test without touching the original. You share results across teams or with partners inside regulatory limits, because what you are sharing is no longer the source record. The blockers look different by industry. DTS rebuilds rare defect data for manufacturing quality models, reconstructs scarce and sensitive records for clinical research, and amplifies thin fraud signals for fraud and AML detection. DTS is the step that moves data from "managed" to genuinely AI-ready data. It pairs with sensitive AI workflow enablement, which clears the confidential-context problem so a workflow can run in the first place. DTS takes the harder case, where the data is not only sensitive but scarce or structurally unusable. Both are capabilities of the CUBIG Syntitan platform, so the transformation engine and the enablement layer share one boundary instead of bolting two tools together. Related: What is Sensitive AI Workflow Enablement? · What is AI-Ready Data? · Syntitan --- title: "What is Sensitive AI Workflow Enablement?" url: "https://cubig.ai/articles/what-is-sensitive-ai-workflow-enablement/" source: live --- Sensitive AI workflow enablement allows enterprises to run AI workflows on confidential data without exposing raw sensitive values to the model. Sensitive values are replaced while the structure and relationships required for the task remain intact. The model works on that substituted representation, and the output is reconstructed inside the enterprise’s controlled environment. It addresses a problem many enterprise teams encounter soon after a promising AI demo. The model can perform the task, but the data it needs cannot be exposed. Contracts, customer records, pricing logic, and internal reports contain exactly the context that makes AI useful. They also contain information that legal, security, or regulatory teams may not approve for use in an external AI service. The result is a familiar stall. The model is ready, but the workflow cannot move forward with real data. Across enterprises that stall is the norm rather than the exception: a 2025 MIT study reported by Fortune found roughly 95% of enterprise generative AI pilots reach no measurable business impact. Sensitive AI workflow enablement removes that barrier by separating confidential values from the structure and relationships the model needs to do the work. The rest of this article explains why the usual workarounds fall short and what changes when enterprises stop treating the values and the work as inseparable. Why it is needed Take one concrete case. An analyst has a renewal contract to review, and the clauses that matter most are the ones naming the counterparty, the negotiated pricing, and the carve-outs that took months to settle. A model could flag the risky terms in seconds. But those exact clauses are the ones the contract forbids her from pasting into an outside tool. The capability is right there. The input is locked. Faced with that, teams usually pick one of two paths. Both fail the same way. The first is to strip the data down before it goes anywhere. You redact the names, blank the numbers, cut the clauses that feel too specific, and send what is left. It feels responsible. The trouble is that you have also removed the context the model was supposed to reason about. A contract with the counterparty and the pricing taken out is no longer a contract; it is a form with the meaningful parts missing. The model reads around the gaps and returns something generic. The output comes back safe and useless, and the analyst is back to doing the review by hand. The second path is to keep everything inside and ban the external model outright. No data leaves, so the compliance question goes quiet. But now you have given up the model that made the project worth starting, and you have not actually stopped the work. People still have a deadline. They paste a "lightly edited" version into a personal account, or they route the task through a browser extension nobody reviewed, and the exposure you tried to prevent happens anyway, off the books and out of view. A ban does not remove the risk. It just moves it somewhere you cannot see it. This pattern shows up in the numbers. In Cisco's 2024 Data Privacy Benchmark Study, 27% of organizations had banned generative AI outright over privacy and data-security risks, while 48% admitted that staff had already entered non-public company information into those same tools. The instinct to shut the door is widespread: a 2023 BlackBerry survey found 75% of organizations were implementing or considering bans on ChatGPT and other generative AI apps at work. The prohibition and the exposure sit side by side, which is what happens when the work is urgent and the only approved answer is no. The reason both paths fail is that they share a hidden assumption: that the confidential values and the useful work are one inseparable object, so any move you make on one is a move on the other. Sensitive AI workflow enablement rejects that assumption. It separates the two. The model gets what it needs to do the job. The confidential values stay where they already live. Both paths take the fear seriously, and they are right to. Where they go wrong is in assuming you must choose between the values and the work. How sensitive AI workflow enablement works It runs in four steps, end to end: Substitute. Sensitive values are replaced while the document's structure stays intact: tables, lists, hierarchy, and the relationships between fields. The model sees a coherent task, not a redacted blank. Execute. An external or on-prem LLM, RAG pipeline, or agent runs on the substituted data. The original values stay inside your environment. Reconstruct. The result is restored to its original context, so it comes back as a usable business document rather than a placeholder someone has to reassemble by hand. Run in place. The whole flow executes inside the environment you already run, whether cloud, on-prem, or air-gapped, instead of routing data somewhere new. The step that decides everything is the first one, so it is worth slowing down on. Plain masking removes a value and leaves a hole. A name becomes a black bar. A price becomes a row of X's. A clause about a specific party becomes a gap the model has to guess around. The data is now safe and also incoherent. Structure-preserving substitution does something different. It swaps the sensitive value for a stand-in that keeps the shape of the original. A company name becomes a different but consistent company-shaped stand-in, used the same way every time it appears. A figure becomes another figure that holds its place in the table and its ratio to the numbers around it. The indentation of a clause, the row-and-column logic of a spreadsheet, the order of steps in an operating note: all of it survives. What leaves is the identifying content. What stays is everything the model reasons from. That is why the model reads a real task instead of a page full of holes, and it is the difference the rest of the mechanism depends on. The full walkthrough lives in substitute, execute, reconstruct. Execution is the step people worry about and the step where the least changes. You run the model you already chose, on the substituted version, and from the model's point of view the task looks complete because the structure is all there. Then reconstruction maps the result back: the stand-ins resolve to the real values in their real context, so what returns is the analyst's own contract, reviewed, not a draft keyed to stand-ins she would have to translate by hand. The result coming back inside the boundary, rather than leaving for good, is what defines a restorable AI data boundary. In practice these four steps are delivered by a Context-Preserving Data Layer for AI in CUBIG's case, LLM Capsule — that sits between your real data and the AI. The layer keeps the structure intact on the way out and rebuilds the business meaning on the way back. The mechanism is the point; the layer is just where it runs. What stays the same The reason this is framed as an enablement layer, and not a migration, is that almost nothing about your setup has to change. The model you wanted is the model you use. The environment you already run in is where it executes. The output lands in the format your team already works with, in the tool they already open. There is no new vendor model to certify, no data lake to relocate, no workflow to rebuild from scratch. One thing changes, and only one: raw values stop being the thing you ship out to get work done. Everything an enablement layer adds is in service of keeping the rest of the picture exactly as it was. How it differs from masking, DLP, or a gateway This is the distinction that matters most, and it is easy to blur because the tools sit near each other. A masking step, a data-loss-prevention filter, and an AI gateway are all controls. Their job is to stop something: to catch the value before it leaves, to block the request that should not go out. You measure a control by how much it prevents. A good one prevents a lot. Sensitive AI workflow enablement is not measured that way. It is an enablement capability, and you measure it by how much work it lets you run that was previously off-limits. The renewal contract that used to sit in the "cannot use AI on this" pile now goes through the workflow and comes back reviewed. That is the metric. Because the original values stay inside the whole time, security, privacy, and legal reviewers do gain a clean basis to approve the workflow, and that approval is real and valuable. But it is the by-product, not the goal. A control exists to say no safely. An enablement layer exists to make the yes possible. If you arrived assuming the answer was masking, the move from one mindset to the other is worth its own read: from PII masking to workflow enablement. If you are weighing this specifically against a data-loss-prevention control, the line-by-line differences are set out in LLM Capsule vs. AI DLP. What "sensitive" actually covers One last thing the practice gets right is the scope of the word "sensitive." Most tooling treats it as a synonym for personal data: names, identifiers, the fields a privacy regulation lists by name. Those matter, but they are not the whole problem. The contract clause that reveals a negotiated discount, the pricing model that encodes years of strategy, the internal metric that would tip off a competitor, the operating note that describes how a system actually fails: none of that is personal data, and all of it is the kind of thing an enterprise cannot afford to leak. Sensitive AI workflow enablement treats business-sensitive context as first-class, because in real workflows it is usually the larger share of what cannot go out. Where it fits Sensitive AI workflow enablement is the entry point to an AI-ready data pipeline. It clears the confidential-context problem first, so the rest of the work has something to operate on. When data is not only sensitive but also scarce or structurally unusable, the next step is DTS, the AI-ready data transformation engine. which rebuilds data that the model cannot learn from in its raw state. Both run on the CUBIG Syntitan platform, so the enablement layer and the transformation engine share one boundary rather than bolting together two separate tools. In the most restricted settings, such as an air-gapped defense network, the same flow runs with no outbound connection at all; running an LLM inside an air-gapped defense environment walks through that case. LLM Capsule is CUBIG's implementation of the layer described here. It runs substitute, execute, and reconstruct so your confidential values stay inside while the model you already chose does the work. --- title: "Agent-Ready Data Needs Semantic Context" url: "https://cubig.ai/articles/agent-ready-data-needs-semantic-context/" source: live --- By Bae Ho, Founder & CEO, CUBIG Corp. · Updated June 2026. Agent-ready data is enterprise data that carries the business meaning, access conditions, and freshness an autonomous agent needs to act on a value, not just text the agent can retrieve. A single model call can tolerate thin context, because a person reads the one answer it returns. An agent that chains ten steps cannot: it carries each misreading into the next decision, with no one in the loop to catch it. Enterprise AI adoption keeps stalling at the same seam: the step from a working demo to AI in production. The agentic wave is hitting that seam harder. Gartner predicts that over 40% of agentic AI projects will be canceled by the end of 2027, pointing to escalating costs, unclear business value, and inadequate risk controls. That sits on top of a problem Gartner had already flagged: it expects organizations to abandon 60% of AI projects through 2026 for want of AI-ready data, and its read of the failures is blunt: the data foundation is usually the issue, not the algorithm. An agent is only as reliable as its reading of the data, and it reads that data on every step of its run, usually with no one checking. This article sets out what semantic context is for an agent, why an agent raises the bar an ordinary model set, why retrieval and a bigger model do not close the gap, and how to tell whether your own data carries enough context for an agent to act on. Why an agent raises the bar a model set A model called once gets a single look at the data and returns one answer. A person reviews it. If the reading was off, the error stops there. An agent works through a loop instead. It reads the data, decides, calls a tool, reads what comes back, decides again, and keeps going, usually with no human between the steps. Every decision rests on how it read the values in front of it. When that reading is wrong because the meaning behind a value was missing, the agent does not stop to check. It acts on the wrong reading and feeds the result into its next step. Picture an agent reconciling invoices. One column holds an amount, and a handful of rows are blank because those invoices were never issued, not because the amount was zero. A human would know that. The agent reads the blanks as zeros, marks the accounts as paid in full, and moves on to issue refunds against the difference. A single misread at step one became a wrong financial action four steps later, and nothing in the chain flagged the moment it went off. Inaccuracy is already the consequence enterprises report most: McKinsey found that 51% of organizations using AI have hit at least one negative consequence, with inaccuracy at the top of the list. An agent multiplies that one failure down its chain rather than containing it to a single output. What an agent needs to read: meaning, permission, freshness Semantic context is the meaning that makes a value safe to act on, not just readable. For an agent, it comes down to three things the data has to carry on its own, because no analyst is standing by to supply them. Business meaning. What a value represents and how it relates to everything around it: what status = 3 stands for, whether amount is in dollars or cents, why a field is blank (never collected, not applicable, or a true zero), that discharge_date must fall after admission_date, and that a negative balance is valid for one account type and impossible for another. A human analyst rebuilds most of this without noticing. An agent reads what is written, and only what is written. Access conditions. Whether the agent is permitted to use this value for this action. A person knows a flagged record is off-limits for an automated decision. An agent sees a value it can technically read and acts on it, unless the permission travels with the data and tells it not to. Freshness. Whether the value is current enough to act on. A model trained last quarter expects a stable snapshot, but an agent acting in real time on a number that went stale two days ago makes a confident, outdated call. Without a freshness signal on the data, the agent has no way to know the difference. Miss any one of the three and the agent still runs. It runs on a reading that meaning, permission, or freshness would have corrected. Where that context disappears Most enterprise data for AI moves through a bronze, silver, gold pipeline. Each promotion makes the data cleaner and more consistent, and along the way it strips out the context an agent depends on. The loss happens in ordinary steps, each defensible on its own: A blank is imputed to the column mean, so "never measured" and "average" collapse into the same number, and the agent can no longer separate a real reading from a filled gap. A field is generalized for compliance, turning a precise code into a broad band. The column clears its privacy check while the distinction the agent needed to route its next action is gone. A column is renamed or merged during a schema cleanup, and the relationship it held to another field is no longer recorded anywhere the agent can read. None of these is a mistake. Each is correct hygiene for a warehouse. The trouble is that what they remove is never written back into the data, so the table still looks complete while the meaning behind it has left. A model trained once might absorb that loss as a few points of accuracy. An agent compounds it across every step of its run. A single model callAn autonomous agentReads the dataOnceRepeatedly, across stepsHuman between stepsUsually, to review the outputUsually noneEffect of a misreadingContained in one answerCompounds into later actionsFills missing context fromThe person reviewing the resultNothing; reads only what is writtenNeeds meaning inside the dataHelpfulRequired Why retrieval and a bigger model don't close the gap The common fix for "the agent lacks context" is to bolt on retrieval: give the agent a vector store and let it pull in documents at run time. Retrieval-augmented generation (RAG) helps, and it is worth doing. It does not solve the problem this article describes. RAG improves what an agent can find. It says nothing about whether what it finds is current, permitted, or trusted by the business. A retriever will return a deprecated schema doc, a number from a table the agent is not cleared to act on, or a definition that was true two reorganizations ago. The text comes back. The conditions that make it safe to act on do not come back with it. A bigger model does not close the gap either. A larger model reasons better over the context it is given; it cannot supply meaning, permission, or freshness the data never carried. When the underlying value is ambiguous or stale, more reasoning produces a more confident wrong answer, not a right one. This is the part that survives better models and better retrieval: the conditions have to be attached to the data itself, upstream of both. A quick check: is your data agent-ready? Before you hand a workflow to an agent, run these on the tables it will read and act on: Can the agent tell, from the data alone, why a given field is blank, with no one there to explain it? Do units, codes, and field meanings travel with the table, or do they sit in a separate doc the agent never opens? Does the data record whether the agent is permitted to act on a given value, or only whether it can read it? Is there a freshness signal on the table, so the agent knows when a value is too old to act on? When the agent acts on a value, can you reproduce the exact data state it read, to judge whether its reading was right? Do you have a number for data readiness across all six axes, or only a "the pipeline is green" feeling? If the answers run to "no," the limit on the agent is the context missing from the data it reads, not the quality of its reasoning. Context is one of the six readiness axes Preparing data for AI agents starts with measurement; "ready" only means something if you can measure it. Syntitan scores enterprise data on six readiness axes, and Context is the one this article is about: does the meaning, permission, and freshness a model or agent needs travel with the data? Clean data often scores well on Integrity and Usability while scoring low on Context, which is the exact profile that passes review and then misleads an agent in production. The full set, including Reproducibility and Traceability, is unpacked in What Is AI-Ready Data?. For an agent, two of those axes work as a pair. Context lets the agent read the data correctly on any single run. Reproducibility lets you return to the exact state it read when a decision it made gets questioned, so you can see whether the data or the agent was at fault. That second axis carries more weight as trust thins: in Stack Overflow's 2025 survey, more developers distrust the accuracy of AI output (46%) than trust it (33%). An agent that cannot show the data behind a decision leaves a skeptical team no way to check it, and the rollout stalls there. Where Syntitan fits Syntitan is the AI-Ready Data Platform that closes this gap. The front of the workflow diagnoses the six axes and rebuilds what blocks execution while preserving the structure and context that cleaning tends to drop, so the meaning an agent needs stays attached to the data. The back fixes that data as a reproducible state, so every agent run binds to something you can diff and return to when you need to explain a decision it made. Making data ready and keeping it reproducible is the job of an AI-ready data operating layer: the missing layer between data management and AI execution. Any performance figure you see is representative until you reproduce it. Validated results come from your own agents, on your own data, in your own environment, not from a benchmark slide. --- title: "AI-Ready Data vs Clean Data: Why Clean Isn't Enough" url: "https://cubig.ai/articles/ai-ready-data-vs-clean-data/" source: live --- Clean data passes type, null, duplicate, and compliance checks. AI-ready data clears one more bar: a model can still learn from what's left, and you can reproduce the exact data behind any result. Clean is a hygiene standard. AI-ready is an execution standard. The two get used as if they mean the same thing, and that confusion is where most enterprise AI pilots stall. "Our data is clean" is usually true. It is also a common reason the AI still fails. The checks that make data clean and the properties that make it ready are different properties, and a dataset can have the first while lacking the second. This article sets out where they diverge, why the gap stays invisible until production, and how to tell which side of the line your own data is on. What "clean" actually certifies Cleaning earns its place. Bad data is expensive. IBM's widely cited estimate put the cost of poor data quality in the U.S. at $3.1 trillion a year, and clearing that mess is what cleaning addresses. So clean is a real and useful standard. When a data team says a dataset is clean, they mean it has passed a set of checks: Types are correct, so a date is a date and a number is a number. Nulls are handled, either dropped or filled, so nothing downstream breaks on a missing value. Duplicates are removed, so a record is counted once. Compliance is satisfied, so identifiers are masked or generalized to clear privacy review. Pass all four and the data is clean. It will load, it will join, it will populate a dashboard, and it will survive an audit of its handling. None of that is in dispute. The trouble is that every one of those checks is about the form of the data, not about whether the signal a model needs survived the process. What cleaning quietly removes Most enterprise data moves through a bronze, silver, gold pipeline. Each promotion makes the data cleaner and more consistent, and each promotion also strips out information a model was relying on. The loss happens in ordinary steps, each defensible on its own: A blank row is dropped or imputed. But a blank often carries meaning. A missing lab value can mean the test was never ordered, which is itself a signal. Drop the row and the model learns nothing was there; impute it to the column mean and the model learns something false. The "reason it was blank" leaves with the row, and nothing records that it left. A field is generalized for compliance. Age becomes a ten-year band; a date becomes a quarter. Each column passes its privacy check in isolation, while the relationship between columns that the model was learning from is broken. A column is normalized to a clean range. It becomes comparable across sources and loses its distribution. A bimodal column that told the model two distinct populations existed flattens into a smooth ramp. None of these is a mistake. Each is correct hygiene. The problem is that what they remove is never written down. The data clears every check and reaches the model stripped of part of what made it useful, and the subtraction stays invisible until a model trained on the residue underperforms in production. Clean is hygiene. Ready is whether a model can learn. This is the core distinction. Clean asks: did the data pass its checks? AI-ready asks: can a model still learn from what's left, and can you reproduce the state it ran on? A dataset can be spotless on the first question and unfit on the second. Clean dataAI-ready dataStandardHygieneExecutionAsksDid it pass its checks?Can a model learn from it, reproducibly?Optimized forDashboards, reporting, audit of handlingModels, LLMs, and agents in productionTreats a blank asA defect to remove or fillPossible signal to preserve and labelCares about distribution?No, tidy categories are fineYes, the full distribution carries the signalCares about reproducibility?No, the latest clean version is enoughYes, every run binds to a fixed, returnable state A dashboard wants tidy categories and suppressed outliers. A model wants the full distribution, including the parts a human would dismiss as noise. So the same gold table that is ideal for a report can be the wrong input for a model, even though it is, by every check, clean. The cost of getting this wrong is not theoretical. Gartner expects organizations to abandon 60% of AI projects through 2026 for want of AI-ready data, and reports that 63% of organizations either lack the data management practices AI needs or are unsure they have them. McKinsey finds that 51% of organizations using AI have already hit at least one negative consequence, with inaccuracy the most common. The mistrust runs downstream to the people building on the output: in Stack Overflow's 2025 survey, more developers distrust the accuracy of AI answers (46%) than trust it (33%). Gartner's own reading of the failures is blunt: the problem is usually the data foundation, not the algorithm. The part "clean" never covers: a state you can return to There is a second half to readiness that hygiene does not touch at all. Clean is a property of a dataset as it is right now. Production is not one moment. The same pipeline, running on data that has shifted underneath it, produces different results next month, and "the data is clean" says nothing about why. When a model that worked last month drifts this month, the team needs to answer one question fast: what changed between the run that worked and the run that didn't. Cleanliness cannot answer it. Only a fixed, reproducible state can. AI-ready data, in this sense, is data sealed to a state you can return to, diff against, and reproduce, so when output is questioned you can rebuild the exact rows, schema, and preprocessing the result came from.Versioning the model is not the same thing: the model version tells you which code ran, not which data state produced the result. This is why clean is necessary but not sufficient. Clean gets the data through its checks today. Ready keeps the result defensible after it ships. How to tell the difference on your own data A quick check before the next planning meeting: In your gold tables, can you still say why a given field is blank? If not, the context is already gone, even though the table is clean. When a model drifts, can you diff the data state to see what moved? If not, you are debugging blind against a moving target. Can you take any past result and reproduce the exact data it ran on? In a regulated setting, if you cannot, you may not be allowed to use that result at all. Do you have a number for readiness across all six axes, or only a "the pipeline is green" feeling? If the answers run to "no," your constraint is not model quality. It is the state of the data reaching the model, and that is a different problem with a different fix. How AI-ready data gets measured "Ready" only means something if you can measure it. Syntitan scores enterprise data on six readiness axes, which turns AI-ready from a claim into a number: Usability, Integrity, Context, Consistency, Reproducibility, and Traceability. Clean data tends to score well on Integrity and parts of Usability while scoring low on Context, Reproducibility, and Traceability. That is the profile of data that passes review and then fails in production. Each axis is unpacked in What Is AI-Ready Data?. Where Syntitan fits Syntitan is the AI-Ready Data Platform that handles both halves of the gap this article describes. The front of the workflow makes data AI-ready, diagnosing the six axes and rebuilding what blocks execution while preserving structure and the context that cleaning tends to drop. The back fixes that data as a reproducible state, so every AI or agent run binds to something you can diff and return to. That arc, make it ready and keep it reproducible, is the job of an AI-ready data operating layer: the missing layer between data management and AI execution. Any performance figure you see is representative until you reproduce it. Validated lift comes from your own model, on your own data, in your own environment, not from a benchmark slide. Related: PHI Masking vs Synthetic Data · What Is DTS? · What Is Sensitive AI Workflow Enablement? --- title: "What Is AI-Ready Data? (And Why Clean Data Isn't Enough)" url: "https://cubig.ai/articles/what-is-ai-ready-data/" source: live --- AI-ready data is enterprise data that scores well across six readiness axes, carries the semantic context a model needs, and is fixed to a state every AI run can be traced back to. Clean data clears type, null, and compliance checks. AI-ready data clears one more bar: a model can still learn from it, and you can reproduce the exact data behind any result. That extra bar is where most enterprise AI stalls. Gartner expects organizations to abandon 60% of AI projects through 2026 for want of AI-ready data, and reports that 63% of organizations either lack the data management practices AI needs or are unsure they have them. Only 37% are confident in those practices. Gartner's own reading of the failures is blunt: the problem is usually the data foundation, not the algorithm. This article sets out what "ready" means in practice, why clean is a lower bar than ready, and what has to be true before a model, an LLM, or an agent can run on enterprise data without breaking. The same gap, felt two ways A vertical-AI vendor feels it as lost revenue. The demo ran on sample data, the customer's real data broke it, and the rollout sat in "pilot" for two more quarters. A data or ML engineer feels it as a technical dead end. The pipeline is green, the warehouse is governed, the dataset is clean, and the model still underperforms. Or it performed last month and slid this month, with nothing in the stack able to say what moved. Both describe the same missing layer: the one between clean and ready, and the one that keeps a run reproducible after it ships. The cost of getting it wrong is not theoretical. McKinsey found that 51% of organizations using AI have hit at least one negative consequence, with inaccuracy the most common. Developers feel it too: in Stack Overflow's 2025 survey, more developers distrust the accuracy of AI output (46%) than trust it (33%). Clean is hygiene. Ready is whether a model can still learn. Most enterprise data moves through a bronze, silver, gold (medallion) pipeline. Each promotion makes the data cleaner, more consistent, more compliant. Each promotion also removes signal a model was relying on. By the time data reaches gold it suits a dashboard, which wants tidy categories and suppressed outliers, far better than a model, which wants the full distribution including the parts a human would dismiss as noise. The loss happens in ordinary steps, each defensible on its own: A blank row is dropped or imputed to the column mean. But a blank often carries meaning. A missing lab value can mean the test was never ordered, which is itself a signal. Drop it and the model learns nothing was there; impute it and the model learns something false. A field is generalized for compliance: age into a ten-year band, a date into a quarter. Each column passes its privacy check, while the relationship between columns that the model was learning from is gone. A column is normalized to a clean range. It becomes comparable and loses its distribution. A bimodal column that told the model two populations existed flattens into a smooth ramp. None of these is a mistake. The trouble is that what they remove is never written down. It disappears, and the disappearance stays invisible until a model trained on the residue fails in production. We go deeper on this in AI-Ready Data vs Clean Data. For now the point is narrow: "the data is clean" and "the data is ready" are different claims. The six readiness axes "Ready" only means something if you can measure it. Syntitan scores enterprise data on six readiness axes, which turns AI-ready from a claim into a number: AxisWhat it asksUsabilityCan a model use the data as it is, in a form that runs, without prep that blocks execution?IntegrityAre values, types, and the relationships between fields consistent and unbroken?ContextDoes the semantic context a model needs, such as why a field is blank, travel with the data?ConsistencyDoes the data behave the same across runs, time, and environments, with no hidden drift?ReproducibilityCan any past result be rebuilt from the exact data state it ran on?TraceabilityCan every value's origin and transformations be verified, and does each run trace back to a fixed state? One low axis is usually enough to block a deployment, which is why a score across all six is more honest than a single quality metric. Each axis is unpacked in The Six Readiness Axes. Why it surfaces only in production A model that passed every check in the lab can still degrade the week after launch, and the reason is rarely the model. Production data differs from the sample the PoC ran on. Schemas shift when an upstream team renames a field. Preprocessing changes when someone updates a default. The data window moves as fresh records arrive and old ones age out. Each change is small. Together they mean the data state behind today's run is not the state behind last month's, and the model's output moves with it. When that happens, the team needs to answer one question fast: what changed between the run that worked and the run that didn't. If the data state behind each run was never fixed, that question has no answer, and the investigation turns into guesswork against a moving target. This is the failure pattern we trace in Why AI Fails After Deployment. A score is the start, not the finish A readiness score tells you the data is ready today. Production is not one day. The same pipeline, running on data that has shifted underneath it, produces different results next month, and a score alone cannot tell you why. So readiness has a second half: the data has to be fixed to a reproducible AI-ready state. Under the umbrella of a Verifiable Data State, that comes down to four operations: Release State seals the exact data state a run used, Run Binding ties each AI or agent run to that state, Diff compares two states to narrow down what changed, and Reproduce returns to the state behind any past result. A score that no fixed state stands behind tells you little the day output drifts. We make that case in A "Ready Score" Stops Too Early. For production AI, the question is not only which model ran. It is which data state and execution conditions produced the result. Syntitan scores enterprise data on six axes, fixes what blocks execution, and binds every AI or agent run to a data state you can diff and reproduce. Where Syntitan fits Syntitan is the AI-Ready Data Platform that handles both halves. The front of the workflow makes data AI-ready, diagnosing the six axes and rebuilding what blocks execution while preserving structure and context. The back fixes that data as a reproducible state, so every run binds to something you can diff and return to. That arc, make it ready and keep it reproducible, is the job of an AI-ready data operating layer: the missing layer between data management and AI execution. Any performance figure you see is representative until you reproduce it. Validated lift comes from your own model, on your own data, in your own environment, not from a benchmark slide. How to tell if your data is AI-ready A quick check before the next planning meeting: In your gold tables, can you still say why a given field is blank? If not, the context is already gone. When a model drifts, can you diff the data state to see what changed? If not, you are debugging blind. Can you take any past result and reproduce the exact data it ran on? In a regulated setting, if you cannot, you may not be allowed to use that result at all. Do you have a number for readiness across all six axes, or only a "the pipeline is green" feeling? If the answers run to "no," the constraint on your AI is the data state reaching the model, not the model. Related: AI-Ready Data: Ready for What? · AI Readiness Assessment: The Six Axes · Agent-Ready Data Needs Semantic Context --- title: "PHI Masking vs Synthetic Data: Which Does Your Healthcare AI Actually Need?" url: "https://cubig.ai/articles/phi-masking-vs-synthetic-data/" source: live --- PHI masking means altering or removing the identifiers inside protected health information, such as names, dates, and record numbers, so the data can no longer be readily traced to a specific patient. That is the definition most teams arrive with. It is also where the trouble starts. Strip the obvious identifiers and you would expect the patient to vanish from the dataset. The research disagrees. A 2019 study in Nature Communications estimated that 99.98% of Americans could be correctly re-identified in any dataset using just 15 demographic attributes. Two decades earlier, Latanya Sweeney had already shown that 87% of the U.S. population was likely unique on nothing more than five-digit ZIP code, gender, and date of birth. Masking removes the name. It does not remove the person. So the question in the title is the right one to ask before you spend a quarter building a pipeline. If you came in certain you needed PHI masking, it is worth knowing what masking actually buys you, where it runs out, and whether synthetic data solves a different problem than the one you have. What PHI masking actually does (and where it stops) In U.S. healthcare, "masking" usually means one of the two de-identification methods written into the HIPAA Privacy Rule. The governing regulation, 45 CFR 164.514, sets out exactly two ways to satisfy the standard, and the HHS Office for Civil Rights names them plainly: Expert Determination and Safe Harbor. Safe Harbor. You remove 18 specified identifiers, labeled (A) through (R) in the rule: names, all geographic units smaller than a state, every date element finer than a year, phone and fax numbers, email, Social Security and medical record numbers, account and license numbers, device and vehicle IDs, URLs and IP addresses, biometrics, full-face photos, and any other unique code. Clear the list and the data is treated as de-identified. Expert Determination. A qualified statistician analyzes the dataset and certifies that the risk of re-identification is "very small." Read that threshold again. The regulation does not say zero. It says very small, "alone or in combination with other reasonably available information." The law itself assumes someone might still link a record back to a person; it just asks that the odds be remote. This is the part most teams miss. Masked data is real patient data with the labels filed off. The records still describe actual people, which is the whole reason the data is useful and also the reason the link never fully disappears. NIST puts it cleanly in NISTIR 8053: "As long as any utility remains in the data derived from personal information, there also exists the possibility, however remote, that some information might be linked back to the original individuals." The same report calls the goals of de-identification and data utility "antagonistic." Push harder on privacy and you lose analytic value; keep the value and you keep some residual risk. None of this means masking is broken. Done to a real standard, it works well. A systematic review of re-identification attacks on health data by El Emam and colleagues found that across studies an average of 26% of records were re-identified, and yet for the one attack on data de-identified to an established standard, the rate fell to 0.013%. Standards-based de-identification dramatically lowers exposure. What it cannot do is change the nature of the asset: a masked record is still a token that, under the wrong conditions, points at someone. What synthetic data is, and why it answers a different question Synthetic data is data generated by a model to reproduce the statistical structure of a real dataset without copying any real record. NIST catalogs the technique in its CSRC glossary and in SP 800-188. The distinction that matters for healthcare is not "fake versus real." It is the mapping. A masked record corresponds one-to-one to a patient who walked into a clinic. A well-made synthetic record corresponds to no one. It carries the distributions, correlations, and edge cases of the source population, but there is no individual at the other end of the row to re-identify. That shift is what makes synthetic data a structural answer rather than a stronger lock. You are not reducing the chance that a real person is exposed; you are generating records where no specific real person is present to begin with. When the generation is governed by differential privacy, the guarantee becomes mathematical. NIST defines it in SP 800-226 (2025) as "a mathematical framework that quantifies privacy loss to entities when their data appears in a dataset," building on the formal definition introduced by Dwork, McSherry, Nissim, and Smith in 2006. Instead of arguing about whether 18 fields were enough, you can bound, with a tunable parameter, how much any single individual could have influenced the output. The fair objection is utility. If the records are invented, do the analyses still hold? In a peer-reviewed comparison published in JMIR Medical Informatics, synthetic patient data reproduced the results of five real observational studies; across 1,000 synthetic iterations the estimate biases stayed small, on the order of −1.3% to 1.9%, and within the 95% confidence limits of the real-data results. Synthetic data is not automatically faithful. Built and validated properly, it can be faithful enough to stand in for the original. PHI masking vs synthetic data: a side-by-side DimensionPHI masking / de-identificationSynthetic dataWhat it producesReal patient records with identifiers removed or alteredModel-generated records with no one-to-one link to a real patientResidual re-identification risk"Very small," never zero; rises as auxiliary data grows (45 CFR 164.514; NISTIR 8053)No source individual in the record; bounded mathematically under differential privacyRegulatory footingMature: explicit HIPAA Safe Harbor and Expert Determination methodsEmerging but real: used by the U.S. Census Bureau, NHS England, and accepted in FDA real-world-evidence guidanceBest forAudits, billing, operational reporting where the actual individuals must be preservedModel training, sharing across boundaries, augmenting rare cases, dev/test environmentsMain limitationUtility falls as privacy rises; link to real people persistsQuality depends entirely on generation and validation; can miss what it was not built to preserve So do you actually need synthetic data? Sometimes masking is the correct and sufficient choice. If your use case requires the real individuals to remain in the data, a clinical audit, a billing reconciliation, a regulatory report tied to actual patients, then you are not looking for synthetic records at all. You need de-identification done to a defensible standard, with the Expert Determination documented. Synthetic data earns its place when the value is in the patterns, not the people. Four signals tend to point that way: You are training or testing models and need volume, including rare events that masked data is too thin to cover. The data has to cross a boundary, to a vendor, a research partner, a cloud region, where moving real records is the bottleneck. You want developers building before they ever touch live PHI. NHS England released SynAE, a synthetic Accident & Emergency dataset, for exactly this: let teams build and test against realistic data first. Your re-identification exposure is already keeping a deal or a launch stuck, and "very small risk" is not clearing legal review. The regulators have started to move with this logic. The U.S. Census Bureau protected the 2020 Census with a differentially private system, its TopDown Algorithm, the largest deployment of differential privacy in a national statistical product. In December 2025 the FDA removed a barrier to using de-identified real-world data from electronic health records, claims, and registries in certain device submissions. And a 2025 comment in The Lancet Digital Health argued for accelerating synthetic-data privacy frameworks for medical research. The direction of travel is clear, even where the standards are still forming. The question underneath the question: what state does your data need to be in? Here is the reframe that saves teams a wasted quarter. Masking and synthesis are both techniques. Neither is the goal. The goal is data your AI workflow can actually run on, share, and answer for later. Picking "masking" or "synthetic data" up front is choosing a tool before you have named the job. The job, in regulated work, has a second half that both techniques tend to ignore. Suppose you generate a synthetic cohort, train a model, and ship it. Six months later an auditor asks which dataset produced which result, and whether you can regenerate it. Can you point to the exact release, the parameters, the lineage back to the source? A synthetic file with no record of how it was made is not audit-ready, however private it is. Privacy without traceability is half an answer. This is the layer CUBIG works on. AI-ready data is not just data that is safe to use; it is data that is usable, has its structure and context intact, and stays reproducible and traceable run to run. DTS, the AI-ready data transformation engine inside Syntitan, turns locked or restricted health data into that state. It preserves the statistical structure, patterns, and correlations of the source, uses differential privacy to control re-identification mathematically, and calibrates the output so models trained on it perform close to models trained on the original. The output can be synthetic; the point is the state, not the label. Syntitan then binds each result to a release you can diff and reproduce, so the privacy decision and the audit trail live in the same place. Put differently: synthetic data may well be the right technique for your project. But "do I need synthetic data?" is the smaller question. "Can I get my restricted data into an AI-ready state I can run on, share, and reproduce on demand?" is the one that decides whether the project ships. A 60-second self-diagnosis Run your use case through these. More boxes on the right means synthetic data, or a transformation that produces it, is the stronger fit. Do you need the actual individuals preserved (audit, billing), or only the patterns (training, analytics)? Does the data have to leave a trusted environment to be useful? Is "very small re-identification risk" still failing legal or partner review? Do you need rare cases or volume the real data cannot supply? Will someone later ask you to reproduce exactly which data produced which result? If the last box is checked, no masking or synthesis technique alone is enough. You need the data in a state that is private and reproducible at once. Bring one restricted dataset and let Syntitan transform it into an AI-ready state: structure preserved, re-identification controlled with differential privacy, and every result bound to a release you can reproduce and audit later. Want proof before you commit? Run a sample and check the utility against your own benchmark before any original data leaves your environment. --- title: "AI Readiness" url: "https://cubig.ai/glossary/ai-readiness/" source: live --- AI readiness is the degree to which an organization's data, systems, and teams are prepared to run AI reliably in production, not just in a pilot. It spans whether data is usable and accessible, whether infrastructure can serve models, and whether results can be trusted and reproduced. Readiness is often framed as a maturity stage you reach once. In practice the conditions that made a system ready can drift: data shifts, access changes, and a result that held in testing breaks in production. The part that gets overlooked is reproducibility. Data is only AI-ready if the exact state behind a result can be restored and re-run, so readiness holds up the next time the system executes rather than only on the day it was assessed. A working definition separates three layers. Core readiness is the general state of the data: usable, consistent, traceable. Target fit is whether that data works for one specific task and model, shown by a comparison run under fixed conditions. Operating evidence is the record that the qualified state was the one actually used, and that the team checked it again when something changed. A dataset can pass the first layer and still fail the second, which is why a general readiness score alone does not settle whether a given AI project is ready. --- title: "AI-Ready Data" url: "https://cubig.ai/glossary/ai-ready-data/" source: live --- AI-ready data is data that has been put into a state an AI system can actually use, trace, and reproduce, not just stored or cleaned. It goes beyond accuracy: the data carries the structure, context, and lineage a model needs, the access conditions that let it run under real constraints, and a record of the exact state behind each result. The distinction matters because most enterprise data is not in this state. It can be well stored and broadly accurate yet still scattered, missing context, or impossible to reproduce after it changes. That gap is where AI projects stall, on the data rather than the model. Practically, AI-ready data is assessed across dimensions such as usability, integrity, context, consistency, reproducibility, and traceability. The first cover whether the data can be used at all; reproducibility and traceability are what keep a result holding once a model is in production. For a fuller explanation of how these dimensions work and why clean data alone is not enough, see What Is AI-Ready Data?. --- title: "Concept Drift" url: "https://cubig.ai/glossary/concept-drift/" source: live --- Concept Drift is when the relationship between a model's inputs and the target it predicts changes over time. The model code stays the same, but what the inputs mean for the outcome shifts, so a model trained on past patterns gradually loses accuracy. It differs from data drift, which is a change in the input data's distribution rather than in the input-to-target relationship. Both are easier to catch when each run is bound to a fixed, released data state you can compare against. --- title: "Context-Preserving Data Layer" url: "https://cubig.ai/glossary/context-preserving-data-layer/" source: live --- A context-preserving data layer for AI is the layer that lets an AI model work on sensitive operational data that cannot leave the environment as-is. The original values stay inside; the model receives DP-based, context-preserving substitutes that keep the structure, relationships, and meaning it needs to reason, and the usable result is reconstructed inside the environment through a protected mapping layer. The point is that the data stays usable. Masking, redaction, and DLP keep the input safe but strip the context a model needs, so the output is hard to use. A context-preserving data layer keeps the work intact: the model runs on a protected working version, and the answer comes back in its original business form. The result is reconstructed through a deterministic internal mapping, not by reversing differential privacy, so it is not a claim of perfect recovery of every value. The original values and the reconstruction mapping never leave the customer's environment. --- title: "Data Readiness" url: "https://cubig.ai/glossary/data-readiness/" source: live --- Data readiness is how prepared a dataset is to be used reliably by AI models or agents, across usability, integrity, context, consistency, reproducibility, and traceability. Data can exist in volume and still not be ready: it may be restricted by compliance, missing the context AI needs, imbalanced, or impossible to trace when results change. Assessing data readiness for AI before a project starts shows the specific gaps blocking model or agent use, so teams fix the data state instead of discovering the problem weeks into cleaning. --- title: "Data versioning" url: "https://cubig.ai/glossary/data-versioning/" source: live --- Data versioning captures the exact state of a dataset at a point in time and labels it, so the same data can be retrieved or rebuilt later. It works the way code versioning does, except the thing under version control is the data itself: its rows, schema, and distributions as they stood when a result was produced. A simple snapshot keeps a copy and stops there. Versioning goes further by binding each model run to a specific, comparable state, which lets a team diff two versions to see what moved and return to the exact one a past result was built on. For AI this is what makes a result reproducible after the underlying data has shifted in production. --- title: "Execution State Layer" url: "https://cubig.ai/glossary/execution-state-layer/" source: live --- Execution State Layer refers to the operational layer that fixes data into a known state before it is used in AI execution. It connects release state, run binding, state comparison, and reproducibility to support more stable and explainable AI operations. Related terms: Execution State · State Diff / State Comparison · Run Binding · Verifiable Data State --- title: "Re-run / Replayability" url: "https://cubig.ai/glossary/re-run-replayability/" source: live --- Re-run or Replayability refers to the ability to execute an AI workflow again under the same or restored conditions. It is a core requirement for validation, incident review, and operational trust. --- title: "Release State" url: "https://cubig.ai/glossary/release-state/" source: live --- Release State refers to an immutable, identifiable data state fixed for a given AI execution. The same release state always points to the same condition, which makes it the reference point for reproducing a result or comparing two of them. Operational data keeps moving. Schemas get adjusted, field definitions change, policies are updated. Without a fixed release state, a result approved last month can return a different answer today with no way to say what changed. This sits at a different layer from model versioning. Model versioning fixes which model ran; release state fixes the condition of the data that model read. Reproducibility needs both. Each execution is tied to a release state through run binding, and differences between states are compared with a diff. --- title: "Reproducibility" url: "https://cubig.ai/glossary/reproducibility/" source: live --- Reproducibility is the ability to obtain the same result when an experiment or computation is repeated under the same conditions. In machine learning it means a model run can be repeated and produce a consistent, explainable outcome rather than a different one each time. In production the hardest part is rarely the model code; it is the data. Reproducing a result means restoring the exact data state the model ran against, its schema, distributions, and transformations, not just rerunning the same weights on whatever data is current. Without that, an AI result cannot be attributed or audited. Reproducibility is the precondition for measuring impact and for explaining why a system behaved the way it did, which is why it is treated as a core property of AI-ready data rather than a nice-to-have. Related terms: Verifiable Data State · Data Reproducibility · Reproducible AI Execution · Release State --- title: "Run Binding" url: "https://cubig.ai/glossary/run-binding/" source: live --- Run Binding refers to linking a specific AI run to a defined release state. It is the connection that makes each execution explainable in terms of the data conditions it ran under. When a run finishes, only the answer remains. Which schema, which field definitions, which applied policy produced it disappears unless it is bound at the time. What an audit asks for is not the result but the conditions behind it, and without that link the result cannot be defended. This differs in purpose from experiment tracking. Experiment tracking records hyperparameters and metrics so training runs can be compared. Run binding ties a production execution to a data state so it can be returned to under the same conditions. A bound run is reproduced through its release state, and if the state has moved, a diff shows what changed. --- title: "Silent Failure in AI" url: "https://cubig.ai/glossary/silent-failure-in-ai/" source: live --- Silent Failure in AI refers to a failure mode in which an AI system continues to run without an explicit error but produces degraded, incomplete, or incorrect outputs. Because the system appears normal, the issue may remain unnoticed until business impact becomes visible. --- title: "State Diff / State Comparison" url: "https://cubig.ai/glossary/state-diff-state-comparison/" source: live --- State Diff or State Comparison refers to the process of identifying meaningful differences between two data states or execution states. It helps teams determine which changes may have affected AI behavior or output. Related terms: State Versioning · Execution State Layer · Release State · Verifiable Data State