DCT

Turning undefined terms into AI data governance you can trust

We unify the sources, turn the documents into records, and settle what each business term means - so the next use case reuses the foundation instead of rebuilding it.

  • 80%

    less manual processing on document-bound work

  • Zero

    extra headcount as volume grew

  • One

    agreed definition per business term

  • Traceable

    every derived number, back to source

What the layer does

Most failed AI programmes are data failures wearing a model's clothes

Ask why an initiative stalled and the answer is almost always upstream of the model. The data existed but sat in five systems. It existed but carried three definitions of "customer". It existed but lived in documents nothing could read. Or it existed and nobody could show where a number came from, so nobody would sign the decision made from it.

Swipe to see more →
SOURCESoperational databasesdocumentsthird-party feedsevent streamsIngestwithout consolidating firstStructuredocuments become recordsValidatechecked at the boundaryDefineone meaning per termServetabular and retrievalGOVERNANCE TRAVELS WITH THE DATAclassification · access control · audit - attached here, not rebuilt in every application aboveCONSUMERSAgents & assistantsReportingApplicationsThe next onereuses this. It does not rebuild it.LINEAGE, END TO ENDsource → transform → rule applied → served → the number on the slideEvery derived figure can be walked back to where it came from andwhat was done to it - which is what makes someone willing to sign it.
Active stage

Ingest

Sources are read where they are, not copied into one system. Each one keeps its own owner and refresh timing.

01 - Against a real workload

Readiness is not a thing you can assess in the abstract

"Is our data AI-ready?" has no answer without naming what it has to be ready for. Ready for a retrieval assistant over policy documents means something entirely different from ready for a churn model. So the assessment maps the sources against the actual use cases they are meant to support, and reports the gap for those - not a maturity score.

  • Sources mapped to named use cases
  • Gaps reported per workload
  • Never an abstract programme
02 - Shared meaning

Three definitions of "customer" is not a data quality problem

It is a governance one, and no amount of pipeline work fixes it. When finance, sales and support each count "active" differently, two correct systems produce two different numbers and the meeting is spent reconciling instead of deciding. Definitions get agreed with the people who own them, written down once, and applied everywhere - so a dashboard and an agent answering the same question give the same answer.

  • Entities, definitions, grain, relationships
  • Agreed with the business, not inferred from schemas
  • One definition, applied in every consumer
03 - Caught at the boundary

A schema change should not surface as a broken dashboard

The usual failure is silent: an upstream team renames a field on Tuesday, nothing errors, and three weeks later someone notices a number looks wrong. Expectations about shape, type, nullability and volume are written down as a contract and checked where the data arrives, so a change is caught at the boundary with the producer named - before it has quietly propagated into every model downstream.

  • Expectations written, not assumed
  • Checked on arrival, with the producer named
  • Freshness and completeness monitored, not spot-checked
The same sources · A different thing built from them

A pipeline per project, and a foundation

  • Every use case starts with the same preparation, done again
    The second use case reuses what the first one paid for
  • Three definitions of the same business term
    One definition, applied in every consumer
  • Access granted by copying, so governance is lost on the first copy
    Classification and access travel with the data itself
  • A field renamed upstream surfaces as a wrong number weeks later
    Caught at the boundary, with the producer named
  • Nobody can say where a reported figure came from
    Walked back to source, transform by transform
  • Most of the useful information sits in documents nothing reads
    Structured, validated and queryable alongside the tables
  • The pilot worked on an extract someone cleaned by hand
    Built and tested on the data production actually has
  • Data quality noticed when a consumer complains
    Freshness and completeness monitored as a property of the layer
Claude Preferred Partner

Claude Partner Network

In the messy middle. Reading a supplier document whose layout changed without notice and pulling the right fields out of it. Proposing which of three columns across two systems are the same thing. Everything downstream of that - the rules, the definitions, the serving - stays deterministic, because a number you cannot reproduce is a number you cannot defend. DCT builds on Claude.

  • Documents into records

    Layouts that vary by supplier and change without warning, handled without a template per variation.

  • Matching across systems

    Candidate matches proposed where the same entity exists several times over. A person confirms the rule.

  • Deterministic downstream

    Validation, definitions and serving are code. The same input produces the same number, every run.

  • Cost per document

    Extraction cost is measured per document from the first batch, because volume is the whole point.

Customer stories

Where the fifth use case cost less than the first

  • Healthcare technology · Data platforms

    From successful pilots to something the organization could actually operate

    The ambition was clear: become an AI-led organization. Pilots had already shown potential, and that is precisely where most programmes stop - because scaling needs a trustworthy data foundation, clear ownership, engineering standards and a structured way to pick what to build next. In a healthcare technology environment, accountability and human oversight are conditions of operating at all, not preferences.

    DCT built the data foundation underneath - unified sources, consistent definitions, quality monitoring and governed serving - alongside the operating model for identifying and scaling use cases, with governance introduced at the start rather than retrofitted.

    • 2nd

      use case costs materially less than the first

    • Day one

      governance, not retrofitted later

    • In-house

      capability, not a permanent dependency

    The output was a capability, not a portfolio of projects. Organizations that build this first find their fifth use case cheaper than their first. Organizations that skip it find the opposite.

  • Travel · Media & data services

    Documents that nothing could read, turned into records everything could

    Information arrived as documents, emails, forms and supplier files, in layouts that varied by supplier and changed without notice. Skilled teams spent their days extracting it by hand, checking it and preparing it for downstream systems. The constraint was explicit and uncomfortable: more volume meant more people.

    DCT built extraction that reads across formats and layouts - including the supplier variations that break template matching - classified and validated the output against business rules before it reached anything downstream, and routed the structured result into the systems that consume it.

    • 80%

      less manual processing effort

    • Zero

      extra headcount as volume grew

    • Validated

      before it reached a downstream system

    Capacity became elastic and accuracy improved with volume. Before, the operating model had a hard ceiling - growth meant cost, and quality fell as throughput rose. Afterwards the business could take on more work without first deciding whether it could afford to.

Common questions

What data and engineering leaders ask first?

DCT Strata is delivered by DCT AI engineers working inside your team, against your own systems. These are the questions that come up before anyone signs anything.

How do you know if your data is ready for AI?

Data readiness for AI depends on the use case, not a single maturity score. Data adequate for a retrieval assistant over policy documents can be nowhere near adequate for a forecasting model. DCT Strata maps your sources against the specific use cases they need to support and reports the gap for those, which is more useful than one number covering everything.

Do you need a complete data foundation before starting AI initiatives?

No, a complete data foundation isn't required before starting an AI initiative. The stronger approach is building the foundation against a real use case, in the order that use case needs, so something ships early instead of a programme that stalls at month nine with nothing to show. DCT Strata leaves behind definitions, contracts, quality checks and lineage the next use case inherits, rather than a one-off pipeline that gets rebuilt each time.

Will I need to replace or consolidate my data warehouse before using DCT Strata?

No, DCT Strata doesn't require a data warehouse consolidation project first. Sources are integrated into an addressable data layer without requiring every system to be replaced or merged, which is fortunate, since consolidation projects are usually the ones that never finish. Where something genuinely needs replacing, that comes out of the assessment as a finding, with the cost of not doing it stated plainly.

What is data governance and why does it matter for AI projects?

Data governance is the set of rules that decide how data is defined, owned and controlled, and it matters for AI because inconsistent definitions produce answers that contradict each other. When finance calls a customer "active" one way and support calls it another, an AI system built on top inherits that disagreement. DCT Strata surfaces every definition in use rather than picking one, lets the owning teams agree the resolution, then applies it everywhere once.

Can AI work with unstructured documents instead of databases?

Yes, AI systems can work directly from unstructured documents when they're extracted and validated into structured records first. Contracts, reports, correspondence and supplier files get read, classified, extracted and validated, including formats where layout varies by sender. DCT Strata treats validation as more important than extraction: values are checked against business rules before anything downstream sees them, because a wrong number that arrives confidently is worse than a missing one.

How does DCT Strata prevent data quality from degrading after the engagement ends?

DCT Strata prevents data quality degradation by monitoring freshness, completeness and schema stability continuously, instead of relying on manual discipline. Expectations about each source are written down as a contract checked at arrival, so an upstream change is caught at the boundary with the producer named, instead of surfacing as a wrong figure weeks later. Every dataset also gets a named owner before it's served.

How does DCT Strata connect to the rest of the AI Suite?

DCT Strata sits underneath the rest of the AI Suite as the shared data foundation. Retrieval answers questions over this layer, process orchestration acts on it, and customer-facing work is only as good as what it knows. It's also reasonable to engage on DCT Strata alone, since plenty of organizations need the data foundation before they need anything standing on it.

How do you measure ROI on an AI data foundation project?

ROI on a data foundation is measured by the cost of the next use case, not the first one: the time from approval to usable data, tracked before and after. DCT Strata also tracks how many datasets have a named owner and a contract, freshness and completeness against agreed thresholds, and how many reported figures can be walked back to source without manual reconciliation.

Start with the use case, not the programme

0/255