Skip to main content
Data Products logo

Starter Services · Data Platform Accelerator

Your data platform probably does not need another rebuild.

Most platform projects begin by asking what to build. We begin by asking what not to rebuild. A diagnostic of your estate, then a foundation on Fabric, Databricks or Snowflake built from what survives it — kept, fixed and built, not rebuilt — in two to four weeks, with the pattern your team runs afterwards.

2 to 4 weeksDiagnostic firstFixed scope · fixed feeFabric · Databricks · Snowflake
Your estate, triaged; the foundation built from what survives Nine components of an existing data estate scatter. A diagnostic sweep passes and each is tagged keep, fix or build. They settle into four layers — sources, ingest, lakehouse, serve — and a foundation card appears: the platform decision, the first pipelines and the runbook. THE DIAGNOSTIC · KEEP / FIX / BUILD · ILLUSTRATIVE SOURCESINGESTLAKEHOUSESERVE ■ KEEP — works, stays■ FIX — stays, repaired▢ BUILD — missing, built ERP extractsKEEPCRM feedKEEPSpreadsheetsFIXNightly batch jobsFIXStreaming ingestBUILDWarehouse (3 yrs)FIXLakehouse layerBUILDBI semantic modelKEEPAccess + lineageBUILD THE FOUNDATION, HANDED OVER 3 kept · 3 fixed · 3 built — not 9 rebuilt Platform decided · 2 first pipelines live · runbook owned by your team Access and lineage from day one · the pattern for every pipeline after Keep, fix, build — the estate triaged Six components of a data estate are tagged keep, fix or build, and a foundation card appears: kept, fixed and built — not rebuilt. KEEP / FIX / BUILD · ILLUSTRATIVE ERP extractsKEEPSpreadsheetsFIXBatch jobsFIXStreamingBUILDWarehouseKEEPLakehouse layerBUILD 2 kept · 2 fixed · 2 builtNot 6 rebuilt. Pattern handed to your team.
Typical accelerator scope

Fixed in the proposal. The constraint is what makes two to four weeks real.

  • 1target platform — Fabric, Databricks or Snowflake
  • 1cloud environment, yours
  • ≤6source systems assessed in the diagnostic
  • 2representative pipelines built — typically one batch, one streaming
  • 1access and governance pattern, with lineage
  • 1agreed first workload the foundation is production-ready for
  • +runbook and handover to the team that will run it
Not includedEnterprise-wide migration · every source · every downstream workload · ongoing operation. Those follow the pattern this proves.

What does your estate look like?

Four ways a foundation fails to carry the work. The diagnostic finds which is yours.

The engagement is the same shape wherever you start. What changes is the keep-fix-build split the diagnostic produces — and therefore how much gets built.

A warehouse nobody trusts

The platform exists. Three teams get three answers to the same question, so nobody builds on it.

The situationThe warehouse is usually fine. What is broken is around it: ingest that fails silently, no lineage, definitions that drifted apart.
What we do differentlyDiagnose before deciding. We trace the numbers from source to dashboard and find where they break — then tell you what to keep, which is often most of it.
What you leave withA foundation people can build on because the path from source to answer is known, repaired and recorded.
Where the numbers breakOne revenue figure leaves the source system. It passes through nightly batch jobs that fail silently and a warehouse with no lineage, and arrives on three dashboards as three different numbers. The diagnostic finds the break at the batch jobs; the finding is to keep the warehouse and fix ingest. WHERE THE NUMBERS BREAK · ONE FIGURE, TRACED SOURCENIGHTLY JOBSWAREHOUSEDASHBOARDS ERP$4.20M FAILS SILENTLY2 of 30 nights SCHEMA OKno lineage $4.20M$3.88M$4.61M Finance · Sales THE FINDINGKeep the warehouse. Fix ingest, add lineage.The break is two nights a month, upstream of a sound schema.Build nothing new yet.

Your estate looks different? Describe what you have, what it has to carry and who owns it. We will tell you what the diagnostic would likely find before you commit to anything.

Diagnose before you build

Keep, fix or build is the conclusion. The diagnostic is how we get there.

Anyone can label a broken thing "fix". The judgment you are paying for is whether the warehouse is actually the problem — and the seven tests are how that judgment is made, component by component, before anything is replaced.

Seven tests on every component. Then one call.
01Reliability
Does it fail — and can anyone tell when it does?
02Ownership
Who is accountable for it, and do they know?
03Lineage
Can a number on a dashboard be traced back to its source?
04Performance and scale
Can it carry the first workload at the volume you are planning?
05Security and access
Can the right people and systems reach it — and only them?
06Unit economics
What does it cost per run, per query, per terabyte at that volume?
07Operability
Can your team run it, extend it and add a source without us?
Passes all seven → KeepFails some, worth saving → FixDoes not exist → Build
01

Diagnose

The existing estate, the workloads it has to carry, and the platform decision if it is not already made — tested, not assumed. Every component gets a keep, fix or build.

You leave withThe diagnostic: every component through the seven tests, with its call and the reason
02

Stand up

The foundation on Fabric, Databricks or Snowflake, with its security and governance model and its cost controls designed in from day one rather than retrofitted.

You leave withA working foundation with the access model it needs
03

Connect

The first sources flowing, built the way every subsequent pipeline should be built — so the pattern is proven on real data before your team repeats it.

You leave withIts first pipelines, live, as the pattern for the rest
04

Hand over

The runbook — how to operate it, extend it and add the next source — and the team that will run it walked through it on the platform itself.

You leave withThe runbook, owned by your team
Under the hoodWhat "stood up properly" is made of — the build manifest and the acceptance criteria the handover is tested against
# foundation/ — illustrative build manifest
platform/        fabric | databricks | snowflake   # decided on the diagnostic
environments/    dev · prod                            # your cloud account
access/          roles.yaml · row-level policies       # designed before first byte
lineage/         capture on every pipeline             # source → table → dashboard
pipelines/
  erp_extracts.batch.yaml      live  schedule 02:00  alert on fail
  events.stream.yaml           live  latency < 60s   dead-letter on
  _template/                   # the pattern your team copies
cost/            budgets · per-run tags · alerts       # unit cost known
runbook.md       operate · extend · add a source · on-call
  • Fails loudly. A broken pipeline alerts within five minutes; nothing fails silently.
  • Traceable. Every table in the first workload traces to its source in the lineage graph.
  • Governed. Role-based access enforced; the wrong role cannot read the data, measured.
  • Costed. Cost per pipeline run and per query known at the planned volume.
  • Operable. Your team adds a third source from the template, without us, before handover.

The diagnostic, as handed over

Every component. Its call. The reason.

A page from the document your team gets, illustrative. The reasons are the point: an engineer should be able to disagree with any line of it, and the foundation on the right is what the calls add up to.

Data foundation — diagnostic and handoverIllustrative example
ComponentCallWhy
ERP extractsKeepReliable, owned, documented. Reconnected as source one.
CRM feedKeepClean and current. Reconnected unchanged.
Nightly batch jobsFixFail silently twice a month; no alerting. Instrumented, not rewritten.
Warehouse (3 yrs)FixSound schema, no lineage. Lineage added; tables kept.
BI semantic modelKeepBusiness definitions are right. Pointed at the new layer.
Lakehouse layerBuildDoes not exist. Built on the chosen platform, governed from day one.
Access and lineageBuildNever designed. Role-based model and lineage capture built in.
Streaming ingestBuildNeeded for the first AI workload; built as the second pipeline.
The foundation, as handed over

3 kept · 3 fixed · 3 built

Platform
Decided on the diagnostic — fit to the estate, the workload and the operating team
First pipelines
ERP extracts (batch) and event stream (streaming), live, as the pattern for the rest
Access model
Role-based, designed before the first byte landed
Runbook
Operate, extend, add a source — walked through with the team that owns it
Still open: whether the spreadsheet reconciliations retire in one step or two once the model carries their logic. Recorded, with the owner's view.
Illustrative. Components, calls and reasons are examples; no client data. Your readout describes your estate in your team's terms.

Who makes the call

Keep, fix or build is a judgment. This is whose.

The wrong call is the expensive one — a warehouse rebuilt that only needed lineage, or a layer kept that could never carry the workload. So the standard for the person making it is fixed, whoever leads the engagement.

Engagement lead: a named senior architect, confirmed at scoping
Named in the proposal, present through handover, supported by platform-specific engineering as the estate requires. Whoever it is meets all four — this is not a migration crew working a template.
  • Has stood up foundations that are still running. Lakehouses and warehouses on Fabric, Databricks or Snowflake, in more than one industry, operated by the client's own team afterwards.
  • Reads an estate before touching it. Traces numbers from source to dashboard and knows which failures are the pipeline, which are the model and which are the people.
  • Designs governance in, not on. Access, lineage and cost controls from the first pipeline — because retrofitting them is what the next rebuild is made of.
  • Stays accountable. The person who leads the diagnostic is the person who scopes whatever the estate needs next, and remains answerable for it.

The lead is there to say which parts of your estate are fine — before anyone is paid to replace them.

That means naming the three-year-old warehouse that only needs lineage, the batch jobs that need alerting rather than rewriting, and the one layer that is the real constraint. We work across Fabric, Databricks and Snowflake; our platform recommendation is not conditioned on an exclusive vendor relationship.

DiagnoseEvery component: keep, fix or build, with the reason
Stand upThe foundation, with governance and cost controls designed in
ConnectFirst pipelines live, as the pattern for the rest
Hand overThe runbook, walked through with the team that owns it

Evidence

Where building on what existed, rather than replacing it, moved the numbers.

Two engagements in the practice behind this service. In both, the estate that was already there stayed — and the layer built on it did the work.

Manufacturing and distribution · layered on an existing planning system

12%fewer delivery miles per order
18%fewer stockouts at key partners
9%better plan adherence

A large food manufacturer and distributor ran on an MRP backbone that worked. We did not replace it. An AI and analytics layer on top — supply-relationship mapping, cleaner demand and lead-time signals into the existing planner, route optimisation, planner dashboards — did the job the rebuild would have been sold to do.

Read the story →

Healthcare · a people-data foundation

30%reporting efficiency gained
45%fewer reporting errors

A multi-site care provider. Scattered HR spreadsheets replaced by an integrated analytics layer over the core people system they already ran, so leaders work from live data instead of weeks-old reports.

Read the story →

The approach is the one the firm argues in public: many AI programmes stall on the data foundation rather than the model. Why AI projects fail after the pilot →

Fit

Right-sized when the foundation is the constraint and the first workload is known.

If you recognise yourself in the first list, the accelerator is the right size. If the second reads truer, we will point you at the engagement that answers your actual next question.

This is a good fit when

  • You know the platform is the constraint on what you want to build, even if you are not sure which part of it.
  • There is a first workload waiting — a model, a product, a reporting layer — that the foundation has to carry.
  • You can give access to the existing estate and the cloud account the foundation will live in.
  • Your team will operate it afterwards and you want the pattern proven before they repeat it.
  • You want someone to tell you what not to rebuild.

Start somewhere else when

  • You are not sure the data is the problem — that is the Assessment Sprint, which scores it as one of six dimensions.
  • The platform decision is really a vendor negotiation — that is procurement, and we do not take a side.
  • Every source and workload has to move — that is a migration programme in the Data Engineering practice, which this is built to lead into.
  • Nobody will own the platform afterwards; a runbook with no reader is a document.
Deliberately outside the windowA full migration of every source and workload · a data strategy or operating model · master data management · BI dashboards and reports · ML pipelines and model serving · ongoing platform operation · vendor selection. Those are the practices, and the accelerator is built to lead into them with the pattern already proven — which is why this is two to four weeks and not two quarters.

Where it sits

Three questions. Three paths.

The accelerator is where a known constraint gets fixed. Before it is finding out whether the foundation is the constraint; after it is whatever the foundation proved the pattern for.

Want a free read first? The Data Maturity Assessment scores lineage, quality, ownership, access and platform in six minutes.

Before you ask

The buying questions that matter.

The six things a VP of Engineering or an IT director asks before a platform engagement is approved, answered the way we answer them in the proposal.

Which platform?

Whichever fits your data, your existing footprint and your team. We work across Microsoft Fabric, Databricks and Snowflake, and the recommendation is not conditioned on an exclusive vendor relationship — it follows the diagnostic. If the decision is already made, the diagnostic tests it rather than reopening it.

What if the diagnostic shows we do not need a new platform?

Then that is the finding, and the engagement is re-scoped around what you do need — which is often a smaller amount of fixing than a rebuild. The diagnostic exists so that decision is made on evidence rather than on the assumption that arrived with the budget.

Is this a migration?

No. It stands up the foundation properly and connects the first sources as the pattern for the rest. A full migration of every source and every workload is the Data Engineering & Platforms practice, and the accelerator is built to lead into it with the pattern already proven.

Who runs it afterwards?

Your team. The runbook is written for them, they are walked through it, and the first pipelines are built the way every subsequent one should be so the pattern is reusable without us. Where you need more hands, that is a separate conversation about staff augmentation, not a dependency built into the platform.

What does it cost?

The fee is fixed in writing before work starts and depends on the number of sources in scope, the state of what already exists and whether the platform decision is made. Work inside that boundary is not billed hourly.

What do you need from us before kickoff?

Access to the existing estate and the cloud account the foundation will live in, the people who own the first sources, and whatever architecture or platform decisions already exist. The specifics are written into the proposal before you commit.

Start from the estate

Tell us what you have, and what it has to carry.

Not a platform name. Describe the estate as it is, the first workload waiting on it, who owns it, and what has already been decided. We will tell you what the diagnostic would likely find — and the window and fixed fee — before you commit to anything.

The estateSources, platform, pipelines, as they are
The workloadWhat the foundation has to carry first
The ownerWho will run it afterwards
Already decidedPlatform, cloud, conventions

Want a read on the foundation first? Take the free Data Maturity Assessment and bring us the result.

Keep what works. Build what is missing. Show us the estate.

Build only what is missing