AI-Driven Discovery for Cloud Data Modernization: A 4-Stage Blueprint for Legacy Data Migration
AK
Ashish Kumar
Vice President - Business Head Data & Analytics at Tarento Group
September, 2026

Summary: A legacy data migration succeeds when the team understands the source estate before it designs the target. AI-driven discovery gives that understanding in four stages:

  1. Inventory: find every pipeline, table and job, including the undocumented ones.
  2. Map purpose: link each asset to the reports and business processes it feeds.
  3. Decide: choose to migrate, redesign, consolidate or retire each asset.
  4. Standardize: rebuild what stays on a small set of standard patterns.

Most migrations focus on the destination platform. The biggest risks sit in the source estate, in dependencies nobody documented and code nobody uses.

In Brief

  • Overruns are the norm, not the exception. Three independent studies found that data and cloud migrations routinely run late and over budget.
  • Documentation is not an inventory. Legacy estates collect years of patches and undocumented jobs. Automated discovery reads the systems themselves.
  • Purpose decides scope. An asset that feeds no active report or system is a candidate for retirement, not migration.
  • Standard patterns cut long-term cost. Fewer, consistent patterns make the new platform easier to test, debug and maintain.

Overview: The 4-Stage Discovery Blueprint

StageQuestion it answersOutputRisk it removes
1. InventoryWhat exists?Complete list of pipelines, tables, jobs and code, with complexity ratingsUndocumented assets found during cutover
2. Map purposeWhat does each asset do for the business?Lineage from each asset to the reports, systems and processes it feedsHidden dependencies that break business processes
3. DecideWhat should happen to each asset?Migrate, redesign, consolidate or retire decision for every assetPaying to migrate code nobody uses
4. StandardizeHow should the assets that stay be built?A small set of standard pipeline patterns for the target platformLegacy inconsistency copied to the cloud
 Legacy estate                                                        Target platform
 ETL tools, SQL, ──▶ 1. Inventory ──▶ 2. Map purpose ──▶ 3. Decide ──▶ 4. Standardize ──▶ Databricks
 schedulers,         metadata,        lineage to         migrate /      standard            Snowflake
 scripts             complexity       reports and        redesign /     patterns,           Microsoft Fabric
                     scores           processes          consolidate /  naming, logging
                                                         retire

Stage 1: Inventory the Legacy Estate

Stage 1 creates a complete, accurate list of every asset in the legacy estate by reading the systems directly, not the documentation. DataVolve's automated discovery agent connects to legacy platforms and extracts metadata across schemas, ETL workflows and code.

Problem: Legacy estates grow over decades through ad-hoc changes, emergency patches and short-term fixes. Documentation falls behind. When the developers who built the pipelines leave, their knowledge leaves with them. Teams that plan from spreadsheet inventories miss assets, and those assets appear during cutover, when changes cost the most.

Solution: Automated discovery produces the inventory from the systems themselves:

  • Every pipeline, table, view, stored procedure and scheduled job
  • A complexity rating for each asset
  • Upstream and downstream relationships and integration points

For a full description of how the discovery agent works and what the Migration Approach Document contains, see DataVolve Discovery: The Automation Advantage That Sets Modernization Up for Success.

Key Insight: "Documentation describes what the estate was supposed to be. Discovery shows what it is."

Stage 2: Map Each Asset to Its Business Purpose

Stage 2 connects each technical asset to the business outcome it supports. Lineage traces every pipeline to the reports, downstream systems and processes it feeds, so the team knows what each asset does and who depends on it.

Problem: An inventory says what exists. It doesn't say what matters. Without lineage, a pipeline that feeds the month-end finance close looks the same as one that feeds a report nobody has opened in two years.

Here is a composite example. A team migrated a small reconciliation job that looked unimportant in the inventory. Nobody knew it prepared the data that the finance close depended on. It was scheduled for a late wave, and the first month-end close after cutover ran a day late.

Solution: DataVolve analyzes how each pipeline runs within the wider ecosystem and maps its dependencies. From that, the team can see for each asset:

QuestionWhat it tells the team
Which reports or systems does this asset feed?Who is affected if it changes or stops
When did anything last read its output?Whether the asset is still in active use
Which assets must move before it?The order of the migration waves
Which business process does it support?How critical it is

The result separates essential business logic from obsolete code, based on evidence instead of opinion.

Key Insight: "A pipeline's importance is set by what depends on it, not by how complex it looks."

Stage 3: Decide What to Migrate, Redesign, Consolidate or Retire

Stage 3 turns discovery into a decision for every asset. With inventory and lineage in hand, the team chooses one of four paths per asset. The target architecture then follows business needs instead of copying the old estate.

Problem: Many programs migrate everything as it is. They pay to convert, test and run obsolete code. They also copy duplicate logic and old design choices to a new platform, so the technical debt moves to the cloud with the data.

Solution: Apply a clear decision rule to every asset:

DecisionWhen to choose itEffect
MigrateThe asset is in active use and its design fits the target platformAutomated conversion with standard patterns
RedesignThe asset is in active use, but its design does not fit the target, for example a row-by-row process that should become a set-based transformationRebuilt to suit the target platform
ConsolidateSeveral assets perform the same functionMerged into one shared pipeline
RetireNo active report, system or process has read its output in an agreed periodArchived and not migrated

The decisions feed directly into the migration roadmap: wave order, effort estimate and risk list.

Key Insight: "The cheapest pipeline to migrate is the one you retire."

Stage 4: Standardize the Pipelines That Stay

Stage 4 rebuilds every remaining asset on a small set of standard patterns. The new platform is then consistent, easier to test and debug, and cheaper to maintain after the migration ends.

Problem: Legacy estates solve the same task in many different ways. Ten engineers over ten years produce ten ways to load a daily file. Every variant needs its own tests, its own monitoring and its own knowledge. That cost continues long after go-live. In Tarento's analysis of enterprise data estates, 53% of engineering effort goes to maintaining existing pipelines, rising to 61% for estates with more than 200 active pipelines (see Beyond Data Migration: Building AI-Certified Pipelines).

Solution: DataVolve applies a standardization framework during conversion:

StandardWhat it sets
Load patternsOne pattern each for full load, incremental load, change data capture and slowly changing dimensions
NamingConsistent names for pipelines, tables, columns and parameters
ParameterizationShared templates configured per source, instead of copied code
Error handlingOne method to catch, log and retry failures
Logging and monitoringThe same run metrics for every pipeline

The result:

  • New engineers learn one set of patterns instead of hundreds of variants.
  • Monitoring works the same way across the estate.
  • Debugging starts from a known structure.

Key Insight: "Standardization is what keeps a migrated estate from becoming the next legacy estate."


Worked Example: What Discovery Changes in a 2,400-Pipeline Estate

Discovery changes the scope and the shape of a migration before any code is converted. The composite example below is based on patterns common in large enterprise estates.

Starting point: 2,400 pipelines across three ETL tools. The documented inventory listed 1,900, so 500 pipelines were undocumented.

Discovery resultPipelinesDecision
Output not read by any active report or system in 12 months430Retire
Duplicate logic across tools and teams290Consolidate into 70 shared pipelines
Design does not fit the target platform160Redesign
Active, and the design fits the target1,520Migrate with standard conversion
Total legacy pipelines2,400
ResultBefore discoveryAfter discovery
Pipelines on the target platform2,400 (planned)1,750
Reduction in pipelines to build, test and runNone650 (27%)
Undocumented pipelines found before cutover0500

Every one of the 650 pipelines removed from scope is a pipeline the team doesn't convert, test, monitor or maintain. The 500 undocumented pipelines were found during planning instead of during cutover.


Where to Start

Start with Stage 1 on one business domain, not the whole estate. Run automated discovery, compare the result with your documented inventory, and count the difference. That number is the size of the risk your current plan doesn't cover.

Across DataVolve engagements, Tarento measures these results against a manual migration baseline of the same scope:

OutcomeResult with DataVolve
Engineering effort30 to 60% lower
Total migration cost20 to 40% lower
Time to valueWithin 50 to 60% of a typical project timeline
Post-migration incidents50 to 60% fewer

Problem: Legacy data estates are opaque. Teams plan from outdated documentation, migrate unused code and find hidden dependencies during cutover. Solution: A four-stage, AI-driven discovery blueprint that inventories the estate, maps each asset to its purpose, decides its future and standardizes what stays. Vision: Migration becomes a planned engineering process with a known scope. The new platform starts smaller, cleaner and fully documented, and every future change begins from a map instead of a guess.

If you want to know what your documented inventory is missing, start an AI-driven discovery assessment with DataVolve.

AK

ABOUT THE AUTHOR

Ashish Kumar
Vice President - Business Head Data & Analytics at Tarento Group
logo
Thor Bot Avatar