That gap is the real subject of legacy system modernization strategies. The usual pitch is a target architecture: cloud, microservices, a new platform. The evidence from audits, technical documentation, and practitioner experience points somewhere else. Successful programs make a separate decision for each system, learn what the old code actually does before replacing it, prove that data and behavior survived the move, and finish by switching the old system off. This guide is for the people who own those decisions, from product and engineering leads to the operators who will live with the result.
An Old System Is Not Automatically a Problem
A system can be fifteen years old and still be the most reliable thing you run. It may process orders correctly, hold a decade of customer history, or encode a pricing rule nobody wants to touch. Age alone does not justify replacing it. One engineer described dropping a planned rewrite of a 2019 Node service once he looked at the numbers: it shipped about one feature a quarter and had gone 18 months without an incident. His summary is worth keeping: technical debt is only debt if it costs you money.
The cost becomes real when the system starts blocking decisions. A new pricing tier takes a quarter to support. A security patch breaks an unrelated feature. One developer with undocumented knowledge has to approve every release. Or the software technically works, but the business around it now runs on spreadsheets, manual approvals, and copy-paste because the system cannot change fast enough. In that last case the pain is operational, and a code rewrite may not fix it.
Maintenance spending is the usual first sign. In Ensono’s 2025 modernization report, 57% of IT teams said they spend most of their time keeping day-to-day operations running, while 25% spend most of their time on new work. Another 49% said legacy maintenance cost more than they had planned the previous year. Ensono sells modernization services, so read the figures as a signal of a pattern rather than a benchmark for your company.
Public-sector numbers show how long this can go on. An earlier GAO testimony on critical legacy systems covered 10 critical federal systems that ranged from about 8 to 51 years old and cost roughly $337 million a year to run and maintain. The lesson is not that every old system needs rebuilding. The lesson is to judge age alongside business dependence, security exposure, operating cost, and failure risk. If you want a plain-language primer on the basics first, F1 Group’s overview of legacy modernisation covers them. Our own legacy application modernization playbook goes deeper on validating changes with real traffic.
The Seven Strategies and What Each One Actually Changes
AWS’s prescriptive guidance groups the options into seven “Rs.” The list is useful because two of the seven involve no rebuilding at all, which reminds you that leaving a system alone or switching it off are real strategies.
| Strategy | What changes | What stays the same | Fits when |
|---|---|---|---|
| Retire | The system is decommissioned | Nothing, though data may need archiving | Usage is low or the capability is duplicated elsewhere |
| Retain | Nothing, for now | Everything | The system is stable, cheap to run, and rarely changes |
| Rehost | Where the code runs | The code and its architecture | The hosting problem is urgent and the code is fine |
| Relocate | The platform underneath, without code changes | The application itself | Moving virtualized workloads to cloud infrastructure |
| Replatform | The runtime, database, or deployment, with light optimization | Core business logic | The foundation causes outages but the rules still work |
| Repurchase | The product, usually to SaaS | Your data and, ideally, your processes | The capability is not a differentiator |
| Refactor | Internal structure and architecture | Observable behavior, if done well | The system is valuable but too hard to change |

No single strategy dominates. In a Red Hat and Konveyor survey of 1,000 respondents fielded in late 2023, replatforming was the most popular single answer at just 20%, and every other option fell between 10% and 19%. The more telling finding was about order. 47% planned to replatform and then refactor, and 38% planned to rehost, then replatform, then refactor. Only 15% planned to go straight to refactoring.
The table leaves one point unclear, so here it is plainly: rehosting is not modernization by itself. Lift and shift changes where the system runs. It does nothing about tangled code, missing tests, or unclear ownership. It can be a sensible first step when a data center contract is ending, but call it what it is.
Two patterns sit beside the seven Rs because they describe how you deliver a change rather than where you end up. Wrapping puts an API or façade in front of the old system so new product work stops touching it directly. The strangler fig pattern, covered below, replaces a system one slice at a time. Most real programs use several of these at once. You might replatform the application, refactor billing, wrap the archive, and strangle the customer portal. We compare these tradeoffs further in our piece on legacy application modernization strategies.
Decide System by System, Not for the Whole Estate
The expensive mistake is not choosing slowly. It is applying one strategy to everything because it is easier to explain in a budget meeting. Systems differ in value, risk, and how often the business asks them to change, and the plan should reflect that.
Start by listing every system that touches revenue or operations: CMS, ecommerce platform, payment flow, CRM, reporting, integrations, data stores, and internal admin tools. Then score each from 1 to 5 on four questions:
- Revenue criticality: How much money or customer activity stops if it fails?
- Failure risk: How exposed is it to security issues, unsupported dependencies, or knowledge held by a single person?
- Maintenance cost: What do licenses, contractors, workarounds, and support actually cost?
- Change frequency: How often does the business ask it to do something new?
A high total does not mean “replace now.” It tells you which system deserves a business case first. A refund service that scores high on risk and criticality may need immediate work, while a dated reporting tool that rarely changes and never blocks a sale can wait. Record one decision per system: keep, wrap, improve, replace, or retire. That record stops the loudest complaint from winning by default.
Retirement is usually the cheapest win, and it comes first. Before Bridgestone migrated its mainframe estate, the team removed applications nobody used. That left four critical applications, covering more than 1.2 million lines of COBOL and JCL, and the migration finished in about seven months. Most of that speed came from cutting the scope.
We had a similar experience with Teton Gravity Research’s platform rebuild. They came to us asking for a redesign, with 10,000 articles locked in an ExpressionEngine install that was barely staying online and earlier migration attempts that had stalled. Before quoting, we asked what the site should actually do. The user-generated content features, once a competitive edge, had become a legal and moderation burden. Retiring them made the migration smaller and the new platform simpler. That was a strategy decision, and it happened before any code was written.
Surveys suggest most organizations end up mixing approaches in this way. Ensono’s 2025 research found 47% of leaders planned to modernize some applications in place, migrate others, and keep core systems intact. Only 11% planned a full retirement and migration.
Discovery Is the Step Most Teams Skip
Every failed modernization one IBM i specialist has seen, he said, “started with replacement instead of learning.” The data backs him up. In Kyndryl’s 2026 survey of 2,000 leaders, only 9% had fully mapped the dependencies of their business-critical applications, and the average enterprise ran 383 of them. When the Social Security Administration’s Inspector General audited its inventory in 2024, it found 819 systems labeled legacy, 415 modern, and 22 unknown. Even those labels were unreliable, partly because some “modern” systems still contained legacy components.
Academic research has the same blind spot. A 2024 mapping study of 109 papers on migration found that most studied the transformation step itself. Far fewer looked at reverse engineering, the work of figuring out what the existing system really does. Yet that is where the hidden rules live: the edge case for a particular client, the rounding quirk finance depends on, the batch job that quietly fixes bad data overnight.
Discovery work that pays off usually includes:
- Talking to the people who know the system, including the long-tenured maintainer and the users who built workarounds around it.
- Characterization tests that record what the system does today, quirks included, rather than what the spec says it should do.
- Tracing inputs and outputs across integrations, so you know which downstream systems read which fields.
- A cost baseline for running the system now, so you can later show whether the change paid off.
Some old bugs are load-bearing. Practitioners who rescue failing migrations say the hardest part is finding the edge cases where behavior changes, and sometimes the right move is to reproduce a bug because a customer or downstream report depends on it. One multi-tenant modernization account made a related point: the painful bugs were not in the frightening modules. They turned up in “the 200th copied CRUD path.”
AI tools can speed up this analysis, but they make a specific mistake. An agent reading undocumented code may decide that unusual edge-case logic is dead code and remove it. Teams using AI for documentation and conversion keep humans accountable for business logic, require approval and audit trails, and do not let agents change production directly. Bridgestone’s migration used automation in the same way, with people overseeing business rules and compliance.
Move First, Redesign Second
AWS’s migration guidance warns against refactoring during a large migration because you take on the complexity of both at once. Practitioners say the same thing more bluntly: migrate as-is, cut the number of variables, and revisit APIs and design once the system is stable in its new home. When something breaks during a combined move-and-redesign, you cannot tell which change caused it.
The tradeoff is real. Separating the steps delays the deeper benefits and can leave architectural limits in place for longer. Some argue that running modernization and new capabilities, such as AI features, one after the other doubles the timeline. That is a fair cost to weigh. For business-critical systems, the evidence still favors controlling risk over saving calendar time.
The same caution applies to microservices. They move complexity rather than remove it. In a 2024 survey of 53 practitioners who had migrated monoliths, 34% reported multiple services writing to the same database tables, which defeats the point of service independence. Respondents also described trading off between eventual consistency and atomic transactions case by case, with no pattern winning everywhere. Split a service when a specific business need justifies the extra network calls, monitoring, and coordination. Do not split one because “microservices” is in the plan. Otherwise you end up with a distributed monolith: everything still depends on everything, and it is now harder to debug.
Strangler Fig: Replacing a System Without a Big-Bang Cutover
A big-bang rewrite tries to replace the whole system in one switchover. One engineer with more than 20 years in the industry could recall only one that succeeded, and it worked because much of the old code was redundant and the underlying problem was simpler than it looked. Rewrites can work when the system is small and well understood, the team knows the domain, and there is a fallback plan. Business-critical systems rarely meet all three conditions.
The strangler fig pattern is the common alternative. You put a routing layer or reverse proxy in front of the old application, build each new capability alongside it, and send traffic to the new path only after it proves itself. Shadow reads let the new system answer the same requests in parallel so you can compare results without users seeing them. A practical sequence is to start with read-only functions and delay anything that writes data until confidence improves.

Traffic ramps should be gradual and reversible. One staged government modernization approach moves from 100% legacy traffic to a 50/50 split and then to roughly 95% on the new path, keeping the old platform as a rollback option and observing for 30 to 60 days before decommissioning. For an ecommerce checkout, measure what customers and the support team will feel at each step:
- Checkout conversion compared with the old path
- Payment and order error rates
- Refunds and duplicate orders
- Support tickets and failed sessions
- Latency and recovery behavior
- Agreement between old and new order records
Agree on rollback triggers before you raise traffic, not during an incident. Conversion falls past the agreed tolerance, payment errors rise, or order records fail reconciliation: any one of these sends traffic back. If the new path depends on an outside service, such as a web data API pulling content from third-party sites, decide before launch who owns it, what happens when it times out, and how its output gets checked.
The pattern has limits. Some systems have no clean module boundary to carve off. One practitioner facing an opaque input system kept it in place and replaced the downstream systems around it instead. Coexistence also costs money, and both systems can survive for years if nobody defines the conditions for finishing. Keep a senior engineer responsible for the legacy path until its traffic reaches zero. Losing that knowledge early turns rollback into guesswork.
Replication Is Not Validation
Migration tools move data. They do not prove the data is right. AWS’s own documentation for its Database Migration Service lists several gaps. Homogeneous migrations between the same database engine have no built-in validation tool. PostgreSQL views arrive as tables. For engines other than MySQL, new tables created on the source during ongoing replication are not picked up automatically. Moving between different engines requires schema and code conversion before any data moves.
None of that is a flaw in the tool. It is a reminder that data assurance is a separate deliverable. A sound plan inventories schema objects, stored procedures, encodings, data types, and in-flight transactions, then checks the result at three levels:
- Records: counts and field-level comparisons for each table or content type.
- Aggregates: totals that finance and operations already trust, such as revenue by month or active members.
- Business outcomes: the same order, invoice, or report produced by both systems from the same input.
Parallel running is the strongest form of proof. Google offers a Dual Run tool that runs workloads on a mainframe and Google Cloud at the same time so outputs can be compared. Meliá kept its source database running while adopting the new one when it modernized a reservations system built on COBOL more than 20 years earlier. Rehearse the cutover and the rollback at least once with production-sized data before the real date.
Publishing migrations follow the same rules at a different scale. When St. Louis Magazine left its legacy CMS, the job was to move 30,000 articles intact, with URLs, media, and editorial workflows still working the morning after. If moving data is the risky part of your plan, our data migration services are built around that testing, redirect, and QA work.

Go-Live Criteria Belong to the Business, Not the Calendar
The most expensive modernization failures usually work technically and still fail the business. Lamb Weston switched to a new ERP system in fiscal Q3 2024. Lower order fulfillment, delayed orders, and facility downtime followed, and the company estimated the hit at $135 million in net sales and about $95 million in net income. Birmingham City Council went live on Oracle in April 2022 and could not reconcile its bank accounts properly. By April 2024, manual reconciliation was costing an estimated £250,000 a month, and projected costs had reached £216.5 million. Press reporting on the audit findings described go-live proceeding despite warnings and testing concerns.
Both are ERP replacements rather than code modernization, but the lesson carries over. Order fulfillment and financial controls should be go-live criteria, signed off by the people who run them. When the calendar overrides readiness, the cost lands on operations.
Cutover should come in waves, each followed by hypercare: a period of heavier support while real users find what testing missed. Dyno Nobel ran its migration in six waves over 11 months across 77 sites, with hypercare after each. Watch the right signals during that period, too. In one AWS case, a rehosted .NET system scaled on CPU, which reacted too late. Each container also kept its own local queue, so new capacity could not pick up waiting work. The fix was a shared queue, scaling on queue depth, and protecting in-flight jobs when capacity scaled down. Clean logs and green builds are not evidence of health. Measure the work the system is supposed to do.
Budget for the Months When You Run Two Systems
Running old and new systems side by side costs more than the new build quote suggests. You pay for parallel infrastructure, duplicate licenses, extra monitoring, reconciliation, user support, and the engineering time to keep two code paths in mind. Ask the team to separate the costs that disappear at cutover from the costs that become permanent. A cheaper hosting quote means little if the old platform keeps running because nobody defined when it can stop.
Overruns are common enough to plan for. Ensono’s 2026 survey found 71% of organizations exceeded planned modernization costs. Practitioner accounts are harsher: a one-year plan that took three, and a project budgeted at six months and $600,000 that ran five years and $9.2 million, even though the replacement eventually worked. No reliable general benchmark exists for duration or cost, and anyone quoting one without seeing your system is guessing.
Dual maintenance also strains the product. One practitioner described a years-long feature freeze on the old version that nearly killed the company. Keep shipping small improvements where you can, and set exit conditions so coexistence has an end.
Staffing usually takes one of three shapes:
- In-house: the most product knowledge, but modernization competes with daily support for the same people.
- Hybrid: internal owners keep business context while a partner adds delivery capacity and migration experience.
- Outside team: faster access to specialists, but you need strong documentation and an explicit handover.
Short-term capacity through staff augmentation services can fill a gap, but extra people will not fix unclear ownership. Every workstream needs one accountable business owner. Protect the people who understand the old system. Practitioners warn that losing migration leads can strand a product halfway between platforms. Invest in the team who will run the new one, as Meliá did with roughly 1,000 hours of cloud training. If maintenance work is crowding out everything else, our guide on how to reduce technical debt faster covers how to tie code cleanup to business decisions.
A Plan That Ends With the Old System Switched Off
GAO’s 2025 audit names three elements a modernization plan must contain: milestones, a description of the work required, and legacy-system disposition, meaning what happens to the old system. Agencies kept deferring their plans until funding or a technical path was clearer. Work started anyway. One HHS system’s expected completion slipped from fiscal 2030 to fiscal 2035. GAO’s conclusion applies well beyond government: incomplete plans make overruns and delays more likely and keep vulnerable systems alive longer.
A workable plan runs in six phases, each with a deliverable and a decision point:
1. Inventory and pain capture
Produce a plain-language map of systems, owners, users, interfaces, data stores, and vendors that says what breaks when each one is unavailable.
2. Portfolio scoring
Apply the four scores with revenue and operations in the room, not just engineering. The output is a short priority list.
3. Strategy selection
Assign an R, or wrap, to each priority system, and write down why it beats the alternatives.
4. Target design
Define behavior, integration boundaries, data ownership, security constraints, and the rollback path before choosing a vendor or framework. A vendor should explain how old and new run together, not only show the new interface. Our guide to choosing a software modernization company covers what to ask.
5. Incremental delivery
Ship one bounded slice, test it against real business rules, and run it beside the old path. Confirm that finance, support, analytics, and compliance still get what they need.
6. Cutover and decommission
Set exit criteria: traffic has moved, data has reconciled, support is trained, documentation has an owner, and rollback is no longer needed. Then retire the infrastructure and contracts, and confirm it actually happened.
Before approving any phase, the people funding it should be able to answer five questions: what business outcome this delivers, which systems and dependencies are in scope, how each phase is tested and monitored, what happens if cutover fails, and how retirement of the old system will be confirmed.
Then close the loop. The SSA audit found the agency could not show legacy operating costs, projected savings, or benefits for the projects it sampled, so nobody could say whether the work paid off. Define success measures before cutover and check them after the system has handled real releases, failures, and support requests. Meliá tracked cost, availability, response time, and request volume. Pair delivery measures such as deployment frequency and change failure rate with numbers leadership already watches: churn, checkout conversion, support volume, and the time it takes to launch a new workflow. Be wary of headline benchmarks. One study of phased modernization programs reports objectives met in 18 to 24 months with availability above 99.9% during the transition. Success rates across studies vary too widely, and are defined too differently, to use any of them as a forecast for your system.
Where to Start This Week
Pick one system. Score it on revenue criticality, failure risk, maintenance cost, and change frequency, then write down the decision you currently believe in: keep, wrap, improve, replace, or retire. Note what you would need to learn to be confident in that call. That list of unknowns is your discovery scope.
The strategy matters less than the discipline around it. Teams that come through modernization in good shape knew what the old system did, kept each change small enough to check, proved the data and behavior survived, kept the business running through cutover, and confirmed the old system was gone. If your plan depends on moving years of records or content safely, Refact’s data migration team starts with that discovery work. It is the same discovery-first process we have used across 200+ projects, and the discovery phase comes with a money-back guarantee.
Building a product and unsure what to scope first? Let’s talk. Free 30-minute call, no pitch.
Masoud Golchin is a backend developer at Refact, working on server-side systems, internal tooling, and infrastructure. He builds and maintains the services that support both client projects and the team’s day-to-day development workflow. His work includes backend logic, developer tools, system reliability, and the technical foundations that allow products to scale and operate consistently. At Refact, Masoud focuses on creating practical engineering solutions that help the team move faster while keeping systems organized, maintainable, and dependable.
More from Masoud Golchin


