, ,

TKGI to VKS Migration Business Case and Phased Timeline (TKGI to VKS Series, Part 4)

TKGI is on a lifecycle clock and VCF 9 is the destination. Here is the business case, the drivers that set your real deadline, and a phased timeline you can hand a budget owner.

TKGI to VKS Series · Part 4 of 26
Key takeaways: VMware Cloud Foundation 9 went generally available on 17 June 2025, and that release, not the health of your current clusters, is what starts the migration clock. TKGI is on a defined lifecycle and will reach End of General Support, so the business case is about landing before support runs out, not about a platform that is failing. Budget the migration as three real cost lines, overlap capacity, Velero object storage, and a parallel run, and treat the overlap as rollback insurance. Anchor the timeline to your slowest stateful application and your hardware lead time, because the support deadline is a finish line, not a start gun.
Who this is for: A platform owner, SRE lead or VMware admin who has accepted that TKGI is winding down, understands why this is a migration rather than an upgrade, and now has to put numbers and a calendar in front of a budget owner. Part 3 made the case that you migrate rather than upgrade. This part turns that framing into a business case, a set of lifecycle drivers, and a phased timeline. Terms on first use: TKGI is Tanzu Kubernetes Grid Integrated, formerly Enterprise PKS; VKS is the vSphere Kubernetes Service on VMware Cloud Foundation 9; VCF is VMware Cloud Foundation; Ops Manager is the tile based console that installs the BOSH managed TKGI platform; Velero is the open source Kubernetes backup and restore tool; End of General Support is the lifecycle phase after which a version no longer receives standard fixes or support cases.

Two dates that set the whole schedule

17 June 2025 is the date VMware Cloud Foundation 9 became generally available, and it started a clock that has nothing to do with how well your TKGI clusters run today. A platform that boots clean every morning still ages on the vendor calendar, and that calendar is the spine of the business case. Two dates matter. One is when the destination became real, VCF 9 with VKS as its Kubernetes runtime, which is now. The other is when the origin stops being supported, the TKGI End of General Support date for the version you run, which is a hard boundary you plan backward from.

Confirm the exact TKGI date on the Broadcom Product Lifecycle Matrix for the version you actually run, because the number drives everything downstream and I will not invent it here. Our running estate sits on TKGI 1.18, and once a TKGI release reaches End of General Support you can still create and delete clusters, but you stop getting fixes and you cannot open a support case against it. Read that carefully: nothing breaks on the deadline, which is exactly why teams let it slide. What you lose is the safety net, right when a large migration is the moment you most want one. [VERIFY: exact TKGI End of General Support date per Broadcom Product Lifecycle Matrix for your running version.]

Set against that origin date is the destination, and here the news is good for the budget conversation. VKS is the CNCF conformant Kubernetes runtime included with VCF 9 at no separate licence line, so the Kubernetes engine itself is not a new purchase. VKS 3.4, which shipped alongside VCF 9, carries 24 months of extended support for its Kubernetes release, which means the platform you land on is not about to age out from under you the way the origin is. That asymmetry, an origin on a countdown and a destination with two years of runway, is the entire reason to move now rather than next budget cycle.

A licensing dimension sits underneath the dates as well. Since Broadcom completed the VMware acquisition in late 2023, the portfolio moved to subscription licensing bundled into VCF, and VKS rides inside that bundle rather than standing as a separate Kubernetes purchase. For a budget owner that changes the shape of the ask. You are not buying a new product to replace TKGI, you are consuming Kubernetes you already pay for as part of VCF 9. State that plainly in the business case, because the reflex assumption is that a platform migration means a fresh licence line, and here it does not. What you fund is overlap and effort, not a new product SKU.

One clause for a platform people fold in by mistake. If you also run TAS, the Cloud Foundry application runtime once sold as Tanzu Application Service, its path is Tanzu Platform for Cloud Foundry, a separate product line, and it does not belong in a VKS plan or its budget. This series and this business case move Kubernetes clusters from TKGI to VKS, and nothing else. Part 3 covered why that boundary is firm, in Why This Is a Migration, Not an Upgrade.

Cost lines a real migration budget carries

A budget owner wants the migration priced as line items, not as a vibe, so give it to them as three concrete costs plus staff time. First is overlap capacity. Because you stand VKS up beside TKGI and move workloads in waves, both platforms draw power and hardware at once, and for a mid size estate that means running with roughly 20 to 30 percent extra capacity through the busiest waves. Second is object storage behind Velero, sized to the backup volume of everything you migrate, which is small next to compute but real and easy to forget. Third is the parallel run itself, the months of two platforms consuming power, cooling and rack space before you reclaim the old hardware.

What is not a new line is the Kubernetes runtime licence, because VKS ships inside VCF 9. That single fact reframes the conversation with finance, since the spend is overlap and effort, not a fresh product purchase stacked on top of a renewal. Put the lines on one page so nobody treats the overlap as waste, because the overlap is the part that buys you a safe rollback at every wave boundary.

Put rough shapes on the two soft lines so they do not get waved away. Staff time is the one that surprises finance, because a wave is not a script you run once, it is a backup, a restore, a validation pass and a traffic cut, repeated per workload group, with a soak in between. Budget it as weeks of focused platform engineering across phases 4 and 5, not as a single weekend. Parallel run overhead is steadier but longer, since two platforms draw power and rack space for the whole coexistence window, which on the running estate is roughly weeks 5 through 24. Neither line is optional, and neither is waste.

Cost lineWhat drives itWhen it lands
Overlap capacity on VCF 920 to 30 percent extra during peak wavesPhase 3 onward
Object storage for VeleroBackup volume of migrated workloads and dataPhase 4
Parallel run overheadTwo platforms drawing power and space for monthsPhase 3 to 6
Staff timeWave execution, validation and cutoverPhase 4 to 5
VKS Kubernetes runtimeIncluded with VCF 9, no separate licenceNone, already owned

Lifecycle drivers, ranked by urgency

Not every reason to move carries the same weight, and a plan that treats them as equal wastes energy on the low ones. Rank them. Support expiry and hardware procurement sit at the top because they gate the calendar itself. Kubernetes version drift and compliance sit in the middle, real but bounded. Staff familiarity sits at the bottom, not because it does not matter but because you buy it down cheaply with an early throwaway cluster rather than a line in the budget. This table is the reference artifact of this part, a driver to action map you can lift straight into a planning deck and defend line by line.

DriverWhy it moves your dateUrgencyFirst action
TKGI End of General SupportFixes and support cases stop for your versionHighConfirm your version date on the lifecycle matrix
VCF 9 capacity procurementHardware lead time gates standing up VKSHighGet a purchase order moving now
Kubernetes version driftOld TKGI Kubernetes falls behind upstream supportMediumList cluster versions against upstream dates
Security and complianceAn unpatched platform fails an audit windowMediumFlag audit dates that land inside the migration
Staff familiarity with VKSLearning curve slows the earliest wavesLowRun a throwaway pilot cluster early
Driver note: Procurement outranks almost everything on this list in practice, even though it never shows up as a technical reason to migrate. You cannot stand up VKS on hardware you have not received, so the purchase order is on the critical path from day one. Raise it before the assessment is finished, not after.

Phased timeline you can hand a budget owner

Six phases carry the whole series, and they map cleanly onto a calendar. Phases 1 and 2 are planning and design, no hardware yet. Phase 3 stands VKS up beside TKGI, which is where the two platforms first run together. Phase 4 moves workloads in waves with Velero, stateless first to build the muscle, then the hard stateful apps. Phase 5 pilots a non production cluster end to end and then shifts production traffic with a soak and a rollback ready at each step. Phase 6 reclaims the TKGI and Ops Manager capacity once the last wave is proven. For a mid size estate of three clusters the whole run lands in four to six months, and the shape holds even when the exact weeks move.

flowchart TB
  subgraph Plan and build, no coexistence yet
    P1[Phase 1 Case for moving] --> P2[Phase 2 Assess and design] --> P3[Phase 3 Stand up VKS beside TKGI]
  end
  subgraph Move and land, both platforms live
    P4[Phase 4 Migrate workloads in waves] --> P5[Phase 5 Pilot and production cutover] --> P6[Phase 6 Decommission TKGI and Ops Manager]
  end
  P3 --> P4
Coexistence begins at phase 3 and ends at phase 6. Everything between those two points runs on two platforms at once by design.

Phases overlap on the calendar even though they read as a sequence, and the overlap is the point. You do not finish all migration before you start the pilot, and you do not finish the pilot before the first production wave. The chart below places the six phases on a 24 week grid for the running estate, with the coexistence window marked underneath. Notice how phases 4 and 5 run alongside each other for most of the project, and how the decommission at the end is short because the risk was already spent in the waves.

Migration timeline, mid size estate of three clustersSix phases across 24 weeks, with the coexistence window marked belowPhase 1 CasePhase 2 AssessPhase 3 Stand upPhase 4 MigratePhase 5 Pilot cutoverPhase 6 DecommissionWeek 0Week 6Week 12Week 18Week 24Coexistence, both platforms live, week 5 to week 24
Planning figures for a three cluster estate. Stateful apps and hardware lead time stretch or compress the grid, but the overlap of phases 4 and 5 is constant.
PhaseFocusTypical durationBoth platforms live
1 Case for movingFraming, approvals, budget sign off1 to 2 weeksNo
2 Assess and designInventory, waves, target architecture3 to 4 weeksNo
3 Stand up VKSVCF 9, Supervisor, first VKS cluster3 to 4 weeksBegins here
4 Migrate workloadsVelero waves, stateless then stateful6 to 10 weeksYes
5 Pilot and cutoverNon prod pilot, production traffic shiftOverlaps phase 4Yes
6 DecommissionReclaim TKGI and Ops Manager capacity1 to 2 weeksEnds here

Where the timeline actually slips

Here is where the obvious plan goes wrong, and it goes wrong the same way almost every time. The tutorial instinct, and the instinct of most change boards, is to anchor the Gantt chart to the TKGI End of General Support date and count backward. That feels disciplined and it is exactly the wrong anchor. Two things set your real earliest finish, and neither is the support date. One is the hardware lead time for the VCF 9 capacity, because phase 3 cannot start until the gear arrives. The other is your slowest stateful application, the one with a large volume and a long soak requirement, because that single app can own weeks of phase 4 on its own.

Anchor the plan to those two constraints and the support date becomes what it should be, a finish line you comfortably clear rather than a start gun you sprint from. Anchor it to the support date instead and you compress the front of the plan, discover the procurement lead time late, and meet your hardest app under deadline pressure with no slack left. Sequence by risk, not by the calendar countdown. Start the throwaway pilot cluster before the assessment is even finished, because the learning it buys is worth more early than a tidy phase boundary.

A second, quieter slip follows the same pattern. Teams schedule the stateful apps last, reasoning that the stateless ones are easier and build confidence, which is correct as far as it goes. Then they find that the hardest app also needs the longest soak and the most validation, so leaving it for the end stacks the riskiest work right against the deadline. Run the stateless waves early for practice, but scope and rehearse the hardest stateful app in parallel, on the throwaway pilot cluster, long before its real wave arrives. Surprises on that one app are the ones that move the finish date.

War story: One team anchored their entire Gantt to the TKGI support date, which sat a comfortable eight months out, and treated the runway as generous. When they finally raised the VCF 9 hardware purchase order, it carried a 12 week procurement lead time that nobody had costed into the plan, and that lead time quietly ate a third of the runway before a single VKS cluster existed. We redrew the schedule starting from the capacity order and from the one PostgreSQL backed app that needed a long soak, and the support date stopped being the anchor. Same deadline, same estate, but the plan finally hung off the two constraints that were actually load bearing.
Contrarian take: The support deadline is the least useful date in your plan. It tells you when you must be done, which you already knew, and nothing about when you can start or how long the hard parts take. Build the schedule from procurement lead time and your slowest stateful app, then check it clears the deadline. A plan built the other way around looks neat on the day it is drawn and falls apart the first time reality touches it.

Anchor the plan to your hardest app, not the deadline

My recommendation is simple and it shapes every part that follows. Build the business case on the lifecycle asymmetry, an origin on a support countdown and a destination with two years of runway that you already own inside VCF 9. Price the migration as overlap capacity, Velero storage and a parallel run, and sell the overlap as rollback insurance rather than waste. Then draw the timeline from your two real constraints, hardware lead time and your slowest stateful application, and let the support date be the finish line you clear rather than the anchor you hang everything from. Teams that do this land early with slack to spare. Teams that anchor to the deadline meet their hardest app in the worst possible week.

Do this on Monday: Open the Broadcom Product Lifecycle Matrix, write down the End of General Support date for your exact TKGI version, then raise the VCF 9 capacity purchase order the same week. Verdict: build the timeline from procurement lead time and your slowest stateful app, and treat the support date as a finish line. Avoid anchoring the plan to the deadline, it is the one date that tells you nothing about how to start.

Next part opens Phase 2 and gets hands on, inventorying the real TKGI estate with discovery commands and a checklist you can run against your own clusters. Bring the phased timeline from this part, because the inventory is what turns those phase boxes into dates. For the target platform you are building toward, the VKS Series covers VKS on VCF 9 in depth.

TKGI to VKS Series · Part 4 of 26
« Previous: Part 3  |  Guide  |  Next: Part 5 »

References

About The Author


Discover more from Journal of Intelligent Infrastructure

Subscribe to get the latest posts sent to your email.

Leave a Reply

Your email address will not be published. Required fields are marked *

Architect’s Toolkit

About the Author

Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.

Discover more from Journal of Intelligent Infrastructure

Subscribe now to keep reading and get access to the full archive.

Continue reading