Two dates that set the whole schedule
17 June 2025 is the date VMware Cloud Foundation 9 became generally available, and it started a clock that has nothing to do with how well your TKGI clusters run today. A platform that boots clean every morning still ages on the vendor calendar, and that calendar is the spine of the business case. Two dates matter. One is when the destination became real, VCF 9 with VKS as its Kubernetes runtime, which is now. The other is when the origin stops being supported, the TKGI End of General Support date for the version you run, which is a hard boundary you plan backward from.
Confirm the exact TKGI date on the Broadcom Product Lifecycle Matrix for the version you actually run, because the number drives everything downstream and I will not invent it here. Our running estate sits on TKGI 1.18, and once a TKGI release reaches End of General Support you can still create and delete clusters, but you stop getting fixes and you cannot open a support case against it. Read that carefully: nothing breaks on the deadline, which is exactly why teams let it slide. What you lose is the safety net, right when a large migration is the moment you most want one. [VERIFY: exact TKGI End of General Support date per Broadcom Product Lifecycle Matrix for your running version.]
Set against that origin date is the destination, and here the news is good for the budget conversation. VKS is the CNCF conformant Kubernetes runtime included with VCF 9 at no separate licence line, so the Kubernetes engine itself is not a new purchase. VKS 3.4, which shipped alongside VCF 9, carries 24 months of extended support for its Kubernetes release, which means the platform you land on is not about to age out from under you the way the origin is. That asymmetry, an origin on a countdown and a destination with two years of runway, is the entire reason to move now rather than next budget cycle.
A licensing dimension sits underneath the dates as well. Since Broadcom completed the VMware acquisition in late 2023, the portfolio moved to subscription licensing bundled into VCF, and VKS rides inside that bundle rather than standing as a separate Kubernetes purchase. For a budget owner that changes the shape of the ask. You are not buying a new product to replace TKGI, you are consuming Kubernetes you already pay for as part of VCF 9. State that plainly in the business case, because the reflex assumption is that a platform migration means a fresh licence line, and here it does not. What you fund is overlap and effort, not a new product SKU.
One clause for a platform people fold in by mistake. If you also run TAS, the Cloud Foundry application runtime once sold as Tanzu Application Service, its path is Tanzu Platform for Cloud Foundry, a separate product line, and it does not belong in a VKS plan or its budget. This series and this business case move Kubernetes clusters from TKGI to VKS, and nothing else. Part 3 covered why that boundary is firm, in Why This Is a Migration, Not an Upgrade.
Cost lines a real migration budget carries
A budget owner wants the migration priced as line items, not as a vibe, so give it to them as three concrete costs plus staff time. First is overlap capacity. Because you stand VKS up beside TKGI and move workloads in waves, both platforms draw power and hardware at once, and for a mid size estate that means running with roughly 20 to 30 percent extra capacity through the busiest waves. Second is object storage behind Velero, sized to the backup volume of everything you migrate, which is small next to compute but real and easy to forget. Third is the parallel run itself, the months of two platforms consuming power, cooling and rack space before you reclaim the old hardware.
What is not a new line is the Kubernetes runtime licence, because VKS ships inside VCF 9. That single fact reframes the conversation with finance, since the spend is overlap and effort, not a fresh product purchase stacked on top of a renewal. Put the lines on one page so nobody treats the overlap as waste, because the overlap is the part that buys you a safe rollback at every wave boundary.
Put rough shapes on the two soft lines so they do not get waved away. Staff time is the one that surprises finance, because a wave is not a script you run once, it is a backup, a restore, a validation pass and a traffic cut, repeated per workload group, with a soak in between. Budget it as weeks of focused platform engineering across phases 4 and 5, not as a single weekend. Parallel run overhead is steadier but longer, since two platforms draw power and rack space for the whole coexistence window, which on the running estate is roughly weeks 5 through 24. Neither line is optional, and neither is waste.
| Cost line | What drives it | When it lands |
|---|---|---|
| Overlap capacity on VCF 9 | 20 to 30 percent extra during peak waves | Phase 3 onward |
| Object storage for Velero | Backup volume of migrated workloads and data | Phase 4 |
| Parallel run overhead | Two platforms drawing power and space for months | Phase 3 to 6 |
| Staff time | Wave execution, validation and cutover | Phase 4 to 5 |
| VKS Kubernetes runtime | Included with VCF 9, no separate licence | None, already owned |
Lifecycle drivers, ranked by urgency
Not every reason to move carries the same weight, and a plan that treats them as equal wastes energy on the low ones. Rank them. Support expiry and hardware procurement sit at the top because they gate the calendar itself. Kubernetes version drift and compliance sit in the middle, real but bounded. Staff familiarity sits at the bottom, not because it does not matter but because you buy it down cheaply with an early throwaway cluster rather than a line in the budget. This table is the reference artifact of this part, a driver to action map you can lift straight into a planning deck and defend line by line.
| Driver | Why it moves your date | Urgency | First action |
|---|---|---|---|
| TKGI End of General Support | Fixes and support cases stop for your version | High | Confirm your version date on the lifecycle matrix |
| VCF 9 capacity procurement | Hardware lead time gates standing up VKS | High | Get a purchase order moving now |
| Kubernetes version drift | Old TKGI Kubernetes falls behind upstream support | Medium | List cluster versions against upstream dates |
| Security and compliance | An unpatched platform fails an audit window | Medium | Flag audit dates that land inside the migration |
| Staff familiarity with VKS | Learning curve slows the earliest waves | Low | Run a throwaway pilot cluster early |
Phased timeline you can hand a budget owner
Six phases carry the whole series, and they map cleanly onto a calendar. Phases 1 and 2 are planning and design, no hardware yet. Phase 3 stands VKS up beside TKGI, which is where the two platforms first run together. Phase 4 moves workloads in waves with Velero, stateless first to build the muscle, then the hard stateful apps. Phase 5 pilots a non production cluster end to end and then shifts production traffic with a soak and a rollback ready at each step. Phase 6 reclaims the TKGI and Ops Manager capacity once the last wave is proven. For a mid size estate of three clusters the whole run lands in four to six months, and the shape holds even when the exact weeks move.
Phases overlap on the calendar even though they read as a sequence, and the overlap is the point. You do not finish all migration before you start the pilot, and you do not finish the pilot before the first production wave. The chart below places the six phases on a 24 week grid for the running estate, with the coexistence window marked underneath. Notice how phases 4 and 5 run alongside each other for most of the project, and how the decommission at the end is short because the risk was already spent in the waves.
| Phase | Focus | Typical duration | Both platforms live |
|---|---|---|---|
| 1 Case for moving | Framing, approvals, budget sign off | 1 to 2 weeks | No |
| 2 Assess and design | Inventory, waves, target architecture | 3 to 4 weeks | No |
| 3 Stand up VKS | VCF 9, Supervisor, first VKS cluster | 3 to 4 weeks | Begins here |
| 4 Migrate workloads | Velero waves, stateless then stateful | 6 to 10 weeks | Yes |
| 5 Pilot and cutover | Non prod pilot, production traffic shift | Overlaps phase 4 | Yes |
| 6 Decommission | Reclaim TKGI and Ops Manager capacity | 1 to 2 weeks | Ends here |
Where the timeline actually slips
Here is where the obvious plan goes wrong, and it goes wrong the same way almost every time. The tutorial instinct, and the instinct of most change boards, is to anchor the Gantt chart to the TKGI End of General Support date and count backward. That feels disciplined and it is exactly the wrong anchor. Two things set your real earliest finish, and neither is the support date. One is the hardware lead time for the VCF 9 capacity, because phase 3 cannot start until the gear arrives. The other is your slowest stateful application, the one with a large volume and a long soak requirement, because that single app can own weeks of phase 4 on its own.
Anchor the plan to those two constraints and the support date becomes what it should be, a finish line you comfortably clear rather than a start gun you sprint from. Anchor it to the support date instead and you compress the front of the plan, discover the procurement lead time late, and meet your hardest app under deadline pressure with no slack left. Sequence by risk, not by the calendar countdown. Start the throwaway pilot cluster before the assessment is even finished, because the learning it buys is worth more early than a tidy phase boundary.
A second, quieter slip follows the same pattern. Teams schedule the stateful apps last, reasoning that the stateless ones are easier and build confidence, which is correct as far as it goes. Then they find that the hardest app also needs the longest soak and the most validation, so leaving it for the end stacks the riskiest work right against the deadline. Run the stateless waves early for practice, but scope and rehearse the hardest stateful app in parallel, on the throwaway pilot cluster, long before its real wave arrives. Surprises on that one app are the ones that move the finish date.
Anchor the plan to your hardest app, not the deadline
My recommendation is simple and it shapes every part that follows. Build the business case on the lifecycle asymmetry, an origin on a support countdown and a destination with two years of runway that you already own inside VCF 9. Price the migration as overlap capacity, Velero storage and a parallel run, and sell the overlap as rollback insurance rather than waste. Then draw the timeline from your two real constraints, hardware lead time and your slowest stateful application, and let the support date be the finish line you clear rather than the anchor you hang everything from. Teams that do this land early with slack to spare. Teams that anchor to the deadline meet their hardest app in the worst possible week.
Next part opens Phase 2 and gets hands on, inventorying the real TKGI estate with discovery commands and a checklist you can run against your own clusters. Bring the phased timeline from this part, because the inventory is what turns those phase boxes into dates. For the target platform you are building toward, the VKS Series covers VKS on VCF 9 in depth.
References
- VMware Cloud Foundation blog, VMware vSphere Foundation 9.0 now available, 17 June 2025
- VMware Cloud Foundation blog, VKS 3.4 with Kubernetes extended support
- Broadcom Knowledge, expected behavior when a TKGI release reaches End of Life or End of Support


DrJha