The second half of March 2020 will be remembered for many things. For infrastructure teams it was the week everyone went home — and the remote access stack designed for twenty percent mobility was asked to carry nearly everyone. This post is field notes from surge engineering: VPN concentrators at multiplies of design load, emergency VDI and WVD capacity, who gets priority when capacity is finite, and decisions made in seventy-two hours that normally take quarters.
We are writing mid-crisis. Some details will age. The operational patterns will not.
What broke first
In order of how often we saw it:
- VPN capacity and licensing — concurrent SSL/IPsec sessions, CPU on concentrators, split vs full tunnel debates that became physical
- DNS and proxy paths — full tunnel hairpinning Office 365 and Teams media through the datacenter
- VDI/RDS session host capacity — pools sized for contractors suddenly hosting the company
- MFA and helpdesk — identity systems held; human password reset queues did not
- Endpoint chaos — personal devices, home Wi-Fi, unmanaged printers as a political problem
Exchange Online and many SaaS apps were the calm center. The path to them was the storm.
Seventy-two-hour decision pattern
When a client called with “everyone remote Monday,” we ran a brutal triage:
Hour 0–6: Measure actual VPN concurrent and drop rates; identify full-tunnel Office impact; list VDI available headroom.
Hour 6–24: Emergency capacity — burst Azure WVD/Citrix Cloud hosts if identity allows; expand on-prem session hosts if hardware exists; enable split tunnel for Office 365 endpoints where security signs off.
Hour 24–72: Prioritize cohorts (clinical, call center, revenue) for guaranteed capacity; everyone else gets best effort; stand up status comms twice daily.
Perfect architecture lost to available capacity. We documented debt to repay later.
Split tunnel and Teams
Forced tunnel + company-wide Teams meetings is how you melt edge firewalls. Microsoft’s guidance to optimize connectivity for Office 365 is not optional during surge — it is oxygen. Security objections are real; answer them with named location controls, CA policies, and logging, not with “we have always forced tunnel.”
VDI and WVD as relief valves
Non-persistent desktops absorb “I only have a Chromebook” better than VPN-only strategies for some app sets. We stood up emergency host pools knowing images were imperfect. Good enough beat perfect. FSLogix share performance became a page-worthy metric overnight.
Citrix estates with warm spare capacity earned their keep. Estates that ran at ninety-five percent utilization in February had nothing to give in March.
Priority is a business decision
When capacity is finite, IT should not invent social policy alone. We forced a business owner to rank: who must work if only half the VPN licenses exist. Clinical and customer-facing roles usually won. Executives sometimes discovered they were not the bottleneck they thought — or that they were, and bought more capacity by Thursday.
What we are telling people not to do
- Open RDP to the internet “temporarily”
- Disable MFA to reduce tickets
- Promise five-nines while building on consumer broadband assumptions
- Skip change logging because it is an emergency — emergencies are when you need the log most
After the surge (already starting)
Track what you built as temporary. Convert emergency WVD pools into managed services or tear them down. Fix split tunnel properly. Buy the VPN capacity you fake-borrowed. Schedule a retrospective while memory is hot.
Extended practice notes (2020)
The remaining gap between a short checklist and a usable field note is usually scenario detail. In practice, the same engagement type described above still requires explicit answers to: what is in scope this quarter, what is deferred with a date, who can halt a wave, and how success is measured in production — not in a lab.
We document those answers before the first production change. When stakeholders disagree, the disagreement is resolved in writing, not on the bridge at cutover. That habit is independent of whether the workload is mail, identity, virtualization, desktop delivery, or security hardening.
Repeatable detail also includes communication: who tells users what changes, when the freeze starts, where status is posted, and how exceptions are requested. Technical excellence without communication still produces an outage from the user's point of view.
If you're facing this
If your remote access design assumed a minority of users off-campus, rebuild capacity and path optimization now — not after the next spike. We help teams stand up emergency VDI/WVD and stabilize VPN/Office paths under load. Bring concurrent counts and a business ranking of who must work.