Load Balancing Load Balancers: Converting to Active Availability
by dasseclab
Overview
I generally have a policy that I don’t talk about “job-specific” stuff on the Internets; however, I decided to write this up for a couple of reasons. The first - as I get further away from the job I did this project at, the more I forget the day to day but this is still a good chunk of my resume, so I want to have a reference to remember what I did. The second - I’ve actually forgotten enough, most likely, that there are very specific details are lost and this is discussed in broad generalities.
In a former position, I was working with load balancing clusters that provided us Large-Scale NAT (LSN) capabilities for our massive data center presences. The idea of these LSN clusters was that access would be granted to specific hosts that met business justifications for having outbound public Internet access - HTTP/S proxies of course but some other internal systems for this access also met this threshold. And like every organization with an IPv4 network, IPv4 address space is a precious, limited resource. And like many large enterprises, IPv6 networks are too new and scary (still). This project I organized allowed us to preserve IPv4 space, improve reliability and mitigate long-term cost incursion. Its success also wouldn’t have been possible without a lot of senior (and senior-plus) engineers answering my questions and reviewing my work.
One last author’s note - in the internal technical documentation for this project, I obviously used the actual IP prefixes and addresses that would be assigned to devices. As this is an example published to public Internet, IP prefix documentation here uses RFC-5737 compliant TEST-NET IP prefixes and RFC-5398 for Autonomous System Numbers.
Problem Statement
We had specific services within our data center architecture that required outbound access to different services. To perform Network address translation (NAT) and meter the connections, we used clusters of load balancers based on internal service type. These load balancer clusters were originally designed with an High-Availability (HA) architecture and adequately scoped with lots of bandwidth for a long runway of opertaion. However, as all runways do, the runway began to run out and capacity through the clusters began to suffer performance hits and trigger alarms for our on-calls. If left alone, the problem would only get worse and the packets that dropped on the floor would eventually cause true perfomance issues for our customers. We explored a few solutions and eventually decided upon a complete redesign of the system functionality.
The Origin Story: High Availability Large Scale NAT
Any number of network devices, by the 2010s, could do address translation. As mentioned, these were for massive data center presenses/campuses, so using a firewall device could cause significant bottlenecks at scale and the security features weren’t necessarily in-scope with primarily outbound initiated connections. Using our edge routers that we were already using for Internet connectivity, we probably could have done address translation, too but that would increase the scope of impact if the routers became overwhelmed. Instead, we opted for new hardware load balancers - new clusters to isolate their function from other LB services - which catered to the CG-NAT (RFC 6598) market, thus having robust address translation software running on it.
We scoped the clusters out into initially two clusters - one focused on external proxy services and services like DNS or NTP and one devoted to a class of storage services that backed up their coldest storage to public cloud storage providers. As our serving infrastructure grew, a third cluster was added as a second resource for cold storage. The fourth and final cluster was initially developed for more latency-sensitive backend services but it also ended up sharing some of the load from the core services proxies.
Obligatory “Use IPv6, You Idiot” Statement and Defense
Of couse, NAT is only a stop gap for conserving IPv4 address space. If we truly, TRULY, wanted to move away from address availability constraints, we would be using IPv6. Anytime anyone mentions anything related to NAT, just about, you can guarantee there will be someone saying you should be using IPv6 instead of NAT.
And they’re not entirely wrong - given that the purposed of address translation was for traversing the Internet and for well known functions and cloud providers, IPv6 at least externally could have been a contender. But depsite the scale of our environment, IPv6 adoption internally was consistenly de-prioritized, often from software engineering teams. I guess having dead::beef being a completely valid IP address was a little too intimidating. Given that, all the IPv6 resoucring was all geared towards peering and transit on the Internet but after routing, all service connectivity was done using IPv4. That’s the funny thing about running company networks, is that engineering often has to make trade offs and pivots around business requirements, even if those requirements are from other engineers.
Proposed Remedies
Proposal: Do Nothing [REJECTED]
What most engineers don’t understand is that ‘do nothing’ is actually an option. If we do not evaluate the cost of ‘do nothing’, it makes it more difficult to assign (usually) already strapped resources to a project and, more importantly, it immediately gives credence that every complaint is a problem that requires engineering a solution to, which isn’t always the case.
In this specific instance, we’d looked over several months of data and there were periodic instances where traffic through the LSN load balancers would peak, causing pages to oncall. During these peaks, we would see some packet loss, though the interfaces were never fully saturated for long. Interfaces would fall back to about 60% utilization after sixty to ninety minutes. Peaks were rarely correlated, such as trying to pinpoint a specific service. Given the near random nature and infrequency of the events - do nothing could have been a real option. It’d have been a minor pain for on-call to have to look at an event but also knowing that it wasn’t severely impactful and would subside on its own could have placed it much lower on the stack of priorities to tackle for another year or so since the runway hadn’t completely run out.
What eventually ruled this proposal out as being an option were capacity requests and predictuons coming down the pike for future product launches and needs that expected to double traffic over 18 months and potentially triple over 24 months. If traffic remained within some level of current levels, we could have ridden this out for another year or so with only minor pain points but with this much traffic forecasted, something had to be done.
Proposal: Expand As-Is == Expensive [REJECTED]
At the time to consider re-designing the large-scale NAT system, we’d already had two clusters dedicated to one service. Another cluster that was recently turned up had a very small customer base and we could re-purpose it to offload some customers from our paging devices. Each of these devices cost tens of thousands of dollars in initial capital plus annual maintenance fees and when a new cluster is turned up, an identical cluster needs to be deployed in each production data center where the backend services live. Easily, the design of one new cluster, grows from not just one pair but three to five pair of new clusters. The operational overhead of the cluster, once deployed and integrated, is negligible but the upfront costs easily made this proposal very unattractive.
Proposal: Do It In Software [REJECTED]
Of all the alternative proposals, this was bandied about and explored a good bit before we reached opted to pursue other avenues. In fact, we ended up putting software load balancing into a parking lot to work on a future service - basically pointing out that this was the direction we should be moving but we’re currently not in a position to execute now to mitigate today’s problems.
We assessed that our team make up at the time was more of a traditional network engineer profile and did not have enough of the cross-over skillsets needed to run our own software load balancer as a distributed microservice successfully (such as the service not dying in production and causing outages). The team make up was planning to shift though - we already had a few other products that could be run as a distributed microservice - and hiring a few heads on a SRE job ladder was starting to happen. But throwing a project like this at them where there was no middle-ground (hardware-based networking to pure software) while our pagers continued to blow up, was not a good way to solve this problem.
Proposal: Load Balance The Load Balancers [ACCEPTED]
This proposal, while being a significant technical undertaking, won out on the financial fronts because it was capital neutral immediately as no new equipment needed to be purchased. High Availability architecture means that there are two devices, one in Primary mode processing traffic and another failover unit should the Primary unit fail. This means that the High Availability architecture limits your bandwidth based on the lowest common demoninator being the Primary member of the cluster. Bandwidth requirements were scoped adequately for future growth at implementation but some unforseen growth over the years meant new clusters needed to be deployed. The Active Architecture though, each member node processes traffic acting independently, which means the available bandwidth increases with the scale of the cluster. Active Architecture also provides a way to grow scalably by being able to add new nodes individually rather than having to build out in pairs.
Technical Designs
Most of this post so far has been going over the technical design of the differing NAT architectures but the chief other networking component to discuss in technical brief is our data center edge routing layer. This section will explain a few of the device or functional layers to illustrate each layer and then go into the architecture designs.
High Availability Architecture
Data Center Edge Routing Layer
Our data center edge routing layer is a network layer that connected each of our production data center facilities via backbone connections to other network PoPs. Unlike other edge designs, our network edge layers actually functioned in two segments. The data center edge did not peer with upstream transit providers or external customers, instead, they peered with other internal routers and load balancers. The other segment of our edge layer, our core router layer, peered with other backbone routers and upstream transit and private interconnection and announced our larger network prefixes. The data center edge announced more specific and anycast prefixes within our network to the core router, which announced the larger prefixes. Production Data Centers started operation with a cluster of two devices for DC Edge routing, which then grew to four routers. Each member of the cluster was peered with another for a full BGP mesh.
In terms of the Large-Scale NAT clusters, the DC Edge routers were connected to each cluster member in a redundant fashion - for instance, a NAT cluster with 80Gb/s of capacity, each NAT device was connected to each edge router at 4x10Gb/s and aggregated via LACP. When moving away from the High Availability architecture to the Active Architecture, this was retained. The publicly routable (non-RFC-1918) address prefixes were shared from the NAT devices to the Data Center Edges, which would then be announced up and summarized at the core router for advertisement to the public Interet.
Large-Scale Network Address Translation Layer
The Large-Scale Network Address Translation Layer (LSN or Large-Scale NAT) was a load balancing product aimed at service providers to perform NAT and CGNAT. In High Availability, the two cluster members were connected with the same 4x10Gb/s connections to each data center edge router with another 10Gb/s connection between each HA device utilizing Virtual Router Redundancy Protocol (VRRP) to designate a cluster leader in the event of a failover. The LSN clusters had the publicly routable IP prefixes configured into NAT groups and an access control list (ACL) of which backend servers were permitted to utilize which NAT group.

Above is a simplified diagram of the High Availability architecture using one of our cold storage clusters. At our legacy sites, there were two of these HA LSN clusters for the cold storage backend and the other backend services that used the large scale NAT load balancers operated on the same architecture.
Network Flow
From the backend service, when it requests an external service, the request would be sent to internal proxy services (owned by an SRE infrastructure team), creating a session on the proxy server. The proxy would forward the request to the LSN load balancer cluster, which was had access control lists for controlling traffic that in- and egressed them. Meeting the ACL requirement, a session for the proxy server was created on the load balancer, source address translated and then forwarded to the data center edge router and then to the intended destination over the Internet.
While simplicity was a goal for our network architecture, the number of hops for this sort of activity was atypical. With High Availability though, there was only one active load balancer per service. Understanding which device was active via VRRP would give us the piece of hardware that we needed to investigate any potential issues. High Availabilty configurations kept identical ACLs and translation pools for ease of failover and resumption of service.
Monitoring
We had our own home-grown monitoring system, using Python, that we could correlate traffic levels, alerts, and the like. With the new architecture, there would need to be some minor changes to this system but these were fairly simple code updates and git pushes once we had validated the fucntionality of the new architecture through Migration/Disaggregation. All of these changes would be made on the LSN load balancer dashboards.
Active Availability
Edge Network Layer
The edge network layers changed little - each load balancer would still be connected to each edge router. As we were trying to improve the runway for the amount of traffic, we increased the capacity of the link aggregation groups (LAGs) connected to each router. As cost controls were a driver of this architecture, particularly with our older hardware in older data centers, we ended up maxxing out the available capacity on those devices but still provided us the increases we needed based on forecasts. We knew that in the High Availability design, each load balancer was peered to each router with BGP via an IP assignment on the VRRP interface that controlled failover between primary and secondary devices but with Active Architecture, each individual load balancer would need to be peered to each edge router for full connectivity. This actually wound up being a lesson learned when we attempted our first migration (described further below).
Large Scale Network Address Translation Layer
To facilitate the active availability of each member of the LSN load balancing cluster, access control lists would include all member hosts for address translation, as in the high availability architecture, but the prefixes that we would use for translation out of RFC-1918 address space, would be unique to the individual load balancer. Given that we’re working with precious IPv4 address resources, we first went for splitting prefixes. Take for example, our cold storage transfer clusters - two clusters (four devices), each cluster having a /25 IPv4 prefix. To address this across four devices, we took each /25 prefix and instead broke it up into two /26 prefixes. Additionally, instead of having a physical link between the LSN load balancers managed by VRRP, it would instead be replaced by the LAGs connected to the edge routers and peered via BGP. The LSN load balancer would announce it’s own /26 translation pool to the edge router, which would then announce aggregated blocks to meet the Internet standard of a IPv4 /24 for the smallest prefix announcement.
| Hostname | HA NAT Prefix (IPv4) | AA NAT Prefix (IPv4) |
| lsn1-1 | 198.51.100.0/25 | 198.51.100.0/26 |
| lsn2-1 | same as above due to VRRP | 198.51.100.64/26 |
| lsn1-2 | 198.51.100.128/25 | 198.51.100.128/26 |
| lsn2-2 | same as above due to VRRP | 198.51.100.192/26 |

While we could cleanly subnet most of our most legacy devices, we were also turning up a new site with new LSN load balancers and had a preference at least for some class of geo (well, site, since all sites were based in the same country in North America) identity, if only for internal recognition. We still had to look for non-RFC-1918 prefixes available which left us taking a /25 vacant from a project over here, a consolidating a few /28s from over there, etc.
Network Flow
The new Active Architecture would not change the flow from backend to proxy to load balancer to edge, but provided density to the translation point. If the pools for address translation were identical across all device members, correlating issues could get very messy as we could have as many as four devices to hunt through. Coupled with our monitoring tools providing records of which device used which prefix, and having records of the session tables, we know which load balancers needed to investigate based upon the translated IP.
Monitoring
As I noted with in above, with Active Availability, there were minor changes that needed to be made to our monitoring system for the LSN load balancers. As clusters expanded, some hostnames changed so we needed to update those references to correctly pull data from all devices. Likewise, we also had monitors for the translated prefixes. As we subnetted the prefixes to smaller aggregates per device, we needed to update and add monitors for the smaller aggregate prefixes.
Migration/Disaggregation
All of the design work was completed, documents and configurations were written and reviewed, upgraded hardware was installed where called for and data center teams were cabling, we finally were ready to work out the details of migration.
For our newest greenfield site, we would deploy as normal. Even though our edge routers were functioning, services were slowly being turned on and routed from the local edge routers to one of our legacy sites via our internal backbone. The most catastrophic of failures in our design would be caught here. It, however; masked a problem we found when disaggregating hardware in our of the legacy sites.
Turning our attention to the legacy sites, as we were breaking high availability clusters, we went through a process keeping the active member running and disengaging the passive member. With the passive member removed from the cluster, we then would re-configure the translation prefix on the LSN load balancer, turn up BGP (while filtering prefix exchanges) between the LSN LB and the edge router and fianlly, release the prefix exchange. After watching the new LB take on traffic, we would then filter traffic on the edge router for the other LB to make the same changes to the LSN prefix lists. So far, our first migration at a legacy site was off to a great start.
When we began filtering the previously primary/active LSN load balancer though, we started seeing TCP failures from connections not being re-established. Causing unwanted impact, we rolled back our changes and placed the devices back into high availability service and took our design back. Each device had a loopback interface, configured with an IP address, used for internal DNS and the like. We used this IP for peering via BGP with the edge routers. Which means when we took a device offline and the BGP status left Established, the traffic did not have another route path to follow upstream, thus getting dropped on the floor. We mitigated this design oversight by adding a second loopback interface to each of the LSN load balancers (cleverly named loopback1) and used THAT for peering with the edges. loopback1 was unique to a cluster, so all of the LSN load balancers that processed the same traffic would share the address, so prefixes would be exchanged seamlessly between all LBs and all edge routers.
We attempted disaggregation again, same target cluster in same legacy site as before, now armed with the second loopback. When removing the passive member of the cluster, made the changes and turned it back up, we saw the same results - BGP Established, prefixes exchanged, traffic came to the new LSN load balancer. We then took the formerly primary member down - and this time, we saw more traffic migrate and fewer reports of performance issues from the backend service team. We made configurations and put the device back into service. Watching the monitoring boards, we saw the devices pick traffic back up and level off between each member as ECMP began to take hold for new sessions. It was a success. We could then repeat across the remaining clusters in our legacy sites and had a good path forward for the deployment of new devices.
Next Steps
With a functioning service, our next steps post-disaggregation we directed more towards optimizing the operation and management of the services. The first big instance was to separate the LSN load balancer clusters on the edge routers to their own unique virtual route forwarder (VRF), rather that keeping them in the standard Internet table (inet.0 for us JUNOS folks). The chief goal was for segmenting the traffic into the VRF out of the global table to ease management of the clusters and simplify routing back to the backend systems. Further down the road, we’d wanted to explore a software-based solution that could be run closer to commodity hardware rather than specialized network hardware. The benefits were two-fold: first, drive down the cost of capital expenditure from service provider class equipment and flatten that expenditure out over smaller cycles rather than punctuated with large purchases. Second, we could control more of our own destiny and aim to provide a lower toil solution.
Conclusion
Overall, it was a success. I had a lot (and I mean a lot) of help getting the design validated and we wouldn’t have been anywhere near as successful without that help. Unfortunately, while I was able to perform most of the disaggregation migrations, there were still one or two that needed to be done when I decided to leave the company. I’d have loved to have stayed to both see the completion of the migrations and work on the VRF implementations and help put the pieces together for the software-based solution but business is business.
tags: