In the previous article, we reviewed the evolution of overlay networks in the AI era, assessing how DPU-based architectures are reinventing the network infrastructure for AI cluster frontend traffic. We covered the journey from kernel networking to DPU/ SmartNIC-based overlay tunnels between host and gateway node.
This publication looks at the implementation of an Overlay Tunnel Gateway Solution based on Juniper PTX Series routers. Although various overlay tunnelling technologies are available e.g. IP over IP, MPLS over GRE, MPLS over UDP, VxLAN (EVPN Type-5) and latest addition is SRv6, but in this blog we will focus on VxLAN (EVPN Type-5) overlay tunnels.
We will cover different design approaches for large-scale front-end fabric for providing connectivity to thousands of DPU nodes, and we will also cover multi-tenancy and integration of management fabric and front fabrics.
Next article will cover North-South and East-West connectivity, firewall integration, redundancy, and DCI over MPLS/RSVP-TE.
Overlay tunnels from DPUs to network gateway can terminate:
at super spine layer (5-stage Clos)
at spine layer in (3-stage Clos architectures),
or it can be a separate border node (without super spine or spine role).
In this blog, we are using the super-spine / spine layer for tunnel termination, and the same device will be used for Data Center Interconnect (DCI) and for controlled public accessibility. We will discuss the reasoning behind this design choice further down the road.
The Juniper PTX Series routers, based on Juniper Express-4 and Express-5 silicon, support tunnel aggregator functionality.
Available in a variety of fixed and modular chassis, these platforms can deliver exceptionally high packet processing rates (billions of packets per second) with ultra-low latency and deep buffer to absorb microbursts. The high port density combined with low power consumption per port makes these routers a perfect choice for power-conscious data centers environments.
With tunnel termination capacity at scale, the PTX Series supports the AI Front-End Network use case where DPU hosted workloads require isolation via overlay networking while maintaining high performance and low latency packet processing demand.
Other features such as line-rate MACsec and native coherent optics support across all interfaces provide secure long-haul connectivity without the need for external transponders. Combined with carrier-grade routing features (bandwidth management and congestion avoidance), make PTX Series Routers an ideal choice for DCI edge roles as well.
Thus, a single PTX Router provides dual functions of tunnel aggregator and DCI edge devices, eliminating the need for separate network devices to perform distinct functions and yielding substantial reductions in capital and operational costs.
HPE Networking has recently announced Juniper PTX12000: 8 and 12 slots modular chassis offering 432x and 648x 800GE ports respectively. Like their predecessors, newly added models also fit well for low latency Front End tunnel aggregation and DCI edge Roles but with much higher port density. More details about PTX120008 can be found in the blog post.
The architecture relies on EVPN Type-5 routes (IPv4/IPv6) propagation and uses VXLAN tunnels for forwarding plane. EVPN Type-5 routes carry along VNI and Route Target which are used to ensure isolation between different tenants at tunnel gateways and software routers running at DPUs. If a reader needs to know EVPN Type-5 implementation details in Junos/Evo, then please visit these two links:
Various routing / switching platforms from Juniper portfolio supports next-hop-based dynamic tunnels. This feature allows overlay networking over IP transport networks, and various types of overlay encapsulation are supported, e.g. (IP over IP, MPLS over UDP and VxLAN etc).
The following text only deals with Juniper PTX series routers (based on express-4/5 chipset) running Junos EVO. Other platforms from Juniper portfolio may or may not match this description.
Once the protocol next-hop (PNH) of a remote prefix is resolved, then Routing Protocols Damon (RPD) creates tunnel composite nexthop (TCNH) with forwarding nexthop from inet.0 (default) routing table. TCNH is chained to Indirect Next-Hop (INH). INH is then further mapped to forwarding next-hop (FNH) using inet.0 routing table. Describing various next-hop types is not in the scope of this document, however more information can be found in this blog post.
In the EVPN Type-5 scenario, once a remote BGP peer (i.e. VTEP) peer sends IPv4 and IPv6 prefixes, then for each address family one TCNH is consumed. We have tested 128K VTEPs with accumulative 2.048 million IPv4 and IPv6 (1.024 million each) EVPN Type-5 Prefixes. It created 256K TCNH (1 for each address family per VTEP) and over all resources consumption was acceptable.
The architecture has three separate layers, each with its own function within the architecture.
The underlay layer uses hop-by-hop eBGP IPv6 to connect all devices in the fabric. Underlay, eBGP uses IPv6 addresses between directly connected neighbors. These sessions are used to advertise loopback addresses. However, the underlay can use either BGP unnumbered with IPv6 link-local addresses or numbered IPv6 /127 subnets. Unnumbered is acceptable as long as all devices support link-local addresses. The purpose of the underlay is to provide loopback addresses with reachability between VTEPs.
The overlay BGP (EVPN signaling) sessions are used to exchange EVPN Type-5 (IPv4/IPv6) routes. Overlay BGP sessions can be hop-by-hop eBGP, or those can be iBGP between DPUs and tunnel aggregator. If the iBGP model is selected for overlay, then Route Reflectors are mandatory for this massively scaled fabric. Overlay BGP sessions need to be configured with a loopback address (IPv4 or IPv6) of each corresponding device.
The data plane uses VxLAN tunnels carrying tenant data with multi-tenancy isolation. VTEP endpoints exist only at the Tunnel Aggregator and at DPU devices. Intermediate forwarding devices (ToRs, spines) forward IP packets through the network using underlay routing but perform no VxLAN encapsulation or decapsulation.
The understanding of the control and data plane flow reveals why this architecture delivers scalability and low latency.
The DPU sends tenant prefixes that are associated with a VNI and a Route Target via EVPN Type-5 route advertisement in the control plane. The Tunnel Aggregator uses the Route Target and imports those routes into corresponding TENANT_VRF.
On the data plane forwarding, when the Tunnel Aggregator receives a packet that is destined for a tenant prefix:
It performs a route lookup which indicates which VNI to use and to which DPU VTEP address is the next hop.
The Tunnel Aggregator encapsulates this packet in VxLAN header with the corresponding VNI and sends this to the DPU VTEP using underlay IPv6 routing for forwarding.
The ToRs and Spines forward these packets based on the destination VTEP address without inspecting or modifying the encapsulation;
the DPU decapsulates the packet based on the VNI that was specified, and forwards this to the tenant's workloads.
This architecture performs VxLAN encapsulation and decapsulation only on the edges of the network (i.e., at the Tunnel Gateway and at the DPU), resulting in low latency; there are no additional hops that need to perform these operations within the fabric.
The underlay, VTEP addressing, and tenant workloads can use different address families, providing significant operational flexibility.
The underlay runs IPv6, either using link-local addresses via BGP unnumbered or numbered /127 subnets.
VTEP loopbacks can be either IPv4 or IPv6. When using IPv4 VTEP loopbacks over an IPv6 underlay, RFC 5549 / 8950 allows advertising the IPv4 addresses with IPv6 next-hops. In our testbed, we have used IPv4 based loopback addresses because Express-4 based Juniper PTX products only support IPv4 VTEPs, but Express-5 based platforms support both IPv4 and IPv6 VTEPs.
The following config snippet shows how to achieve RFC 5549/ 8950 functionality in Junos / Junos EVO.
group underlay {
neighbor fd7e:92a1:cb3d::3e:10 {
family inet {
unicast {
extended-nexthop;
}
}
peer-as 65101;
}
rfc8950-compliant;
}
When it comes to tenant workloads or service prefixes, they don't really depend on the underlay or VTEP addressing choices.
This means they can use IPv4, IPv6, or even both. For example, you might have a setup where the underlay uses IPv6, but the VTEP loopbacks use IPv4, and yet the tenant workloads still support both IPv4 and IPv6. This kind of flexibility is really useful because it allows network infrastructure team to update the underlay to IPv6 without disrupting any existing IPv4 based workload / applications, and at the same time, they can offer support for multiple address families to their tenants.
Now that we have a good grasp of the basic ideas behind the architecture, such as having two BGP sessions i.e underlay and overlay, putting VTEPs in the right places, and being flexible with address families, let's dive into the real-world decisions we need to make when setting up the underlay network.
The main goal of the underlay network is straightforward: it's supposed to make sure that devices in each layer can communicate using loopback IPs. But when we are dealing with a large-scale network, then IP address management / assignment to each link and underlay BGP configuration can be a cumbersome effort.
The scale of a large AI cluster exposes a critical problem with traditional numbered BGP peering. With thousands of DPUs connecting to hundreds of ToR switches, you need tens of thousands of link IP addresses just for inter-switch connections. This consumes multiple RFC 1918 private address ranges and creates an enormous configuration burden. Every interface must be manually configured with a unique IP address, tracked in IP address management systems, and documented for troubleshooting.
We have two options to either use BGP unnumbered or BGP numbered addresses. Let's dive deep into each option and compare the pros and cons of each approach.
Option 1: BGP Unnumbered with IPv6 Link-Local Addresses¶
BGP unnumbered, supported on Junos OS Release 21. 21.1R1 and later (Junos BGP Unnumbered Configuration Guide), eliminates the need of explicit IP address configuration on physical interfaces while still providing full underlay functionality.
The device auto-generates a link-local IPv6 address (on each interface with family inet6 enabled) on every interface; these addresses use the MAC address to generate IPv6 link local addresses. IPv6 Router Advertisements (RAs) are sent on fabric interfaces, and neighbors perform IPv6 Neighbor Discovery (ND, RFC 4861) resolution of link-local IPv6 addresses.
BGP peer auto-discovery. Each device has a list of allowed AS numbers instead of configuring explicit IP addresses and AS numbers for each BGP neighbor. When a device receives an RA advertisement from a directly connected neighbor, it attempts to establish a BGP session using the auto-generated link-local IPv6 address. When the AS number of the peer matches one of the ranges specified in the allow list, the session is accepted.
Configurations are vastly simplified. Instead of managing thousands of /30 or /31 subnets, family inet6 is enabled on all fabric interfaces in your fabric topology. An AS number allow list is populated with the AS numbers of peers that can auto-peer with each other. Router advertisement is enabled on fabric interfaces. A single BGP group is defined with peer-auto-discovery enabled. Export policies are defined for advertising loopback addresses over BGP sessions. In a large-scale deployment, this reduces configuration from tens of thousands of IP addresses and explicit BGP in neighbor statements.
Option 2: IPv6 Underlay with Explicit /127 Addressing¶
A second strategy retains explicit addressing control with an IPv6 data plane (Juniper AI DC EVPN Multi-tenancy Design). Explicit /127 IPv6 prefixes are assigned to point-to-point links yielding exactly two usable addresses without wasting a /64 for each link.
Each physical interface is assigned an explicit IPv6 /127 address from a dedicated pool. Each device has a loopback address (IPv4 or IPv6) for VTEP establishment. BGP sessions are explicitly set up using these /127 addresses. The underlay eBGP session transmits IPv6 unicast routes to establish loopback reachability.
The method enables complete visibility. Network operators know exactly which IP is assigned to each interface. Troubleshooting is easier with addressable pings and traceroutes with known source IPs. IPAM solutions can track assigned IPs. Some customers require explicit assigning for regulatory reasons. Multi-vendor fabrics benefit from this approach since all vendors might not support BGP unnumbered with the same effectiveness.
The complexity of configuration is still a challenge. At a huge scale, you are looking at tens of thousands of interfaces with IPv6 addresses as well as loopback addresses for each device. This is mitigated in fabric management solutions. In our testbed, we have numbered IPv6 underlay address.
The operational challenge of numbered addressing is greatly mitigated through fabric management platforms such as HPE Networking Apstra (HPE Networking Apstra Data Center Design Guide). Intentional networking changes the game for large scale fabrics.
Apstra does IP address planning, device configuration and day zero through day one scale operations, instead of deploying a ZTP solution for tens of thousands of interfaces one by one for a network engineer. A network architect defines the intent, the desired topology, AS number scheme, and IP address scheme. Apstra converts this into device specific configs and implements it with ZTP. Interface IPs are precomputed and assigned when the device is provisioned.
For customers who want to see numbered IPv6 addressing for compliance or visibility reasons, Apstra bridging the gap offers the operational simplicity of BGP unnumbered (no manual assignment of IPs for deployment) while offering the visibility that numbered addressing (traceability of IPs, integration with IPAM solution, interface level troubleshooting) provides.
BGP unnumbered simplify the provision with minimal dev config and no physical interfaces addressing to deal with. There is no need for readdressing for topological changes. It has a common config for all devices facilitating mass deployments at scale.
Numbered IPv6 has visibility with interface addresses in routing tables and in packet captures. You can ping individual addresses and do traceroutes with known addresses. IPAM tools can pinpoint precise assignments. Some admin policies require explicit addressing.
The bottom line: both approaches enable the same EVPN Type-5 overlay architecture detailed in the High-Level Architecture Overview. The decision is purely operational, not architectural, and life cycle management at scale is easy with Apstra.
As stated in the High-Level Architecture Overview, EVPN Type-5 routes are distributed through the network topology hop-by-hop using multi-hop eBGP peering. At each layer we need ensure next-hop for EVPN Type-5 route is maintained, this can be achieved by adding overlay eBGP group using "multihop no-nexthop-change"
An alternate approach is to use tunnel aggregators as Route Reflectors (RR) and each DPU would be RR client, while EVPN signaling is configured between RRs and their clients.
We have seen that both architectures are adapted by different customers in host-based overlay routing architecture. Each approach has its pros and cons e.g in Hop-by-Hop eBGP signaling each Spine and ToR switch will get copy of each EVPN type-5 (subject to AS path loop prevention mechanism) and overlay route (EVPN Type-5) scaling will be subject to allowed scaling limit of Spine and ToR hardware models. In the RR approach, careful consideration would be required on how many clients have peered with a RR set.
Management Fabric Integration with Tunnel Aggregator¶
A dedicated management network is required for AI Compute Infrastructure out-of-band (OOB) connectivity. The management network connects to the management ports of all data center servers with DPUs that need to be accessed remotely for configuration, monitoring, troubleshooting, and firmware updates.
On tunnel aggregator, tenant isolation is required, and for that purpose each tenant will have its own management and front end VRF. We will explore per-tenant intra-VRF communication and management fabric connectivity with the tunnel aggregator in the next section.
Management fabric is EVPN-VxLAN based Clos network but, unlike the Front End Network, serves a different purpose. Each physical server management NIC connects to a standard ToR switches in the management fabric. The management ToRs connect to management spines which peer with the tunnel aggregator. If multi-tenancy is required for out-of-band management infrastructure, then multi-tenancy must be maintained throughout the management infrastructure. Hence, management of NICs on compute nodes cannot run routing stack, so overlay tunnels cannot be established directly to the management of NICs.
Option-1: Inter-AS Option A Handover Between Tunnel Aggregator and Management Fabric Border node.¶
It requires separate eBGP sessions (Inter-AS Option A style) between each tenant's management VRF on the tunnel aggregator and the management fabric border device and additional config is required for each VRF Inter-AS Option-A hand over towards Management Fabric Border node as well for route manipulation on the tunnel aggregator.
Option-2: EVPN Signaling Between Tunnel Aggregator and Management Fabric Border Node.¶
Instead of per-tenant eBGP sessions, one single eBGP session with EVPN signaling between the management of fabric border devices and the tunnel aggregator.
The management fabric advertises tenant management prefixes as EVPN Type-5 routes, each tagged with its appropriate VNI and Route Target. For instance, Tenant A's management network is 192.168.254.4/32.
The management spine advertises this as an EVPN Type-5 route with VNI 100 and Route Target 1:100.
On the tunnel aggregator, Tenant A's management VRF imports routes with the Route Target of 1:100. This automatically imports the route into the correct VRF based on matching Route Targets.
We have tested both approaches in test bed and witnessed both approaches in production environments as well. Each approach has its own pros and cons e.g:-
Option-1: Requires more configuration between management fabric border nodes and tunnel aggregators. It also requires route manipulation on tunnel aggregator so that Front End VRF and Management VRF on tunnel aggregator have access to each other's routes. Beside config overhead once Front End VRF (EVPN type-5) routes are leaked into Management VRF, it will consume additional tunnel composite next hops.
Option-2: It's easier to implement as very few configs are required for each tenant as compared to approach-1, but management fabric will establish direct VTEPs with DPUs connected with Front End Fabric. In large deployments (16K+ DPUs), VTEP scale limits on management border nodes should be considered.
In front end fabric hop-by-hop eBGP (EVPN signaling) is configured to exchange type-5 routes between DPUs and Tunnel Aggregator.
All eBGP sessions with EVPN signaling are configured as multi-hop using lookback interface IP addresses. Every tenant has 2 VRFs at the tunnel aggregator, i.e. management (TENANT_X_MGMT_VRF) and Front End (TENANT_X_ FE_VRF).
For management fabric connectivity with the tunnel aggregators, we have tested both approaches described in section above Management Fabric Topology. For sake of brevity traffic flows and verifications for "Option-1" is only presented in this blog i.e. management fabric has Inter-AS Option-A connectivity towards tunnel aggregator for each TENANT_X_MGMT_VRF.
The tunnel aggregator establishes VxLAN tunnels directly with DPU and for cross fabric connectivity route manipulation is required on tunnel aggregator.
1. EVPN Type-5 prefixes (i.e 192.168.254.3/32 and 2001::100/128 ) are sent by DPU and arrives at the tunnel aggregator with RT value target:2:200 and VNI 200.
2. RT extended community is examined for each route.
3. Routes are copied into matching VRFs i.e TENANT_1_MGMT_VRF & TENANT_1_FE_VRF
4. Tunnel aggregator's TENANT_1_MGMT_VRF further re-advertise these routes to management fabric TENANT_1_MGMT_VRF via eBGP session.
5. Management fabric TENANT_1_MGMT_VRF routes (192.168.254.1/32 and 2001:102/128) arrive at Tunnel aggregator TENANT_1_MGMT_VRF via eBGP session.
6. Tunnel aggregator TENANT_1_MGMT_VRF converts these IP prefixes into EVPN prefixes and place into BGP-EVPN.RIB.0.
7. EVPN prefixes placed into BGP-EVPN.RIB.0 are further advertised towards Front End fabric via EVPN based control plane , along with RT and VNI of TENANT_1_MGMT_VRF.
8. Software router running on DPU receive these routes and import into respective VRF i.e TENANT_1_FE_VRF.
9. Now management host connected on management fabric can communicate with workload hosted in DPU based software router.
DUT is PTX10001-36MR running Junos EVO (25.2R1-S2.3-EVO). Let's examine the control plane route exchange before verifying forwarding plane.
DUT> show route 192.168.254.3/32 ## Route installed into mgmt.inet.0 and tan.inet.0
mgmt.inet.0: 4 destinations, 4 routes (4 active, 0 holddown, 0 hidden)
@ = Routing Use Only, # = Forwarding Use Only
+ = Active Route, - = Last Active, * = Both
192.168.254.3/32 *[EVPN/170] 1d 00:49:02
> to fd7e:92a1:cb3d::3f:10 via et-0/0/6.0
to fd7e:92a1:cb3d::3e:10 via et-0/0/2.0
tan.inet.0: 1 destinations, 1 routes (1 active, 0 holddown, 0 hidden)
@ = Routing Use Only, # = Forwarding Use Only
+ = Active Route, - = Last Active, * = Both
192.168.254.3/32 *[EVPN/170] 1d 00:49:02
> to fd7e:92a1:cb3d::3f:10 via et-0/0/6.0
to fd7e:92a1:cb3d::3e:10 via et-0/0/2.0
DUT> show route 2001::100/128
mgmt.inet6.0: 6 destinations, 6 routes (6 active, 0 holddown, 0 hidden)
@ = Routing Use Only, # = Forwarding Use Only
+ = Active Route, - = Last Active, * = Both
2001::100/128 *[EVPN/170] 1d 00:51:40
> to fd7e:92a1:cb3d::3f:10 via et-0/0/6.0
to fd7e:92a1:cb3d::3e:10 via et-0/0/2.0
tan.inet6.0: 1 destinations, 1 routes (1 active, 0 holddown, 0 hidden)
@ = Routing Use Only, # = Forwarding Use Only
+ = Active Route, - = Last Active, * = Both
2001::100/128 *[EVPN/170] 1d 00:51:40
> to fd7e:92a1:cb3d::3f:10 via et-0/0/6.0
to fd7e:92a1:cb3d::3e:10 via et-0/0/2.0
DUT> show route 2001::100/128 ## Route installed into mgmt.inet6.0 and tan.inet6.0
mgmt.inet6.0: 6 destinations, 6 routes (6 active, 0 holddown, 0 hidden)
@ = Routing Use Only, # = Forwarding Use Only
+ = Active Route, - = Last Active, * = Both
2001::100/128 *[EVPN/170] 1d 00:51:40
> to fd7e:92a1:cb3d::3f:10 via et-0/0/6.0
to fd7e:92a1:cb3d::3e:10 via et-0/0/2.0
tan.inet6.0: 1 destinations, 1 routes (1 active, 0 holddown, 0 hidden)
@ = Routing Use Only, # = Forwarding Use Only
+ = Active Route, - = Last Active, * = Both
2001::100/128 *[EVPN/170] 1d 00:51:40
> to fd7e:92a1:cb3d::3f:10 via et-0/0/6.0
to fd7e:92a1:cb3d::3e:10 via et-0/0/2.0
Let's verify route exchange between Management fabric to Tunnel Aggregator (DUT)
DUT> show platform pfetokend type5-tunnel-db-information
Source Ip Destination Ip Source Mac Destination Mac Token RefCount
10.187.0.78 10.38.96.16 9c:5a:80:30:a9:66 9c:5a:80:30:d6:66 2 4
DUT> show platform idmd summary TYPE5
Idmd Server Chunk Allocation Statistics:
Name Clients ChunkSize Total Free Global Local GlobalUsage
TYPE5 1 100 640 639 1 0 1 ## One VTEP is created
Let's verify tunnel composite next-hops on DUT.
Hence remote VTEP is sending IPv4 and IPv6 EVPN Type-5 prefixes and each address family consumes 1 TCNH.
Part 2 addressed the foundational architecture for EVPN Type-5 tunnel aggregation at scale. We explored the BGP topology connectivity design using unnumbered BGP and hop-by-hop EVPN signaling, how management and frond fabrics integrate, and the multi-tenancy VRF architecture which supports strict tenant isolation. We also covered forwarding plane mechanisms involved in East-West traffic, i.e. Management Fabric and Front Fabric, and vice versa.
Part 3 will cover external connectivity (Internet access, cloud access, DCI for cross-data center connectivity).
During the solution writing and validation process, I had many candid discussions and troubleshooting sessions with Bala Murali Krishna Sanka (Test Engineer), Rajesh Pillai and they heled me during various trouble shooting session. Kaliraj Vairavakkalai (Distinguished Engineer) also provided insight on BGP overlay design approaches.
Special thanks to Wen Lin (Distinguished Engineer) for reviewing the document contents. Her feedback on design approaches for management-fabric and front-end fabric integration was very crucial.
Acknowledgment to Sunessh Babu (Principal TME) and Pezhmon Sadjady (Test Engineer) for sharing scale testing results which are included in section above "Tunnel Composite Next Hop"