AI Training Workloads Network Deployment
Budget / Salary₹1,000,000–2,500,000
TypeFreelance project
LocationRemote
Posted1 hour ago
Technical Summary
Architecting, deploying, and validating a non-blocking 2-tier Spine-Leaf Clos fabric designed for distributed AI training workloads. The network utilizes BGP-EVPN with VXLAN overlay alongside end-to-end RoCEv2 lossless Ethernet parameters (PFC, ECN, DCBX) to ensure zero packet loss and sub-2 microsecond latency across NVIDIA GPU clusters.
Technical Details
Topology & Core Fabric
- Architecture: Non-blocking 2-tier Spine-Leaf Clos topology
- Underlay Routing: eBGP with IPv4 loopback reachability
- Overlay Control Plane: BGP-EVPN (Type-2 MAC/IP & Type-5 IP Prefix routes)
- Data Plane Encapsulation: VXLAN
- Global MTU: 9216 bytes (Jumbo Frames end-to-end)
QoS & Lossless Ethernet Mechanics (RoCEv2)
- Class of Service Mapping: CoS 3 reserved for RDMA traffic (DSCP 26 / AF31)
- Priority-Based Flow Control (PFC): IEEE 802.1Qbb enforced on CoS 3
- Explicit Congestion Notification (ECN): RFC 3168 RED marking
- K_min (Min Threshold): 150 KB
- K_max (Max Threshold): 1500 KB
- PFC Headroom Trigger: 3000 KB
- Parameter Negotiation: DCBX (IEEE 802.1Qaz)
Hardware & Platform Specifications
- Compute & Offload: NVIDIA A100/H100 GPUs, NVIDIA ConnectX-6/7 SmartNICs (GPUDirect RDMA / GDR)
- Switching Infrastructure: Arista EOS, Cisco NX-OS, and disaggregated NOS platforms
- Testing Infrastructure: Ixia IxNetwork traffic generators (line-rate 100G/200G stress testing)
Architecting, deploying, and validating a non-blocking 2-tier Spine-Leaf Clos fabric designed for distributed AI training workloads. The network utilizes BGP-EVPN with VXLAN overlay alongside end-to-end RoCEv2 lossless Ethernet parameters (PFC, ECN, DCBX) to ensure zero packet loss and sub-2 microsecond latency across NVIDIA GPU clusters.
Technical Details
Topology & Core Fabric
- Architecture: Non-blocking 2-tier Spine-Leaf Clos topology
- Underlay Routing: eBGP with IPv4 loopback reachability
- Overlay Control Plane: BGP-EVPN (Type-2 MAC/IP & Type-5 IP Prefix routes)
- Data Plane Encapsulation: VXLAN
- Global MTU: 9216 bytes (Jumbo Frames end-to-end)
QoS & Lossless Ethernet Mechanics (RoCEv2)
- Class of Service Mapping: CoS 3 reserved for RDMA traffic (DSCP 26 / AF31)
- Priority-Based Flow Control (PFC): IEEE 802.1Qbb enforced on CoS 3
- Explicit Congestion Notification (ECN): RFC 3168 RED marking
- K_min (Min Threshold): 150 KB
- K_max (Max Threshold): 1500 KB
- PFC Headroom Trigger: 3000 KB
- Parameter Negotiation: DCBX (IEEE 802.1Qaz)
Hardware & Platform Specifications
- Compute & Offload: NVIDIA A100/H100 GPUs, NVIDIA ConnectX-6/7 SmartNICs (GPUDirect RDMA / GDR)
- Switching Infrastructure: Arista EOS, Cisco NX-OS, and disaggregated NOS platforms
- Testing Infrastructure: Ixia IxNetwork traffic generators (line-rate 100G/200G stress testing)
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.