Network Engineer – AI Data Center Networking
- AI
- TCP/IP
- Linux
- Fabric
- Devops
- VLANs
- Ansible
- Agile
- Kanban
- CCNA
- Configuration Management
- VPC
- BGP
- Kubernetes
- Python
You have an entrepreneurial spirit. You enjoy working as a part of well-knit teams. You value the team over the individual. You welcome diversity at work and within the greater community. You aren't afraid to take risks. You appreciate a growth path with your leadership team that journeys how you can grow inside and outside of the organization. You thrive upon continuing education programs that your company sponsors to strengthen your skills and for you to become a thought leader ahead of the industry curve.
You are excited about creating change because your skills can help the greater good of every customer, industry and community. We are hiring a talented<Job Title>to join our team. If you're excited to be part of a winning team, CirrusLabs (http://www.cirruslabs.io) is a great place to grow your career.
Network Engineer – AI Data Center Networking (High-Performance Infrastructure)
Location: LATAM
Engagement Type: Long-term Consulting / Contract
Position Summary
We are seeking a mid-levelNetwork Engineer to support and operate modern AI data center networking environments used for performance-sensitive workloads. This role is best suited for someone with strong hands-on experience in production network operations, solid L2/L3 and TCP/IP fundamentals, strong Linux comfort, and proven ability to troubleshoot end-to-end traffic issues across physical, switching, routing, and transport layers.
The engineer will supporthigh-speed switch-based fabric environments where traffic must be traced logically and physically from source to destination across multiple switching points. This is not a traditional CLI-heavy network engineering role focused on one-off device administration. Instead, the role requires an operator who understandstraffic paths, underlay behavior, observability, and repeatable network operations in modern leaf-spine environments.
The ideal candidate will be comfortable working in data center environments built aroundleaf-spine architectures, ECMP, VLAN/VRF segmentation, VXLAN underlay/overlay concepts, and performance-focused monitoring. The person should also be able to useautomation and repeatable playbooks to reduce manual switch operations, while collaborating effectively with DevOps and infrastructure teams.
This role is a strong fit for someone who can become productive quickly on core networking operations and can be trained within a short period on environment-specific tools such asNVIDIA Spectrum andNetQ.
Key Responsibilities
- Operate and support production data center network infrastructure, including switches, optics, cabling, transceivers, and physical connectivity.
- Monitor and troubleshoot end-to-end traffic paths across the network fabric from point A to point B.
- Support day-2 operations for modern data center networking environments using leaf-spine architectures, ECMP, VLANs, VRFs, and VXLAN-based overlay/underlay models.
- Investigate and resolve L1-L3 issues including interface errors, optics faults, cabling issues, forwarding failures, routing issues, and segmentation problems.
- Troubleshoot performance issues across the traffic path, including latency, packet loss, congestion, MTU mismatches, retransmissions, and ECMP pathing behavior.
- Use Linux CLI, system logs, and standard network troubleshooting tools to validate host and network behavior during incidents.
- Use monitoring, telemetry, and alerting tools to isolate faults and identify performance bottlenecks in production environments.
- Support observability practices for network health, path validation, and operational troubleshooting; experience with NetQ or similar tools is valuable.
- Work closely with compute, storage, platform, and DevOps teams to determine whether issues originate in the network, host, or application layer.
- Execute network changes, maintenance tasks, and operational validation activities with strong focus on stability, repeatability, and uptime.
- Contribute to network automation efforts using Ansible-like playbooks, templates, and basic scripting to reduce manual configuration tasks.
- Follow documented operational procedures, change controls, and Agile/Kanban ways of working.
- Contribute to continuous improvement through documentation, standardization, and operational best practices.
Must-Have Qualifications (simple to complex order)
- 3–5+ years of experience in data center or enterprise network operations/support in production environments.
- CCNA-level knowledge or equivalent hands-on networking foundation.
- Strong understanding ofL2/L3 networking and TCP/IP fundamentals.
- Proven troubleshooting ability acrossL1-L3, including physical, switching, routing, and forwarding issues.
- Experience supportingon-prem physical network infrastructure, including switches, optics, cabling, transceivers, and port-level fault isolation.
- StrongLinux operational comfort, including CLI, logs, interfaces, routing tables, and common troubleshooting tools.
- Strong experience withnetwork monitoring, observability, and alerting in production environments.
- Ability to troubleshootperformance issues across the traffic path, including latency, packet loss, congestion, MTU, retransmissions, and ECMP pathing.
- Understanding ofmodern data center architecture, including leaf-spine, ECMP, VLAN/VRF segmentation, and east-west traffic patterns.
- Familiarity withunderlay/overlay concepts in VXLAN environments and ability to support day-2 operations.
- Foundational understanding ofRDMA/RoCE concepts and why low-latency traffic behavior matters in AI/HPC-style environments.
- Working exposure toautomation/configuration management, including Ansible-like playbooks, repeatable execution, and basic scripting for network operations.
- Ability to work effectively withDevOps teams and align with existing tooling and workflows.
- Experience working inAgile/Kanban delivery environments.
- Basic awareness ofVPC-to-cloud connectivity concepts.
Nice-to-Have Qualifications (simple to complex order)
- Experience withNVIDIA NetQ or equivalent monitoring/telemetry platforms such as NMX.
- Exposure toNVIDIA Spectrum switching platforms or similar modern data center switching ecosystems.
- Experience withhybrid networking, including on-prem to cloud connectivity and routing/policy considerations.
- Production experience withEVPN-VXLAN fabrics, including MP-BGP EVPN, VNIs, anycast gateway, VRFs, and day-2 troubleshooting. BGP (Border gateway protocol)
- Strongerstructured automation practices, including templates, version control, peer review, and rollback discipline.
- Familiarity withKubernetes networking concepts, including CNI, overlays, and service networking.
- Experience withhost-side NIC (Network Interface card) tuning.
- Experience withhigh-performance or parallel storage networking.
- Prior exposure toAI/HPC networking environments, including rail-optimized, dual-plane, CLOS, or GPU-cluster traffic patterns.
- Exposure toGPU communication patterns such as NCCL.
- BasicPython capability for operational scripting and tooling.
Preferred Candidate Profile
The ideal candidate is astrong network operations engineer first, not a pure architect and not a DevOps platform engineer. They should be confident in troubleshooting production network issues, comfortable in Linux, effective with monitoring tools, and able to follow repeatable playbook-driven operational practices instead of relying on manual switch-by-switch CLI work.
They do not need to be an expert in NVIDIA technologies on day one, but they should have the technical foundation and learning agility to become productive quickly in a modern AI fabric environment.
Network Engineer – AI Data Center Networking · CirrusLabs