IW
RL Lab Network Architect
Info Way Solutions LLC
๐จ๐ฆ Canada
On-site
2 weeks ago
- Fabric
- BGP
- Python
- YAML
- JSON
- Git
- CI/CD
- Ansible
- Terraform
- Arista
- Cisco
- iOS
- Juniper
- OSPF
- VLANs
- SNMP
- Linux
- Tcpdump
- Wireshark
- Node.js
- DHCP
2 weeks ago
Network Engineer (Config-as-Code)
Job Title: RL Lab Network Architect
Location: Menlo Park (100% onsite at Client office)
Need W2 candidate who is ready to work on Client FT
About the Team
You will be joining the Engineering Labs group inside Network Infrastructure Services. The team builds and runs the lab network that hardware, firmware, and network engineering teams depend on to do their work: spine-leaf lab fabrics across multiple sites, the handoff between those fabrics and the wider corporate network, and the provisioning, access control, and telemetry that keep them usable.
The defining thing about this team is that the network is managed as code. Configuration lives in a version-controlled repository, changes go through peer code review and automated validation, and deployment happens through an automated pipeline with canaries, health checks, and a rollback path. Nobody logs into a switch to make a change stick. If that sounds like how you already want to work, you will fit here. If you have only ever worked device by device, this will be a hard adjustment.
The work is a mix of design, software engineering, and operations. Expect to spend real time writing and reviewing code, and real time on incidents and fleet hygiene. You will be part of an oncall rotation (within regular business hours) and you will own fabric design and deployments end to end rather than being handed tickets.
What You Will Be Doing
A realistic picture of the work, based on what the team actually ships:
- Designing and iterating on lab fabric traffic paths. Taking an east-west traffic optimization design through successive versions: aligning physical ports, adjusting BGP peering groups and route policy, adding route suspenders to stop unwanted routes leaking, and moving port-role definitions out of policy files and into intent data where they belong.
- Writing and refactoring the templates that generate device config. Adding new data structures to the template framework so a whole class of sites can render a feature, rather than special-casing one site at a time.
- Managing access control as code. Adding service groups and rules so specific lab and compute environments can talk to each other, in both directions, and proving the flow works before it ships.
- Large-scale fleet cleanup. Removing references to decommissioned devices across policy files, templates, exception files, and cluster definitions, site by site, often dozens of devices at a time. Retiring orphaned templates and dead platform blocks. This is refactoring on a live network and it is a standing part of the job, not a one-off project.
- Onboarding new devices and racks through the provisioning and IP allocation pipeline, and debugging it when a provisioning step fails.
- Building the dashboards and alerting that tell you whether the fabric is healthy, and being the person who gets paged when it is not.
By 30 days
- Development environment set up, and you can render a device's config locally and diff it against what is live
- First change landed: small, reviewed, and deployed through the normal pipeline end to end
- You can explain the lab fabric topology and the roles in it, and read an existing template and say what it will produce
Shipping changes independently across intent data, policy files, and templates, with test plans that show before and after renders
- Reviewing other engineers' changes and catching real problems, including scoping and rollback gaps
- Completed at least one multi-site cleanup or migration wave without a self-inflicted outage
- Onboarded into the oncall rotation and handled incidents with support
- Comfortable stopping a rollout mid-flight and rolling back without asking for help
- Owning a fabric or a workstream: you are the person others come to for it, and its design docs are current because you keep them current
- Delivered a design change of your own, not just executed someone else's, and took it from proposal through review, staged rollout, and verification
- Removed toil rather than absorbing it: at least one thing the team used to do by hand is now automated, templated, or validated automatically
- Oncall without a safety net, and the runbooks are better than when you arrived
- Reduced risk in a measurable way: fewer manual touches, less config drift found by audits, or better coverage on the monitoring for a failure mode that used to be silent
- Minimum 3 years of network engineering experience in a lab, campus, or data center environment
- Hands-on experience authoring and shipping network configuration as code through a version-controlled repository, with peer code review, automated validation, and staged deployment (HARD REQUIREMENT). Candidates who only configure devices through a CLI, GUI, or controller UI do not meet this bar.
- Ability to write and maintain production-quality code, not just scripts. Python (object-oriented) required. Working knowledge of Jinja2 templating, YAML/JSON config schemas, and a Python-based configuration DSL (HARD REQUIREMENT).
- Repository management and code review experience (e.g., Git or equivalent industry-standard tools), including branching, rebasing, stacked changes, and revert
- Deployment automation and continuous delivery experience (e.g., Ansible, Terraform, NAPALM, or similar; client uses proprietary equivalents)
- Experience with intent-based configuration: keeping device intent data separate from rendering templates, and generating vendor-specific config from a single source of truth
- Experience managing ACL and firewall policy as code, rendered out to more than one vendor syntax
- Multi-vendor experience across Arista EOS, Cisco IOS-XE/NX-OS, and Juniper JunOS
- Working knowledge of spine-leaf fabric design, including leaf, spine, border leaf, gateway, and access roles
- Routing and switching fundamentals: BGP, OSPF, VXLAN/EVPN, VLANs, IRB interfaces, prefix-lists, route policy, ECMP
- Solid IPv4 and IPv6, including SLAAC, DHCPv4/DHCPv6, and CIDR planning
- Experience with IPAM tooling and programmatic IP and prefix allocation (e.g., BlueCat or similar; client uses proprietary equivalents)
- Familiarity with telemetry and monitoring: SNMP, gNMI, NetFlow/IPFIX, sFlow, syslog, ICMP probes
- Comfortable working from a Linux development environment and a source-control CLI day to day
- Exposure to open network operating systems and whitebox switching
- Experience with 802.1X / MAB, TACACS+, or RADIUS in a lab or campus context
- Prior work supporting hardware or firmware engineering teams in a test lab
- Configuration as Code: Own network configuration in the code repository rather than on the device. Author and review changes to intent data, config policy files, and Jinja2 device templates. Build template logic that renders correct vendor CLI for Arista, Cisco, and Juniper platforms from one shared intent model. Keep templates parameterized and reusable across sites instead of hand-writing per-device config. Never hardcode IPs or hostnames: pull them from the source of truth or from service discovery.
- Code Quality and Review: Write Python 3 with type hints on function signatures. Add unit tests for new logic, and mock device interactions so tests run offline. Cover the edge cases that actually happen: device unreachable, empty result set, malformed input. Run linting and the pre-submit checks before every submission. Submit clean, single-purpose changes with a real test plan that shows before and after rendered config, and call out breaking changes in the summary. Review other engineers' changes to the same standard.
- Validation Before Deployment: Render locally and diff against what is live before anything ships. Use dry-run and preview modes, compare generated config to the archived config, and confirm the change touches only the devices intended. Validate ACL changes by tracing the specific flow, source, destination, protocol and port, rather than assuming. Prove a change on one device before rolling it to a role, site, or the whole fleet.
- Staged Deployment and Rollback: Push config through the automated deployment path with canaries, slow-rolls, and health checks rather than by hand. Use feature-flag style gating to slow-roll template changes that touch many devices. Know how to stop a rollout mid-flight using push-disable controls and killswitches scoped by role, region, or site, and document the rollback path in the change itself. Guard destructive operations behind confirmation prompts and dry-run flags, and scope fleet-wide operations with explicit target limiting.
- Lab Fabric Design: Design and extend spine-leaf lab fabrics and their handoff to the wider corporate network. Plan BGP peering groups, route policy, prefix-lists and route suspenders, IRB interfaces, VLAN and sub-interface layout, and port roles. Handle east-west traffic paths between lab fabric and the systems labs depend on. Plan IPv4 and IPv6 addressing for the fabric and keep it registered in IPAM.
- Device Onboarding and Fleet Hygiene:Bring new devices and racks online through the provisioning and IP allocation pipeline. Sweep and remove references to decommissioned devices across policy files, templates, port and node exception files, and cluster definitions, at scale and without breaking neighbours. Retire orphaned templates and dead platform blocks. This is a large, ongoing part of the role: it is refactoring work on a live network and it needs the same care as any other code change.
- Access Control as Code: Define and maintain ACL policy in code: named network groups, service groups, and rule terms that render to Juniper, Cisco, and Arista formats. Wire dynamic network definitions to authoritative sources so groups stay current without manual edits. Support centralized authentication and port-based access control for lab devices. Never commit keys, credentials, or tokens, and keep destructive network operations behind access control.
- Monitoring and Telemetry: Build and maintain time-series dashboards and SLI definitions for reachability, interface utilization, throughput, and cross-site RTT and loss. Query structured event data to debug authentication, DHCP, port state, and provisioning failures. Add or update monitoring and alerting whenever a change introduces a new failure mode. Report against per-role availability targets for leaf, spine, border leaf, gateway, and access switches.
- Troubleshooting and Support: Support the engineers who depend on the labs. Troubleshoot to root cause on connectivity, routing, reachability, and access issues, including packet capture and log analysis. Take part in an oncall rotation, work incidents to resolution, and feed what you learn back into automation, monitoring, and runbooks so the same issue does not come back.
- Documentation and Enablement: Keep design docs current for the fabrics you own. Update oncall runbooks when operational procedures change, and update CLI help text and internal tooling docs when you change behaviour other engineers depend on. Make your automation usable by the rest of the team, not just by you.
RL Lab Network Architect ยท Info Way Solutions LLC