
Platform Engineer
Base-2 Solutions, LLC
๐บ๐ธ United States
10 hours ago
- Kubernetes
- AI
- IaC
- Terraform
- Ansible
- Bash
- Python
- CI/CD
- GitLab CI/CD
- Linux
- RHEL
- Ubuntu
- Oracle
- Docker
- AI/ML
- Argo
- Airflow
- Kubeflow
- REST API
- Server Components
- AWS
- Slurm
- CCNA
- GSEC
- CCNP
- CISA
- CISSP
- GCIH
- CCSP
10 hours ago
Position Summary
Base-2 Solutions is seeking a Platform Engineer to design, develop, and optimize Kubernetes platforms supporting enterprise artificial intelligence capabilities for mission customers. The role requires deep expertise in operating systems, hardware, GPUs, and high-speed networking and is performed 100% on-site at the customer site in Bethesda, Maryland, at the Intelligence Community Campus.
Essential Duties and Responsibilities
- Design, configure, and maintain enterprise Kubernetes platforms and collaborate with multidisciplinary teams to optimize Kubernetes architecture for performance, efficiency, and feature requirements.
- Develop and manage Infrastructure as Code using tools such as Terraform, Salt, Ansible, Bash, Python, or similar frameworks.
- Collaborate with development teams to design and implement secure, automated, and repeatable CI/CD pipelines, including GitLab CI/CD.
- Troubleshoot complex systems issues across cloud, network, and platform layers.
- Maintain technical documentation, architectural specifications, and Linux best practices; support Authority to Operate activities and compliance with federal security standards.
Required Qualifications
- Platform Engineering or Systems Engineering experience meeting an applicable education and experience pathway in the equivalency section.
- Strong expertise with Linux distributions, including RHEL, Ubuntu, Oracle Linux, and Rocky Linux.
- Experience administering Kubernetes clusters, including deploying, scaling, and maintaining containerized workloads.
- Hands-on experience creating, managing, and troubleshooting Docker containers and container images throughout the software development lifecycle.
- Experience with Kubernetes cluster management and AI/ML workflow orchestration, including Argo, Airflow, and Kubeflow.
- Experience consuming and troubleshooting RESTful APIs for platform integration and automation.
- Excellent problem-solving skills and the ability to collaborate within a team.
- U.S. citizenship.
Preferred Qualifications
- Experience managing NVIDIA GPU data center platforms, including DGX, HGX, H200, H100, 200, B300, and L40S.
- Experience with NVIDIA enterprise tools such as Base Command Manager, Run:AI, and NVIDIA AI Enterprise.
- Knowledge of enterprise server components, including storage and network controllers, HBAs, and SSDs.
- Familiarity with GPU virtualization and cloud computing.
- Experience developing and deploying infrastructure in AWS.
- Knowledge of distributed resource scheduling systems such as Slurm, LSF, and Open MPI.
Required Education and Experience Equivalency
- Bachelor's or higher degree in Computer Science, Electrical Engineering, or a related field with 5+ years in Platform Engineering or Systems Engineering experience.
- Additional years of Platform Engineering or Systems Engineering experience in lieu of a degree.
Required Certifications
- Candidates must, at a minimum, meet DoD 8570.11 IAT Level II certification requirements, currently Security+ CE, CCNA-Security, GICSP, GSEC, or SSCP, along with an appropriate computing environment (CE) certification. An IAT Level III certification would also be acceptable, including CASP+, CCNP Security, CISA, CISSP, GCED, GCIH, or CCSP.
Required Security Clearance
- Active Top Secret/SCI
Platform Engineer ยท Base-2 Solutions, LLC