PT
Senior Software Engineer, AI Platform Engineering
PRI Technology
πΊπΈ United States
On-site
Senior
2 months ago
- AI
- Incident Response
- Kubernetes
- OpenAI
- Gemini
- AWS Bedrock
- Terraform
- Python
- Java
- AWS
- EC2
- IAM
- IaC
- OpenTelemetry
- Prometheus
- Grafana
- Datadog
- EKS
- VPC
- Disaster Recovery
- Claude Code
- Cursor
- Machine Learning
- Bedrock
- PyTorch
- TensorFlow
- scikit-learn
2 months ago
Contract
RESPONSIBILITIES
β Improve platform reliability and resilience by designing for failure, defining and meeting SLOs, leading incident response, and reducing operational toil.
β Design, build, and operate Kubernetes-based PaaS frameworks, AI/LLM gateways, APIs, and self-service tools for AI applications.
β Develop model-agnostic gateway capabilities for providers such as OpenAI, Anthropic, Gemini, and AWS Bedrock, including routing, fallback, retries, rate limiting, and cost controls.
β Build observability systems covering metrics, logs, traces, dashboards, and alerting to detect and resolve issues before they affect clients.
β Develop networking solutions that connect applications across public-cloud and on-premises environments.
β Provision and manage cloud infrastructure using Terraform and modern software engineering practices.
β Keep platforms secure and current through dependency patching, runtime upgrades, migrations, and provider-integration updates.
β Create frameworks, templates, and workflows that improve developer productivity and reduce operational overhead.
β Evaluate emerging AI technologies and adapt the platform to support new development patterns and use cases.
QUALIFICATIONS
β 6+ years of professional software engineering experience.
β Strong Python skills and experience developing production-grade backend services and APIs; Java experience is a plus.
β Experience designing and operating distributed systems in public-cloud environments, with a strong understanding of failure modes and resilient design patterns.
β Hands-on AWS experience, including services such as EC2, S3, IAM, and container-based workloads.
β Experience with Infrastructure as Code, preferably Terraform.
β Experience with production operations, including metrics, logging, tracing, alerting, SLOs, and incident response.
β Strong knowledge of software architecture, databases, networking, cloud infrastructure, and modern application development.
β A degree in computer science, engineering, or a related field, or equivalent practical experience.
Preferred Qualifications
β Experience building or operating API gateways, LLM gateways, or similar proxy layers with routing, fallback, rate limiting, caching, and cost tracking.
β Experience with OpenTelemetry, Prometheus, Grafana, Datadog, or similar observability tools.
β Experience with Kubernetes and autoscaling technologies, preferably Amazon EKS and Karpenter.
β Knowledge of AWS networking and security, including VPC, Direct Connect, IAM, and cloud security controls.
β Experience with chaos engineering, load and failure testing, capacity planning, or disaster recovery.
β Experience developing AI-powered applications, agent-based systems, model-inference services, or AI serving platforms.
β Familiarity with AI development tools such as Claude Code, Cursor, or GitHub Copilot.
β Working knowledge of machine learning concepts and the ML development lifecycle. Experience with SageMaker, Bedrock, PyTorch, TensorFlow, or scikit-learn is a plus.
β The ability to learn quickly and independently lead large technical initiatives from concept through production.
Senior Software Engineer, AI Platform Engineering Β· PRI Technology