ES
Senior AI Agent Engineer
EPAM Systems
π¦π² Armenia | β€οΈπ€β€οΈ Belarus | π°π¬ Kyrgyzstan | π°πΏ Kazakhstan | πΉπ· Turkey | πΊπΈ United States | πΊπΏ Uzbekistan
Remote
Senior
1 day ago
- AI
- Bedrock
- Azure AI
- Foundry
- Vertex AI
- MCP
- CI/CD
- Penetration Testing
1 day ago
Design and deliver AI agent systems that remain reliable beyond a successful demonstration. In this role, you will own technical decisions and delivery for a defined system or substantial component, from requirements and design through production operation. You will combine strong software engineering with practical agent development, make trade-offs explicit, and help other engineers deliver high-quality work.
We welcome engineers from different programming backgrounds. We value depth in your existing language and the ability to learn the languages and tools needed for the role.
Responsibilities
- Design, build, deploy, and operate agent systems across execution, memory, identity, tool integration, and observability within your area of responsibility
- Work directly with customers and stakeholders to turn needs into functional requirements and measurable acceptance criteria. Personally build working prototypes and production-grade slices of the solution, from design through implementation, testing, and deployment
- Develop agent workflows and production deployments using cloud managed services for AI hosting (e.g., Bedrock, Azure AI Foundry, Vertex AI), selecting services and development tools against the system's requirements
- Implement and operate an LLM gateway & routing layer covering model tiering, quota/cost control, and observability (e.g., LiteLLM, APIM AI Gateway)
- Design memory and context-management approaches for multi-step workflows, including state persistence, retrieval quality, retention, and isolation between users or engagements
- Build and maintain MCP servers and tool integrations for authorized security workflows, including reconnaissance, scanners, controlled exploit tooling, and internal services. Define contracts, permissions, approval boundaries, and recovery behavior
- Establish automated tests and agent evaluations for your components. Distinguish model-quality issues from software defects, and cover tool failures, interrupted execution, and unintended repeated actions
- Diagnose production issues using logs, metrics, and traces. Improve reliability, latency, and cost, and verify that mitigation and recovery work as intended
- Contribute to CI/CD, release checks, rollback plans, code review, and technical documentation. Mentor less experienced engineers and communicate decisions clearly to technical and non-technical stakeholders
Requirements
- Substantial production software engineering experience, typically five or more years, including at least one year delivering LLM-based agents to production. You can explain your contribution, the trade-offs you made, and the operational results
- Strong programming skills and production experience in at least one general-purpose language, together with hands-on agent development using SDKs, frameworks, or direct model APIs. You can independently learn an unfamiliar language or runtime and validate your implementation through tests and diagnostics
- Depth in at least one technical area and the ability to evaluate language, runtime, and framework trade-offs against the system's requirements, including concurrency, performance, maintainability, and integration needs
- Experience designing systems or substantial components, evaluating alternatives, and independently making technical decisions within an agreed scope
- Practical experience with MCP, tool use, prompt and context engineering, memory, and agent evaluation, including testing behavior across multiple steps and failure conditions
- Practical experience with cloud managed services for AI hosting (e.g., Bedrock, Azure AI Foundry, Vertex AI), including runtime, state, identity, integration, and observability concerns, and the ability to adopt an unfamiliar platform through working implementations
- Production engineering skills in API design, asynchronous processing, persistence, authentication and authorization, observability, retries, and security boundaries. You can investigate issues that span multiple components
- Experience with automated testing, CI/CD, code review, and AI-assisted development, together with a disciplined approach to validating generated output. You can clarify requirements, communicate designs, and support the growth of other engineers
- Proficiency in English at a B2+ level
Nice to have
- Experience in application security, web penetration testing, or red-teaming within an authorized scope
- Experience with sandboxed code execution, browser agents, or autonomous tool orchestration
- Experience improving evaluation coverage, deployment safety, or operating costs for agent systems
- Experience conducting technical interviews or contributing to engineering standards and knowledge sharing
Senior AI Agent Engineer Β· EPAM Systems