Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
Image Frame Investment (UK) Limited logo

Sr. AI infra engineer

Image Frame Investment (UK) Limited
  • 🇸🇬 Singapore
  • On-site
  • Senior
  • 22 hours ago
  • AI
  • Devops
  • Fabric
  • Windows
  • Integration Testing
  • TCP/IP
  • VLANs
  • BGP
  • OSPF
  • InfiniBand
  • Performance Testing
  • Disaster Recovery
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

About the Hiring Team

Tencent Overseas IT has the mission to empower Tencent’s rapid global growth with future ready, global IT platforms, applications and services. We are chartered to lead the Overseas IT strategy, architecture, roadmap and execution. Satisfying our internal/external customers and becoming a world class global IT team are our top aspirations.

What the Role Entails

About the Team

The AI Compute Centre sits within Tencent's Overseas IT department, acting as the bridge between internal AI infrastructure demand and the external resources that fulfil it. We play key roles in the full lifecycle of AI infrastructure clusters — from requirement gathering and capacity planning, through architectural design and development, to delivery, operations, and DevOps — across regions worldwide.


About the Role

You will lead the cloud/edge cloud infrastructure projects you are responsible for delivering. Your focus is to turn business and technical requirements into deliverable plans, hold suppliers accountable, coordinate data centre readiness, and ensure that new capacity is accepted and handed over to operations with clear ownership.


This is a cross-functional role. You will work with specialist engineers on compute, networking, storage, facilities, and security decisions; you are not expected to personally configure every system or run every benchmark. You should, however, be able to challenge assumptions, identify gaps between vendor commitments and delivered outcomes, and drive issues to resolution.


What You Will Do

● Translate demand into infrastructure requirements. Work with AI business and platform teams to define compute capacity, network bandwidth and latency, connectivity, data access, scalability, availability, and operational needs. Convert these into supplier requirements, evaluation criteria, delivery milestones, and acceptance criteria.

● Lead network design and delivery reviews. Review data centre fabric, inter-data-centre and WAN connectivity, routing, IP planning, network segmentation, resilience, and expansion plans with internal architects and suppliers. Ensure designs address both AI workload requirements and day-to-day operability.

● Manage network providers and delivery dependencies. Coordinate network equipment vendors, carriers, colocation partners, and system integrators on circuits, cross-connects, cabling, configuration handoffs, implementation windows, and fault escalation. Track responsibilities and resolve gaps between suppliers.

● Drive data centre readiness. Coordinate rack space, power, cooling, physical access, cabling, carrier connectivity, and installation prerequisites. Identify site or network dependencies that could delay cluster deployment.

● Own integration and acceptance governance. Coordinate compute, network, storage, and platform teams through installation, integration, testing, remediation, and sign-off. Require evidence for connectivity, bandwidth, latency, failover, and agreed cluster-level outcomes—not just confirmation that equipment has been installed.

● Manage delivery and service accountability. Maintain integrated plans, risks, decision logs, and supplier action trackers. Define handover documentation, monitoring expectations, incident escalation, change procedures, support boundaries, and service review cadence.

● Communicate technical risks clearly. Explain design trade-offs, unresolved defects, capacity constraints, and schedule impacts to both engineering teams and business stakeholders.

Who We Look For

Must-haves

● Substantial experience in infrastructure architecture, network engineering, technical delivery, or service management, typically10+ years in relevant roles.

● Strong data centre or enterprise networking experience, with the ability to review network topology, routing, connectivity, redundancy, and capacity plans—not only manage project schedules.

● Working knowledge of TCP/IP, L2/L3 networking, VLANs, routing protocols such as BGP and OSPF, and common data centre network design principles.

● Experience delivering or operating networks across data centres, cloud environments, or multiple sites, including coordination with carriers, colocation providers, and network equipment vendors.

● Ability to review network test plans and evidence, identify gaps in connectivity or failover, and work with engineers and suppliers to resolve issues.

● Experience with a data centre migration, site deployment, colocation delivery, or comparable infrastructure transition involving multiple technical teams and external providers.

● Demonstrated experience managing suppliers through design review, delivery milestones, acceptance, escalation, and production handover.

● Ability to assess infrastructure proposals across compute, network, storage, availability, and operational support, while engaging specialists for detailed validation.

● Strong documentation and stakeholder communication skills.


Nice to Have

● Experience evaluating or sourcing AI server capacity, AI infrastructure services, or large-scale compute platforms.

● Deeper experience with Spine-Leaf/Clos fabrics, EVPN-VXLAN, ECMP, MLAG, or inter-data-centre connectivity.

● Familiarity with InfiniBand, RoCE, RDMA, AI server cluster network design, high-performance storage, or cluster-level performance testing such as NCCL.

● Experience with high-density data centre deployments, including power and cooling constraints.

● Exposure to supplier SLAs, commercial evaluation, cost optimisation, disaster recovery, or service governance.

● Relevant networking, cloud, AI infrastructure, project management, or IT service management certifications.

Equal Employment Opportunity at Tencent

As an equal opportunity employer, we firmly believe that diverse voices fuel our innovation and allow us to better serve our users and the community. We foster an environment where every employee of Tencent feels supported and inspired to achieve individual and common goals.

Sr. AI infra engineer · Image Frame Investment (UK) Limited

Auto apply with Likeremote