Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
Uber logo

Sr Software Engineer - Engineer

Uber
  • 🇺🇸 United States
  • On-site
  • Senior
  • 1 day ago
  • $202,000 – $224,000
  • AI
  • Machine Learning
  • Java
  • C++
  • Python
  • Flink
  • Incident Response
  • System Design
  • Elasticsearch
  • OpenSearch
  • Apache Solr
  • vLLM
  • Triton
  • Large Language Models
  • Design Systems
  • Equity
  • Pension
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

Senior Software Engineer 

About the Role

We are seeking talented Senior Software Engineers to join our Search Engineering team and help build the next generation of AI-powered search experiences.

In this role, you will design and build large-scale backend and model-serving infrastructure that powers search retrieval, ranking, personalization, and emerging LLM-based search experiences. You will work on systems that operate at high request volumes with strict latency and reliability requirements, spanning traditional search infrastructure, machine learning model serving, and modern LLM inference.

You will collaborate closely with backend and ML engineers, data scientists, product managers, and platform teams to evolve the search stack toward more intelligent, real-time, and AI-native architectures. This includes integrating LLM-based ranking and retrieval, optimizing GPU inference and serving efficiency, incorporating real-time marketplace signals, and building scalable infrastructure that enables rapid experimentation while maintaining production-grade performance and reliability.

Basic Qualifications

  • 5+ years of professional software engineering experience building large-scale backend or distributed systems.
  • Strong programming skills in Go, Java, C++, Python, or a similar language.
  • Strong understanding of distributed systems, service-oriented architectures, concurrency, networking, caching, and data consistency.
  • Experience building high-throughput, low-latency online serving systems.
  • Experience with search, recommendation, ranking, machine learning serving, or other large-scale data-intensive systems.
  • Experience diagnosing and optimizing system performance across latency, throughput, reliability, and infrastructure efficiency.
  • Familiarity with distributed data processing and streaming technologies such as Kafka, Flink, Spark, or similar frameworks.
  • Experience operating production systems, including observability, monitoring, capacity planning, incident response, and reliability engineering.
  • Strong system design, problem-solving, and analytical skills with the ability to work across multiple layers of a complex production stack.

Preferred Qualifications

  • Experience building search and recommendation systems, including retrieval, ranking, query understanding, indexing, and personalization.
  • Hands-on experience with search technologies such as Elasticsearch, OpenSearch, Solr, Vespa, Lucene, or large-scale proprietary search systems.
  • Experience with LLM or ML model-serving infrastructure, including frameworks such as vLLM, Triton, or similar inference platforms.
  • Experience optimizing GPU-based inference workloads, including batching or micro-batching, request scheduling, model parallelism, memory management, and GPU utilization.
  • Familiarity with LLM serving concepts such as prefill/decode, KV caching, prefix caching, streaming generation, speculative techniques, and distributed inference.
  • Experience with embeddings, semantic retrieval, approximate nearest-neighbor search, semantic IDs, or generative retrieval.
  • Familiarity with constrained decoding or integrating real-time business and marketplace constraints into AI-powered serving systems.
  • Experience integrating near-real-time features and signals into latency-sensitive ranking or inference systems.
  • Experience designing ML/LLM systems with strong reliability, graceful degradation, experimentation, and launch-safety mechanisms.

What the Candidate Will Do

  1. Design and build highly scalable search and AI serving infrastructure with a focus on latency, throughput, reliability, and infrastructure efficiency.
  2. Develop the backend architecture for next-generation AI-powered search, including LLM-based retrieval, ranking, personalization, and generative search experiences.
  3. Integrate large language models and machine learning models into production search serving paths while meeting stringent latency and reliability requirements.
  4. Optimize GPU inference and model-serving performance through techniques such as batching, micro-batching, request routing, caching, streaming, and efficient resource utilization.
  5. Build infrastructure for advanced LLM-serving patterns such as context prefill, KV/prefix-cache reuse, progressive or streaming generation, and efficient multi-turn or paginated search experiences.
  6. Develop scalable retrieval and ranking systems spanning lexical retrieval, semantic retrieval, embeddings, structured signals, and ML/LLM-based ranking.
  7. Work on semantic-ID and constrained-decoding infrastructure that allows generative models to interact safely and efficiently with large-scale search catalogs and real-time marketplace signals.
  8. Build and optimize real-time feature retrieval, hydration, caching, and data-processing systems that provide fresh signals to ranking and LLM models.
  9. Partner closely with ML engineers and data scientists to productionize new ranking and relevance models, improve experimentation velocity, and shorten the path from model development to production.
  10. Continuously improve end-to-end search performance by identifying bottlenecks across retrieval, feature serving, model inference, orchestration, networking, and presentation hydration.
  11. Design systems for graceful degradation, observability, capacity management, experimentation, and safe production rollouts of new AI and search capabilities.
  12. Analyze production and experiment metrics to understand latency, relevance, reliability, and business trade-offs and use those insights to guide system architecture.
  13. Contribute to the long-term technical architecture of the Search platform as it evolves from traditional multi-stage retrieval and ranking toward increasingly unified, AI-native search systems.
  14. Write high-quality, maintainable production code and provide technical leadership through design reviews, code reviews, mentoring, and cross-team collaboration.
  15. Troubleshoot complex production issues across distributed search and AI-serving systems and drive improvements that prevent recurrence.

For San Francisco, CA-based roles: The base salary range for this role is USD $202,000 per year - USD $224,000 per year.


For Sunnyvale, CA-based roles: The base salary range for this role is USD $202,000 per year - USD $224,000 per year.


For all US locations, you will be eligible to participate in Uber's bonus program, and may be offered an equity award & other types of comp. All full-time employees are eligible to participate in a 401(k) plan. You will also be eligible for various benefits.

Sr Software Engineer - Engineer · Uber

Auto apply with Likeremote