Skip to main content

Sr Software Engineer - Engineer

Location
Sunnyvale, California / San Francisco, California
Team
Engineer
Subteam
Software Engineering
Posted on
Oct 8, 2026

Senior Software Engineer 

About the Role

We are seeking talented Senior Software Engineers to join our Search Engineering team and help build the next generation of AI-powered search experiences.

In this role, you will design and build large-scale backend and model-serving infrastructure that powers search retrieval, ranking, personalization, and emerging LLM-based search experiences. You will work on systems that operate at high request volumes with strict latency and reliability requirements, spanning traditional search infrastructure, machine learning model serving, and modern LLM inference.

You will collaborate closely with backend and ML engineers, data scientists, product managers, and platform teams to evolve the search stack toward more intelligent, real-time, and AI-native architectures. This includes integrating LLM-based ranking and retrieval, optimizing GPU inference and serving efficiency, incorporating real-time marketplace signals, and building scalable infrastructure that enables rapid experimentation while maintaining production-grade performance and reliability.

Basic Qualifications

  • 5+ years of professional software engineering experience building large-scale backend or distributed systems.
  • Strong programming skills in Go, Java, C++, Python, or a similar language.
  • Strong understanding of distributed systems, service-oriented architectures, concurrency, networking, caching, and data consistency.
  • Experience building high-throughput, low-latency online serving systems.
  • Experience with search, recommendation, ranking, machine learning serving, or other large-scale data-intensive systems.
  • Experience diagnosing and optimizing system performance across latency, throughput, reliability, and infrastructure efficiency.
  • Familiarity with distributed data processing and streaming technologies such as Kafka, Flink, Spark, or similar frameworks.
  • Experience operating production systems, including observability, monitoring, capacity planning, incident response, and reliability engineering.
  • Strong system design, problem-solving, and analytical skills with the ability to work across multiple layers of a complex production stack.

Preferred Qualifications

  • Experience building search and recommendation systems, including retrieval, ranking, query understanding, indexing, and personalization.
  • Hands-on experience with search technologies such as Elasticsearch, OpenSearch, Solr, Vespa, Lucene, or large-scale proprietary search systems.
  • Experience with LLM or ML model-serving infrastructure, including frameworks such as vLLM, Triton, or similar inference platforms.
  • Experience optimizing GPU-based inference workloads, including batching or micro-batching, request scheduling, model parallelism, memory management, and GPU utilization.
  • Familiarity with LLM serving concepts such as prefill/decode, KV caching, prefix caching, streaming generation, speculative techniques, and distributed inference.
  • Experience with embeddings, semantic retrieval, approximate nearest-neighbor search, semantic IDs, or generative retrieval.
  • Familiarity with constrained decoding or integrating real-time business and marketplace constraints into AI-powered serving systems.
  • Experience integrating near-real-time features and signals into latency-sensitive ranking or inference systems.
  • Experience designing ML/LLM systems with strong reliability, graceful degradation, experimentation, and launch-safety mechanisms.

What the Candidate Will Do

  1. Design and build highly scalable search and AI serving infrastructure with a focus on latency, throughput, reliability, and infrastructure efficiency.
  2. Develop the backend architecture for next-generation AI-powered search, including LLM-based retrieval, ranking, personalization, and generative search experiences.
  3. Integrate large language models and machine learning models into production search serving paths while meeting stringent latency and reliability requirements.
  4. Optimize GPU inference and model-serving performance through techniques such as batching, micro-batching, request routing, caching, streaming, and efficient resource utilization.
  5. Build infrastructure for advanced LLM-serving patterns such as context prefill, KV/prefix-cache reuse, progressive or streaming generation, and efficient multi-turn or paginated search experiences.
  6. Develop scalable retrieval and ranking systems spanning lexical retrieval, semantic retrieval, embeddings, structured signals, and ML/LLM-based ranking.
  7. Work on semantic-ID and constrained-decoding infrastructure that allows generative models to interact safely and efficiently with large-scale search catalogs and real-time marketplace signals.
  8. Build and optimize real-time feature retrieval, hydration, caching, and data-processing systems that provide fresh signals to ranking and LLM models.
  9. Partner closely with ML engineers and data scientists to productionize new ranking and relevance models, improve experimentation velocity, and shorten the path from model development to production.
  10. Continuously improve end-to-end search performance by identifying bottlenecks across retrieval, feature serving, model inference, orchestration, networking, and presentation hydration.
  11. Design systems for graceful degradation, observability, capacity management, experimentation, and safe production rollouts of new AI and search capabilities.
  12. Analyze production and experiment metrics to understand latency, relevance, reliability, and business trade-offs and use those insights to guide system architecture.
  13. Contribute to the long-term technical architecture of the Search platform as it evolves from traditional multi-stage retrieval and ranking toward increasingly unified, AI-native search systems.
  14. Write high-quality, maintainable production code and provide technical leadership through design reviews, code reviews, mentoring, and cross-team collaboration.
  15. Troubleshoot complex production issues across distributed search and AI-serving systems and drive improvements that prevent recurrence.

For San Francisco, CA-based roles: The base salary range for this role is USD $202,000 per year - USD $224,000 per year.

For Sunnyvale, CA-based roles: The base salary range for this role is USD $202,000 per year - USD $224,000 per year.

For all US locations, you will be eligible to participate in Uber's bonus program, and may be offered an equity award & other types of comp. All full-time employees are eligible to participate in a 401(k) plan. You will also be eligible for various benefits.

Ready to Ride?

This isn't the kind of place where you follow a playbook — it's where you help write one. If you're driven by impact, energized by challenge, and ready to shape how the world moves — we'd love to hear from you.

You may be eligible for bonuses, equity, and other compensation, as well as a range of benefits. Explore our benefits.

Offices remain key to collaboration and Uber's culture. Unless approved for full remote work, employees must spend at least 50% of their time in-office. Some roles, like those at greenlight hubs, require full-time in-office presence. Ask your Recruiter for details about this role's requirements.

Uber is proud to be an Equal Opportunity employer. All qualified applicants will receive consideration for employment without regard to sex, gender identity, sexual orientation, race, color, religion, national origin, disability, protected Veteran status, age, or any other characteristic protected by law. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. If you have a disability or special need that requires accommodation, please let us know by completing this form.

Related jobs

Sunnyvale, California + 1 location
Engineer
Seattle, Washington + 1 location
Engineer
San Francisco, California + 1 location
Engineer