Back to all jobs

Tech Lead, Gemini Inference Performance, DeepMind

Google

LondonOn-siteFull Time10+ years
Posted 4 hours ago

Minimum qualifications:

  • Bachelor’s degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 5 years of experience in a people management, supervision/team leadership role.
  • 5 years of experience in a technical leadership role and overseeing projects.

Preferred qualifications:

  • Master’s degree or PhD in Engineering, Computer Science, or a related technical field.
  • 5 years of experience working in a complex, matrixed organization.

About the job:

Artificial intelligence will be one of humanity’s most transformative inventions. At Google DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority.

We are pushing the boundaries across multiple domains. Our global teams offer diverse learning opportunities and varied career pathways for those driven to achieve exceptional results through collective effort.

Responsibilities:

  • Set the technical roadmap for a team of performance engineers, support and develop direct reports and drive prioritization across competing optimization opportunities — staying in the highest-leverage problems.
  • Leverage roofline analysis, hardware-level profiling, and systems analysis to identify and eliminate performance bottlenecks across ML frameworks, compilers, custom kernels, and serving infrastructure on hardware accelerators.
  • Apply a first-principles understanding of Transformer and Mixture-of-Experts model components to identify and prioritize opportunities to optimize their execution efficiency and memory footprint.
  • Collaborate with research teams early in the development lifecycle to evaluate inference implications, modeling how architectural choices impact latency, memory footprint, and serving costs.
  • Guide and contribute to the development of custom kernels and serving optimizations. Conduct deep performance profiling using hardware tracing tools to analyze accelerator utilization, memory bandwidth saturation, and interconnect latency across large-scale topologies.