View All Jobs 158939

Software Engineer, LLM Inference

Develop scalable, high-performance LLM inference software for multiple platforms
Shanghai
Senior
5 hours agoBe an early applicant
NVIDIA

NVIDIA

A leading designer of graphics processing units (GPUs) for gaming and professional markets, as well as system on a chip units (SoCs) for the mobile computing and automotive market.

NVIDIA Llm Inference Framework Developer Engineer

NVIDIA has continuously reinvented itself over two decades. NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world.

This is our life's work — to amplify human imagination and intelligence. AI becomes more and more important in Auto Driving and AI City. NVIDIA is at the forefront of the Auto Driving and AI City revolution and providing powerful solutions for them. All these solutions are based on GPU-accelerated libraries, such as CUDA, TensorRT and V/LLM inference framework etc. Now, we are now looking for an LLM inference framework developer engineer based in Shanghai.

What You'll Be Doing

  • Craft and develop robust inferencing software that can be scaled to multiple platforms for functionality and performance
  • Performance analysis, optimization and tuning
  • Closely follow academic developments in the field of artificial intelligence and feature update
  • Collaborate across the company to guide the direction of machine learning inferencing, working with software, research and product teams

What We Need To See

  • Masters or higher degree in Computer Engineering, Computer Science, Applied Mathematics or related computing focused degree (or equivalent experience)
  • 3+ years of relevant software development experience.
  • Excellent C/C++ programming and software design skills, including debugging, performance analysis, and test design.
  • Strong curiosity about artificial intelligence, awareness of the latest developments in deep learning like LLMs, generative models
  • Experience working with deep learning frameworks like PyTorch
  • Proactive and able to work without supervision
  • Excellent written and oral communication skills in English
  • Strong customer communication skills, powerfully motivated to provide highly responsive support as needed
+ Show Original Job Post
























Software Engineer, LLM Inference
Shanghai
Engineering
About NVIDIA
A leading designer of graphics processing units (GPUs) for gaming and professional markets, as well as system on a chip units (SoCs) for the mobile computing and automotive market.