Machine Learning Engineer, Inference Infrastructure
ai breaking wireBrooklyn, NY
today
Occupations
Computer Systems Engineers/ArchitectsSoftware DevelopersComputer and Information Research ScientistsIndustries
Computer Systems Design ServicesComputing Infrastructure Providers, Data Processing, Web Hosting, and Related ServicesCustom Computer Programming ServicesAbout the role
About the Role We are seeking a Machine Learning Engineer to join our Inference Infrastructure team. You will build and optimize the high-throughput, low-latency distributed systems that power models like GPT-4 and Sora for millions of users worldwide.
Responsibilities Architect, build, and scale low-latency distributed inference serving systems for massive generative models.
Optimize GPU utilization, memory management, and kernel execution for state-of-the-art transformer architectures.
Collaborate with research teams to ensure smooth transition of new model architectures into production.
Monitor system performance, troubleshoot bottlenecks, and implement robust reliability measures.
RequirementsBS, MS, or Ph. D. in Computer Science or related technical field.
3+ years of industry experience building large-scale distributed systems or ML infrastructure.
Deep proficiency in C++ and Python.
Extensive experience with CUDA, Triton, or deep learning hardware accelerators.
Familiarity with distributed training and inference frameworks (vLLM, TensorRT-LLM, Megatron).
Benefits Top-tier compensation including equity.
Full medical, dental, and vision coverage with zero employee contribution options.
Unlimited paid time off and flexible hybrid work policy.
Catered daily lunches and wellness stipends.
Matching similar jobs
JOB OVERVIEW
Experience level
Lead
Location
Brooklyn, NY
Occupation
Computer Systems Engineers/Architects
Industry
Computer Systems Design Services
Posted
today
Tired of running searches?
Rank the roles you'd take once, and matches like these arrive on their own.
CREATE PROFILE