ML Infrastructure Engineer

ML Infrastructure Engineer

Job Type:

Direct-Hire

Location:

San Francisco - CA

Industry:

Technology & Software

Category:

Python

Compensation Range:

$ - $ Per Year

Additional Compensation Info:

Base salary plus equity

Job ID

26308

Rich Text Widget

Job Title: Machine Learning Infrastructure Engineer

Location: San Francisco, CA Metro Area (100% On-Site)

About the Opportunity

An ultra-high-growth artificial intelligence platform company is seeking a Machine Learning Infrastructure Engineer to help architect the compute, training, and execution frameworks powering next-generation model performance. Backed by top-tier venture capital firms and serving elite technology enterprises, this team is scaling rapidly to solve complex unstructured data challenges. In this role, you will hold direct ownership over scaling distributed systems across massive hardware clusters while collaborating directly with core platform teams.

Responsibilities

  • Architect and sustain high-throughput execution and training pipelines optimized for rapid experimentation, flexibility, and ultra-low latency serving.

  • Establish comprehensive benchmarking suites across system stacks to identify performance bottlenecks and optimize resource efficiency.

  • Evaluate cutting-edge developments in model optimization and integrate state-of-the-art research into production systems.

  • Engineer resilient, highly observable multi-node environments that support seamless parallelized training tasks.

  • Maximize hardware efficiency and operational reliability while running large-scale workloads across extensive GPU clusters.

  • Build developer abstractions, internal tooling, and telemetry frameworks that streamline the transition from research prototype to production.

Requirements (Must-Have)

  • 3+ years of production experience in software development, backend systems, or machine learning platforms (exceptional early-career candidates with extraordinary technical foundations will be evaluated).

  • Advanced proficiency in Python alongside a solid background in low-level systems architectural design.

  • Hands-on experience working with container orchestration (e.g., Kubernetes) and modern distributed compute frameworks.

  • Demonstrated ability to build from first principles and take full ownership of projects from strategic planning to deployment.

  • Comfortable operating in a fast-paced, high-rigor setting with continuous iteration cycles.

Preferred Qualifications (Nice-to-Have)

  • Prior experience working within an early-stage or rapidly scaling tech startup.

  • Meaningful open-source contributions to established distributed training or inference engines.

  • Practical experience managing multi-node inference setups across hundreds or thousands of GPUs.

  • Strong passion for aligning deep technical excellence with direct business results.

Compensation & Benefits

  • Work Arrangement: Full-time, 5 days per week on-site in San Francisco, CA.

  • Health & Wellness: Fully covered medical, dental, and vision coverage, plus a $150 monthly stipend for wellness and fitness expenses.

  • Time Off & Flexibility: Flexible PTO policy and adaptable parental leave programs.

  • Perks: Daily catered lunches in the office and full commuter reimbursement.

  • Equal Opportunity Employer: All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or protected veteran status.

Apply Now
Apply Now

Share this job

SCHEMA MARKUP ( This text will only show on the editor. )
Back to Job Search