Inference Software Engineer

Etched is an AI-hardware company building rack-scale frontier inference clusters.

San Jose, United States
About Etched

Etched co-designs chips, racks, software, and manufacturing systems for efficient inference of frontier AI models, targeting throughput, latency, cost, and power efficiency across prefill and decode workloads.

View jobs by Etched

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will port state-of-the-art models and build abstractions and testing capabilities for rapid iteration. You will build and scale runtime features for multi-node inference, intra-node execution, state management, and error handling. You will optimize collective communication layers and use profiling and debugging tools to resolve performance bottlenecks and correctness issues.

Requirements

  • Proficiency in C++ or Rust
  • Understanding of performance-sensitive or complex distributed software systems
  • Familiarity with PyTorch or JAX
  • Experience porting applications to non-standard accelerator hardware or platforms

Responsibilities

  • Support model porting and build programming abstractions and testing capabilities
  • Build, enhance, and scale runtime capabilities for multi-node inference, intra-node execution, state management, and error handling
  • Optimize routing and communication layers using collectives
  • Use performance profiling and debugging tools to identify bottlenecks and correctness issues

Benefits

  • Medical, dental, and vision coverage with generous premium coverage
  • USD 500 monthly credit for waiving medical benefits
  • USD 2.5k monthly housing subsidy for eligible nearby residents
  • Relocation support to San Jose
  • Wellness benefits covering fitness and mental health
  • Daily office lunch and dinner
  • Unlimited compute budget subject to ROI justification
Inference Software Engineer at Etched | JobStash