Inference Software Engineer
Etched is an AI-hardware company building rack-scale frontier inference clusters.
Funding history
About Etched
Etched co-designs chips, racks, software, and manufacturing systems for efficient inference of frontier AI models, targeting throughput, latency, cost, and power efficiency across prefill and decode workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will port state-of-the-art models and build abstractions and testing capabilities for rapid iteration. You will build and scale runtime features for multi-node inference, intra-node execution, state management, and error handling. You will optimize collective communication layers and use profiling and debugging tools to resolve performance bottlenecks and correctness issues.
Requirements
- Proficiency in C++ or Rust
- Understanding of performance-sensitive or complex distributed software systems
- Familiarity with PyTorch or JAX
- Experience porting applications to non-standard accelerator hardware or platforms
Responsibilities
- Support model porting and build programming abstractions and testing capabilities
- Build, enhance, and scale runtime capabilities for multi-node inference, intra-node execution, state management, and error handling
- Optimize routing and communication layers using collectives
- Use performance profiling and debugging tools to identify bottlenecks and correctness issues
Benefits
- Medical, dental, and vision coverage with generous premium coverage
- USD 500 monthly credit for waiving medical benefits
- USD 2.5k monthly housing subsidy for eligible nearby residents
- Relocation support to San Jose
- Wellness benefits covering fitness and mental health
- Daily office lunch and dinner
- Unlimited compute budget subject to ROI justification
