Head of Inference Performance Visibility
Etched is an AI-hardware company building rack-scale frontier inference clusters.
Funding history
About Etched
Etched co-designs chips, racks, software, and manufacturing systems for efficient inference of frontier AI models, targeting throughput, latency, cost, and power efficiency across prefill and decode workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will define how performance is measured and understood across accelerators, hosts, racks, and clusters. You will build telemetry architectures, counter taxonomies, event-correlation frameworks, time-alignment mechanisms, and analysis tools. You will connect low-level hardware signals with distributed workload behavior, identify bottlenecks, and guide cross-layer instrumentation and performance decisions.
Requirements
- Experience building complex systems spanning hardware and software
- Experience building profiling, tracing, or observability systems
- Experience translating hardware signals into production telemetry and analysis infrastructure
- Experience correlating time-series events across distributed systems
- Systems programming expertise in C++ or Rust
- Experience designing distributed correlation, timestamp-alignment, or performance-modeling frameworks
- Experience with distributed tracing or observability platforms at scale
- Experience with high-performance computing systems and large AI training clusters
- Experience with hardware-counter design and instrumentation strategy
- Experience with large-scale ML performance modeling
- Experience leading cross-functional hardware and software architectural initiatives
Responsibilities
- Define telemetry collection and structuring across CPUs, drivers, interconnects, and accelerators
- Design scalable models for correlating performance events across devices and hosts
- Develop mechanisms that correlate hardware counters, runtime activity, communication phases, and workload semantics
- Implement distributed time-synchronization and trace-alignment strategies
- Define counter taxonomies and derived performance models
- Influence instrumentation strategies for future hardware generations
- Build tools for identifying multi-accelerator workload bottlenecks
- Build cluster-scale performance analysis for distributed inference
- Contribute to analysis engines and developer-facing performance tooling
- Shape performance intelligence for engineers debugging large-scale AI systems
Benefits
- Medical, dental, and vision coverage with generous premium coverage
- USD 500 monthly credit for waiving medical benefits
- USD 2.5k monthly housing subsidy for eligible nearby residents
- Relocation support to San Jose
- Wellness benefits covering fitness and mental health
- Daily office lunch and dinner
- Unlimited compute budget subject to ROI justification
