Senior AI Compute Infrastructure Engineer
Payward, Inc. is the parent company and operating platform behind Kraken and a portfolio spanning trading, custody, payments, lending, tokenized assets, onchain finance, and benchmarks.
Maintainer signals as of 9/25/2026
Funding history
Investors
Projects
About Payward, Inc.
Payward operates unified financial infrastructure across crypto and traditional markets, including shared liquidity, risk and margin systems, collateral and settlement, compliance, and licensing. Its first-party portfolio includes Kraken, Kraken Pro, NinjaTrader, Breakout, xStocks, Payward Services, CF Benchmarks, Reap, Krak, Kraken Prime, Kraken OTC, Kraken Custody, and Bitnomial.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Operate GPU and accelerator clusters for AI training, inference, evaluation, and experimentation. Build scheduling and utilization systems, optimize inference performance and cost, develop observability, lead reliability practices, integrate new hardware and runtimes, and support long-term AI infrastructure architecture.
Requirements
- 5+ years of infrastructure engineering experience
- GPU compute or ML infrastructure experience
- Distributed systems
- High-performance computing
- Production platform engineering
- GPU cluster operations
- Accelerator-backed infrastructure
- Scheduling and orchestration
- Utilization monitoring
- Cost optimization
- Linux
- Networking
- Storage
- Containers
- Kubernetes
- Distributed runtimes
- Production debugging
- vLLM
- Triton Inference Server
- TensorRT
- Python
- Performance optimization
- Observability
- Incident response
Responsibilities
- Operate GPU and accelerator clusters
- Configure drivers, runtimes, kernels, and device plugins
- Build scheduling, orchestration, placement, and quota systems
- Optimize inference pipelines for latency, throughput, reliability, and cost
- Partner with ML engineers and researchers to remove infrastructure bottlenecks
- Build GPU observability and reporting
- Drive reliability, incident response, alerting, and runbooks
- Evaluate and integrate hardware, accelerators, runtimes, and serving frameworks
- Build tooling for visible and accountable GPU usage
- Contribute to AI infrastructure architecture decisions
Hiring Process
Candidates may be asked to complete job-related skills or work-style assessments. Results are considered alongside experience and interviews and are not the sole basis for an employment decision.
