Search...

Senior AI Compute Infrastructure Engineer

Payward, Inc. logo
Payward, Inc.

Payward, Inc. is a global financial infrastructure company and the parent organization behind Kraken. It provides trading, custody, payments, lending, staking, tokenized assets, derivatives, and market data infrastructure to consumers, professional traders, institutions, enterprises, fintechs, banks, exchanges, asset managers, and onchain platforms.

Recently fundedCompany intelligence
Distributed

Funding history

About Payward, Inc.

Payward, Inc. operates a unified financial infrastructure platform powering a portfolio of products including Kraken, NinjaTrader, Breakout, xStocks, CF Benchmarks, and Payward Services. Its shared architecture provides global liquidity, risk and margin management, collateral and settlement, compliance and licensing, and operational infrastructure across crypto, tokenized assets, and traditional markets. Through Payward Services, the company offers APIs and infrastructure for crypto trading, custody, on/off-ramps, tokenized equities, derivatives, staking and yield, payments, and benchmark data. Payward operates across more than 190 jurisdictions and serves consumers, professional traders, institutional investors, enterprises, fintechs, banks, exchanges, asset managers, and DeFi/onchain protocols.

View jobs by Payward, Inc.

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will operate GPU and accelerator clusters for AI training, inference, evaluation, and experimentation. You will build scheduling and utilization systems, optimize inference performance and cost, develop observability, lead reliability practices, integrate new hardware and runtimes, and support long-term AI infrastructure architecture.

Requirements

  • 5+ years of infrastructure engineering experience
  • GPU compute or ML infrastructure experience
  • Distributed systems
  • High-performance computing
  • Production platform engineering
  • GPU cluster operations
  • Accelerator-backed infrastructure
  • Scheduling and orchestration
  • Utilization monitoring
  • Cost optimization
  • Linux
  • Networking
  • Storage
  • Containers
  • Kubernetes
  • Distributed runtimes
  • Production debugging
  • vLLM
  • Triton Inference Server
  • TensorRT
  • Python
  • Performance optimization
  • Observability
  • Incident response

Responsibilities

  • Operate GPU and accelerator clusters
  • Configure drivers, runtimes, kernels, and device plugins
  • Build scheduling, orchestration, placement, and quota systems
  • Optimize inference pipelines for latency, throughput, reliability, and cost
  • Partner with ML engineers and researchers to remove infrastructure bottlenecks
  • Build GPU observability and reporting
  • Drive reliability, incident response, alerting, and runbooks
  • Evaluate and integrate hardware, accelerators, runtimes, and serving frameworks
  • Build tooling for visible and accountable GPU usage
  • Contribute to AI infrastructure architecture decisions
Senior AI Compute Infrastructure Engineer at Payward, Inc. | JobStash