HPC Infrastructure Operations Lead
Jump Trading is a global trading firm where traders, engineers, and researchers develop trading strategies, models, infrastructure, and systems across asset classes and time horizons.
About Jump Trading
Jump Trading is a global trading firm focused on research-driven trading and the engineering of scalable models, tools, infrastructure, and execution systems. Its operations combine trading, technology, AI/ML, and quantitative research, and it also runs research and talent programs including conference travel grants and a fellowship program.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You lead data center site leads and their teams across multiple HPC facilities while owning operational reliability, standards, preventative maintenance, monitoring, incident response, hardware break-fix, inventory, capacity planning, vendor relationships, budgets, and AI-driven operational improvements. You regularly travel to HPC data center sites and partner with engineering teams.
Requirements
- Minimum 7+ years of data center operations experience
- At least 3 years leading teams in 24/7 critical infrastructure environments
- In-depth knowledge of data center power systems and redundancy architectures
- In-depth knowledge of air, liquid, and hybrid cooling technologies
- Experience building preventative maintenance programs and operational procedures
- Experience with DCIM, BMS, environmental sensors, and monitoring strategies
- Deep knowledge of server and network switch hardware
- Hardware break-fix experience at the component level
- Strong networking knowledge including L2/L3, VLANs, BGP, OSPF, LACP, ECMP, and HPC topologies
- Strong Linux systems proficiency
- Experience managing inventory and spares programs
- Structured cabling standards expertise
- Programming or scripting experience, preferably Python
- Heavy professional use of AI tools
- Strong project management skills
- Knowledge of ASHRAE and TIA-942 standards
- Excellent written and verbal communication skills
- Ability to work on ladders and elevated platforms and lift up to 50 pounds
- Willingness to travel regularly
- Bachelor's degree preferred
Responsibilities
- Lead data center site leads and their teams across multiple HPC facilities
- Recruit, mentor, and develop team members
- Direct onsite contractors and validate completed work
- Develop and enforce HPC data center operational standards
- Own the preventative maintenance program
- Drive process improvement and automation
- Maintain expertise in power distribution and cooling architectures
- Own data center monitoring and alerting strategies
- Lead incident response and root cause analysis
- Track operational KPIs
- Evaluate hardware platforms and drive qualification testing
- Own hardware break-fix across HPC sites
- Manage inventory and critical spares
- Conduct capacity planning
- Plan new hardware installations
- Manage colocation providers and hardware vendors
- Develop and manage operational budgets
- Use AI tools across operational workflows
- Champion AI adoption within the team
- Partner with engineering teams to align operations with business needs
- Ensure compliance with safety, security, and regulatory requirements
Benefits
- Discretionary bonus eligibility
- Medical insurance
- Dental insurance
- Vision insurance
- HSA
- FSA
- Dependent Care options
- Employer-paid group term life insurance
- Employer-paid AD&D insurance
- Voluntary life insurance
- Voluntary AD&D insurance
- Paid vacation
- Paid holidays
- Retirement plan with employer match
- Paid parental leave
- Wellness programs
