HPC Production Engineer
Jump Trading is a global trading firm where traders, engineers, and researchers develop trading strategies, models, infrastructure, and systems across asset classes and time horizons.
About Jump Trading
Jump Trading is a global trading firm focused on research-driven trading and the engineering of scalable models, tools, infrastructure, and execution systems. Its operations combine trading, technology, AI/ML, and quantitative research, and it also runs research and talent programs including conference travel grants and a fellowship program.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You design, implement, maintain, and support high-performance compute and storage systems. You build software and operating system tooling, monitor performance and faults, collaborate on code and testing infrastructure, improve documentation, support production systems, optimize HPC infrastructure with researchers, manage vendors, and participate in maintenance operations and operational support rotations.
Requirements
- 5+ years of professional experience in high-performance computing
- 5+ years of Linux systems administration experience
- Experience with parallel filesystems such as Lustre or GPFS
- Experience with batch systems such as Slurm or Grid Engine
- Proficiency in Go, Python, C, or another programming or scripting language
- Experience designing, building, and maintaining distributed systems
- Experience profiling and debugging application stacks
- Experience with SaltStack, Ansible, Puppet, or similar configuration management tools
- Root cause analysis
- Reliable and predictable availability
Responsibilities
- Design, implement, maintain, and support high-performance compute and storage systems
- Implement and support performance monitoring and fault monitoring systems
- Monitor systems, storage, and network performance
- Build tooling to compile, package, install, and upgrade software and operating system components at scale
- Write code and testing infrastructure across multiple programming languages
- Develop and improve systems and user documentation
- Participate in coordinated maintenance operations
- Collaborate with researchers to optimize HPC infrastructure
- Develop and monitor production computing environment tools
- Provide operational support on a rotating basis
- Manage relationships with outside vendors
- Follow company cybersecurity and IT policies
Benefits
- Private medical insurance
- Vision insurance
- Dental insurance
- Travel medical insurance
- Group pension scheme
- Group life assurance
- Income protection schemes
- Paid parental leave
- Parking benefits
- Commuter benefits
