Search...

Production Engineer

Movement Labs logo
Movement Labs

Movement Labs builds Move-based blockchain infrastructure that bridges the security of the Move programming language with Ethereum's ecosystem and liquidity.

Distributed
About Movement Labs

Movement creates a network for developers to build applications using the Move programming language. By connecting with Ethereum, projects gain both security and performance benefits. Movement's platform supports a variety of applications including DeFi, gaming, and NFTs, with tools that make building safer and faster for developers of all experience levels.

View jobs by Movement Labs

Skills

About the Role

You will build production software, infrastructure platforms, tooling, and automation. You will own the reliability of the systems you ship, improve monitoring and observability, and work with protocol, runtime, storage, and external engineering partners to make systems operable, debuggable, upgradeable, and resilient. You will define reliability practices such as SLOs, capacity planning, failure testing, and safe rollouts, lead incident response and root-cause analysis, and deliver long-term fixes.

Requirements

  • Strong software engineering background
  • Experience owning production distributed systems or critical infrastructure
  • Ability to write code in Rust, Python, or similar languages to solve infrastructure problems
  • Understanding of Linux, networking, and cloud infrastructure
  • Understanding of containers, Kubernetes, and modern CI/CD systems
  • Understanding of observability, performance analysis, and capacity planning
  • Ability to design for failure
  • Focus on automation and eliminating manual, error-prone work

Responsibilities

  • Design, build, and own production infrastructure, tooling, and automation
  • Use AI to improve the technical simplicity and economic efficiency of network node infrastructure
  • Own monitoring, alerting, and observability across the stack
  • Partner with protocol, runtime, and storage engineers to improve operability, debuggability, and safety
  • Drive reliability-by-design through SLOs, capacity planning, failure testing, and safe rollouts
  • Lead incident response, root-cause analysis, and long-term fixes
  • Treat production as a first-class product