Founding Infrastructure Engineer
Metagov is a 501(c)3 nonprofit research-to-infrastructure laboratory focused on community sovereignty, network governance, and statecraft in the age of AI. It cultivates tools, practices, and communities that enable self-governance in the digital age, working with researchers, practitioners, and online communities (e.g., DAOs, online governance experiments) rather than traditional commercial clients.
Projects
About Metagov
Metagov is a nonprofit research-to-infrastructure laboratory, founded in 2019 out of a Harvard Law School seminar on "Governing Virtual Worlds" and spun out as an independent 501(c)3 nonprofit in January 2020. Its mission is to cultivate tools, practices, and communities that enable self-governance in the digital age, working toward a governance layer for the internet that is empowering, creative, interconnected, and accountable. The organization operates through three interdependent pillars: Community (a membership-based network of researchers and practitioners), Research (grant-funded projects led by Research Directors), and Operations (participatory organizational infrastructure). Metagov runs a weekly research seminar, community calls, and orientation sessions, and incubates projects such as DAOstar (a standards body for the DAO ecosystem), KOI Pond/GovBase (a research dataset project), Interoperable Deliberative Tools, and Public AI (focused on a new political economy for AI). It serves researchers, practitioners, online communities, DAOs, and protocol ecosystems, and is funded by a range of foundations and crypto-ecosystem organizations including the Ethereum Foundation, Filecoin Foundation, Optimism Collective, Mina Foundation, Arbitrum, Aragon, ENS, and NEAR, among others.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Technical owner of the operational backbone for a public AI inference utility, responsible for reliability, observability, routing transparency, production ML serving infrastructure, and coordination across cloud and HPC partners.
Requirements
- Significant experience operating production inference or ML serving infrastructure
- Strong distributed systems and SRE instincts
- Experience with observability, incident response, fallback design, and capacity planning
- Comfort working with cloud providers, sovereign HPC centers, and institutional IT
- Experience orchestrating multiple stacks and open-source projects
- Maintainer and integrator experience
- Ability to work autonomously in a small team
- Ability to travel occasionally for team workshops
Responsibilities
- Harden the platform for major launches
- Perform load testing and build fallback routing
- Set up monitoring and end-to-end observability
- Ship downtime warnings and fallback behavior
- Implement routing transparency and endpoint provenance
- Improve public endpoint performance
- Integrate MCP and programmatic infrastructure interfaces
- Operate production inference and ML serving infrastructure
- Coordinate with cloud providers and HPC centers
- Orchestrate and integrate open-source stacks
