Software Engineer Infrastructure and Reliability
CrewAI is an enterprise agent build-and-runtime platform, with an open-source framework for building multi-agent systems and commercial controls for deployment, governance, and observability.
About CrewAI
CrewAI provides tools to design agents, orchestrate crews, automate flows, and operate production agentic workflows. Its current enterprise platform offers a centrally governed control plane, while its open-source offering provides the underlying multi-agent framework.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build and operate cloud and enterprise infrastructure across hyperscalers. You will improve CI/CD, observability, deployment safety, incident response, security, and self-hosted installation tooling while automating operational workflows and supporting production reliability.
Requirements
- Infrastructure or platform engineering experience in production SaaS environments
- AWS, Docker, CI/CD, GitHub Actions, and containerized-services experience
- ECS or Kubernetes experience
- Experience operating PostgreSQL, Redis, background job systems, queues, and web services
- Debugging skills across application, infrastructure, network, deployment, and dependency layers
- Knowledge of IAM, secrets, workload identity, vulnerability management, and production access
- Ability to write automation in Python, Ruby, Go, Bash, or similar languages
- Incident, rollback, migration, and production change-management experience
Responsibilities
- Own and improve platform infrastructure across AWS, containers, networking, secrets, databases, and Redis
- Build and maintain CI/CD pipelines for builds, tests, publishing, migrations, promotion, rollbacks, and deployments
- Improve reliability through health checks, alerting, incident response, capacity planning, recovery paths, runbooks, and on-call support
- Partner on Celery, FastAPI, Redis, Rails, queue, and PostgreSQL production workloads
- Manage logs, metrics, traces, dashboards, telemetry, and actionable alerts
- Harden IAM, workload identity, secrets management, vulnerability scanning, and least-privilege access
- Build automation and tooling for self-hosted installations
- Automate recurring operational workflows
