System Engineer GIOC
Vultr is an active independent cloud-infrastructure company providing global compute, GPU, bare-metal, storage, Kubernetes, and serverless AI inference services.
About Vultr
Vultr provides globally available cloud infrastructure for developers, enterprises, and AI innovators, including Cloud Compute, Cloud GPU, Bare Metal, Cloud Storage, managed Kubernetes, and serverless inference.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will monitor compute, storage, operating system, and virtualization health signals; triage and classify incidents; execute approved runbooks; document actions; and escalate issues outside L1 scope. You will maintain shift logs, communicate incident status, contribute to incident reviews, and identify opportunities for alert tuning and automation.
Requirements
- Relevant graduate or engineering degree
- 3–5 years of systems administration experience
- Linux and Windows knowledge
- Operating system processes, services, and log analysis
- Compute, storage, virtualization, VM, hypervisor, and basic networking knowledge
- Incident-management and severity-based triage knowledge
- Monitoring, observability, alerting-console, and ITSM exposure
- Written documentation and escalation communication skills
- Ability to work rotational 24x7 shifts, including nights, weekends, and holidays
- English verbal and written communication
Responsibilities
- Monitor systems-estate health signals across compute, storage, operating systems, and virtualization
- Acknowledge alerts and incidents within first-response SLA targets
- Review alerts, recent changes, and the CMDB before acting
- Classify incidents by severity, customer impact, and owning tower
- Route tickets to the correct tower or escalation path
- Execute documented runbooks and permitted L1 actions
- Document actions and supporting evidence in incident tickets
- Escalate unresolved or out-of-scope issues to Senior Engineers with complete handoffs
- Initiate and support major-incident bridges
- Maintain shift logs, ticket updates, incident timelines, and handover notes
- Communicate incident status to stakeholders and customers
- Identify recurring alerts for tuning and automation
- Improve runbooks and knowledge-base articles
- Participate in post-incident and trend reviews
Benefits
- Annual medical insurance stipend
- 9 company-paid holidays
- Generous leave policy
- One month paid sabbatical every 5 years
- Anniversary bonus each year
- Professional development reimbursement
- Internet reimbursement
- Fitness membership reimbursement
- Company-paid Wellable subscription
