Software Engineering
Site Reliability / Gitops Engineer
Canonical
Description
Canonical publishes Ubuntu and has been remote-first since 2004, with 1,200+ colleagues across 75+ countries and very few office-based roles.
**The role**
This position sits in Canonical's Information Systems team, which supports and maintains every one of Canonical's IT production services. The team runs services used by **over 60 million Ubuntu users**, which is an unusually large blast radius for an internal platform group.
The listing frames it as a role for an "automation-first" technologist with a genuine interest in Linux. You would push operations automation further across both Canonical's private clouds and the public clouds, using open source infrastructure-as-code tooling, CI/CD practice, and Canonical's own operations automation products.
There is a second dimension worth noting. Beyond defining infrastructure as code, you would improve Canonical's products and the open source technologies underneath them by feeding critical operational insight back to developers — filing bugs, sometimes writing pull requests, and collaborating on design with other teams. In other words, running the systems at scale directly shapes what gets built.
You would join a global team of SREs supporting one another across timezones.
**What you would do**
- Apply infrastructure-as-code experience to steadily increase automation and improve IaC practice within IS
- Automate software operations for reusability and consistency across private and public clouds, accounting for distributed-systems complexity
- Develop new features and improve resilience and scalability across Canonical's cloud and container portfolio
- Maintain operational responsibility for all core services, networks and infrastructure
- Build troubleshooting, capacity planning and performance investigation skills, and run observability tooling including Prometheus, Grafana and Elasticsearch
- Design and maintain monitoring and alerting across systems and services
- Collaborate with development teams on service architecture, documentation, playbooks, policies and operational procedures
- Support globally distributed engineering, operations and support peers
- Share know-how through design sessions, mentorship and pairing
- Carry final responsibility for time-critical escalations
Canonical also states you would be given **uninterrupted development time** to focus on larger projects and automating manual work — a specific commitment that is rare in SRE postings.
**What they are looking for**
- Deep experience defining operations in code, using version control, peer review and CI/CD to roll out changes to both applications and infrastructure
- A strong modern engineering background: peer review, unit testing, SCM, CI/CD, Agile
- Python development experience on large projects
- Practical knowledge of Linux networking, routing and firewalls
- Affinity with Linux storage, from Ceph through to databases
- Hands-on enterprise Linux server administration
- Extensive knowledge of cloud computing concepts
- Bachelor's degree or above, preferably computer science or a related engineering field
- Clear English communication across email, chat, video and in person
- Willingness to troubleshoot from kernel to web, and to ask for help when appropriate
- Comfort in fast-changing environments and distributed teams
- Genuine enthusiasm for open source, particularly Ubuntu or Debian
**Benefits**
- Distributed work with twice-yearly in-person team sprints
- USD 2,000 personal learning and development budget per year
- Compensation reviewed every six months, plus annual bonuses
- Recognition rewards, annual holiday leave, maternity and paternity leave
- Employee Assistance Programme
- Priority Pass and travel upgrades for long-haul company events
No salary figures are published on this listing.
**Location**
Header, og:description and body text all agree: **Home based - Worldwide**, and the body adds that the role is "available remotely in any timezone." The application even asks which region and timezone you would work from, rather than treating a non-standard location as an exception.
**Before applying**
**AI use disqualifies your application.** You must confirm you used only your own words.
Expect the standard Canonical academic screening, plus **four written or select answers**:
- Your experience of site reliability engineering on large-scale applications or systems
- Your experience of infrastructure defined in code
- Your overall proficiency with Python
- What is the first thing you look at when a Linux system performs slowly
That final question is a straight technical test. Have a real answer ready.
Required Skills
Similar Jobs
4 more roles like this one.
See roles like this one
A free account opens every listing on EnRoute Jobs, saves the ones worth keeping, and scores each against your skills.
Create a free accountAlready have one? Log in
.png)