Devops/On call
Ahrefs · REMOTE ok (Americas time zones)
AI Summary
Site Reliability Engineer role focused on maintaining Ahrefs' distributed crawler infrastructure across 2,000 bare-metal servers and 25 petabytes of storage. You'll handle 24/7 on-call responsibilities, work with custom OCaml code and technologies like ELK and Puppet, and balance automation with manual incident resolution for a large-scale distributed system.
Job Description
Our system is big part custom OCaml code and also employs third-party technologies - Debian, ELK, Puppet, and anything else that will solve the task at hand. In this role, be prepared to deal with 25 petabytes storage cluster, 2,000 bare-metal servers, experimental large-scale deployments and all kinds of software bugs and hardware deviations on a daily basis.
If you possess a healthy desire to automate everything while being able to quickly resolve urgent issues manually, then we want you! We strive to keep humans away from doing repetitive jobs that can be done by computers and focus instead on foreseeing problems and defining programmatic means to handle them. If there is any new technology that will make our life easier - no doubt, we'll give it a try. We rely heavily on opensource code (as the only viable way to build maintainable system) and contribute back [1]. Occasionally we track down CPU bugs [2].
Our motto is "first do it, then do it right, then do it better". Drop an email to connect@ahrefs.com
[1] https://github.com/ahrefs [2] https://tech.ahrefs.com/skylake-bug-a-detective-story-ab1ad2...
See how well your resume matches this job before you apply
Run a free ATS check