Senior Site Reliability Engineer New Dublin, Ireland

Details of the offer

Reddit is a community of communities. It's built on shared interests, passion, and trust and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 82M+ daily active unique visitors, Reddit is one of the internet's largest sources of information.
As a Senior Site Reliability Engineer on Reddit's Infrastructure SRE team, you'll use your knowledge of distributed systems and architecture to improve the reliability and performance of Reddit's engineering platforms and services. This team will work closely with the Compute, Traffic, and Observability infrastructure teams. They will own a suite of tools for allowing engineers to understand their creations, based primarily on open-source solutions at scale.
In this role, you will also take ownership of risk management, ensuring the reliability and performance of our systems. You will collaborate with cross-functional teams to identify, assess, and mitigate risks, implementing best practices to enhance system resilience. Your expertise will drive proactive measures to maintain uptime and optimize service delivery, making a significant impact on our operational excellence.
Responsibilities:Advise: Work closely with engineering teams in designing and developing systems that are resilient and highly performant at a tremendous scale, and maintaining the foundational platform for running Reddit's infrastructure.Amplify: Identify and build capabilities into our foundational Infrastructure and Platform services, which are used by Reddit engineering teams to build, deploy, and operate Reddit.Deliver software to improve the availability, scalability, latency, and efficiency of observability components.Identify and engineer away risk across Reddit's systems.Automate: Take repetitive, manual, or risky tasks and automate them out of existence. Build tools and integrate systems to support Reddit's evolution.Automate critical aspects of the event-driven development process.Diagnose: Draw on your knowledge of distributed systems to identify and fix network, system, and service-level issues. Practice sustainable incident response, and drive structural improvement with blameless postmortem.Share on-call responsibilities.Optimize: Observe and improve performance, reduce cost, and improve the experience for millions of users.Contribute upstream changes to the open source projects we use.Qualifications:5+ years of experience in Software Engineering, Site Reliability Engineering, or a development-focused DevOps role.Proficiency in one or more programming languages. We're predominantly writing code in Go and Python.Experience with Kubernetes and Cloud systems.Familiarity with distributed systems development, bonus if familiar with any of the specific tools (Prometheus, Thanos, Grafana, Vector, Clickhouse, Otel, Loki).Experience with the development and operation of high-traffic backend systems.A demonstrated ability to debug, fix, and optimize code.Troubleshooting skills that span applications, networking (TCP/IP), and systems.Strong working knowledge of Linux and containers.Excellent communication and collaborative skills.Apply for this job
#J-18808-Ljbffr


Nominal Salary: To be agreed

Source: Jobleads

Requirements

Network Engineer

About Post Consult International (PCI) PCI is a wholly owned subsidiary of An Post, providing specialist IT Strategy, Management and Development services acr...


An Post - County Dublin

Published a month ago

Data Scientist - Pharmacy Analytics

Data Scientist, Pharmacy Analytics – Dublin, HybridOptum is a global organization that delivers care, aided by technology to help millions of people live hea...


Unitedhealth Group - County Dublin

Published a month ago

Azure Devops Engineer

About the Role The ideal candidate will possess extensive DevOps experience and will join our UK&I Technology team. You will be responsible for developing to...


Lexisnexis Risk Solutions Uk Ltd. Company - County Dublin

Published a month ago

Saas Technical Manager

The most trusted digital enabler team.blue is a leading digital enabler for companies and entrepreneurs. It serves over 3.3 million customers in Europe and h...


Team.Blue Global - County Dublin

Published a month ago

Built at: 2024-11-15T13:27:06.995Z