Description
Sr. Director, Site Reliability Engineering
Coupang operates one of the largest and most complex technology platforms in the world. We are seeking a Senior Director, Site Reliability Engineering (Head of SRE) to define and lead company-wide reliability, resilience, scalability, and operational excellence. This leader will transform reliability from a collection of team-specific practices into platform mechanisms that services inherit by tier, while advancing incident response toward an intelligent, AI-assisted, and increasingly autonomous operating model. We are looking for a visionary, industry-recognized technology leader who has previously conceived, built, and scaled a comparable SRE, production engineering, resilience, or autonomous-operations organization at a leading global technology company. The successful candidate must combine deep technical credibility with the organizational leadership required to align executives, influence architecture across the company, and build a world-class leadership bench.
Key Responsibilities
- Set a bold, multi-year vision for company-wide reliability, resilience, and autonomous operations, and translate that vision into an executable roadmap with measurable business outcomes.
- Define and own the SRE strategy, operating model, engineering standards, and reliability governance across Coupang.
- Build platform mechanisms that allow services to inherit reliability requirements based on service tier rather than recreate them independently.
- Lead initiatives that materially improve availability, resilience, scalability, performance, and operational readiness.
- Partner with engineering, product, infrastructure, security, finance, and business leaders to align reliability investments with customer and business priorities.
- Own executive reliability metrics, including availability, detection and recovery performance, change risk, incident recurrence, capacity readiness, and operational toil.
- Build and scale a world-class SRE organization capable of influencing engineering practices across the company.
Reliability Strategy, SLOs & Engineering Governance
- Establish and evolve service-tier definitions, SLOs, SLAs, error budgets, reliability scorecards, and objective certification mechanisms such as RBD/RBO.
- Create clear reliability requirements for Tier 0, Tier 1, and Tier 2 services, including redundancy, load testing, disaster recovery, observability, and incident response.
- Ensure reliability governance is embedded in architecture, development, release, and production operations rather than applied as a final review.
- Drive systematic reduction of recurring incidents, reliability risks, operational debt, and unsafe change patterns.
- Influence company-wide architecture for graceful degradation, fault isolation, load shedding, circuit breaking, and failure containment.
Incident Management & Autonomous Operations
- Transform incident management into a fast, disciplined, data-driven, and increasingly autonomous operating model.
- Enable AI-assisted detection, event correlation, triage, escalation, root-cause drafting, remediation recommendations, and selected guardrailed auto-remediation.
- Improve incident command, on-call quality, escalation mechanisms, communication, post-incident learning, and corrective-action completion.
- Reduce noisy alerts, manual on-call work, repeated diagnosis, and time spent coordinating across fragmented systems.
- Use incident and telemetry data to continuously improve platform standards, testing, capacity models, and engineering roadmaps.
Disaster Recovery, Resilience & Capacity
- Own the strategy and execution model for disaster recovery, regional resilience, availability-zone loss, capacity-constrained recovery, and critical business continuity.
- Build reusable DR and failover mechanisms that services inherit from the platform rather than implement as bespoke projects.
- Establish objective RPO/RTO targets, automated readiness gates, regular game days, fault injection, and evidence-based recovery certification.
- Drive proactive and intelligent capacity management using forecasting, reservations, workload prioritization, and automated response to demand and failure scenarios.
- Partner with compute, traffic, networking, storage, and application leaders to enable safe zone evacuation, regional failover, and surge readiness.
Observability, Testing & Reliability Intelligence
- Partner with Observability and TestOps leaders to integrate logs, metrics, traces, continuous profiling, testing, and incident intelligence into one reliability feedback loop.
- Ensure every critical service has actionable telemetry, meaningful SLOs, release-quality signals, and production-readiness evidence.
- Use production incidents and operational patterns to drive targeted integration, load, resilience, and regression testing.
- Establish executive reliability dashboards that provide trusted views of service health, risk, capacity, and operational effectiveness.
Talent Leadership & Organization
- Lead multiple layers of SRE leaders, including senior managers, directors, principal engineers, and senior individual contributors.
- Own organizational design, global hiring strategy, leadership development, succession planning, and the creation of a strong leadership bench.
- Attract exceptional SRE, distributed systems, resilience, incident-management, and capacity-engineering talent from best-in-class technology organizations.
- Build an empowered organization with clear accountability, strong technical judgment, high execution velocity, and a company-wide perspective.
- Act as a force multiplier by mentoring technical and organizational leaders and raising reliability capabilities across engineering.
Technical Leadership & Architecture
- Own reliability architecture decisions across large-scale distributed systems and cloud-native infrastructure.
- Define resilient patterns for redundancy, failover, traffic management, data recovery, workload prioritization, and dependency isolation.
- Guide architecture reviews and platform standards for safe scaling, fault tolerance, and operational simplicity.
- Balance availability, customer impact, engineering velocity, cost, and operational complexity in major technical decisions.
- Maintain sufficient technical depth to challenge assumptions, guide principal engineers, and make high-quality decisions during critical incidents.
Execution & Impact
- Deliver measurable improvements in availability, time to detect, time to mitigate, time to recover, incident recurrence, change-failure rate, and on-call burden.
- Create disciplined operating rhythms, milestones, ownership models, and quarterly targets for strategic reliability programs.
- Increase adoption of common reliability mechanisms and reduce team-specific implementations and manual operations.
- Demonstrate business impact through improved customer experience, reduced outage exposure, stronger peak readiness, and more efficient use of infrastructure capacity.
- Build credibility through predictable delivery, transparent risk management, and objective evidence of reliability improvement.
Essential Qualifications
- Leadership experience in a best-in-class SRE, production engineering, infrastructure reliability, or cloud operations organization at hyperscaler, major cloud provider, global marketplace, leading fintech, or similarly scaled technology company.
- Experience building an SRE practice comparable in maturity to leading industry organizations, rather than operating a traditional support or operations function renamed as SRE.
- Experience with Kubernetes, service mesh, cloud-native platforms, traffic engineering, and large-scale capacity management.
- Experience with chaos engineering, fault injection, regional resilience, and automated disaster recovery.
- Experience building AI-assisted operations, incident intelligence, predictive reliability, or self-healing systems.
- 15+ years of experience in software engineering, infrastructure engineering, distributed systems, or site reliability engineering.
- 8+ years leading large-scale, multi-layer engineering organizations, including senior managers, directors, and senior individual contributors.
- Demonstrated experience personally defining the vision and leading the architecture, build-out, launch, and scaled adoption of a company-wide SRE, reliability, resilience, or autonomous-operations program.
- Prior experience building reliability systems and operating practices for high-scale, high-availability, customer-critical distributed systems.
- Deep expertise in SLOs, error budgets, observability, incident management, disaster recovery, capacity planning, and resilience engineering.
- Proven ability to lead through major incidents while also creating durable mechanisms that prevent recurrence.
- Recognized as a visionary technology and organizational leader who can influence executive stakeholders, align multiple engineering organizations, and attract exceptional talent.
- Proven ability to convert long-term strategy into measurable execution and company-wide adoption.
Similar jobs
Sr. Director, Site Reliability Engineering Coupang operates one of the largest and most complex technology platforms in the world. We are seeking a Senior Director, Site Reliability Engineering (Head of SRE) to define an…
Job Description – Sr. Director, TestOps Coupang operates a large distributed technology ecosystem where changes must be validated quickly, safely, and repeatedly across infrastructure and application boundaries. We are s…
Job Description – Sr. Director, TestOps Coupang operates a large distributed technology ecosystem where changes must be validated quickly, safely, and repeatedly across infrastructure and application boundaries. We are s…
본 공고는 재직 임직원을 대상으로 하는 사내 공모 전용입니다. (임직원 추천은 Link로 진행) This posting is exclusively for internal employees. (Employee referrals are submitted via the Link) 지원 시에는 반드시 첨부된 영문 ‘사내 공모 지원서 양식’을 작성한 후, 쿠팡 이메일 계정으로 접수해 주시기 바랍니다.…
We are representing a leading force in the decentralized exchange (DEX), and seeking a high-caliber technical leader to architect the backbone of a global financial ecosystem. In this role, you will bridge the gap betwee…
Est. 120,000 USD
We are representing a leading force in the decentralized exchange (DEX), and seeking a high-caliber technical leader to architect the backbone of a global financial ecosystem. In this role, you will bridge the gap betwee…
Est. 95,000 USD
SummaryWe are seeking an experienced Site Reliability Engineering (SRE) / DevOps Manager to lead our teams and drive reliability roadmaps. As a key player in our leading crypto tax and portfolio tracking platform, you wi…
We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did we ever live without Coupang?” Born out of an obsession to make shopping, eating, and living easier than ever, w…
Est. 124,000 USD
Application Support Engineer (Site Reliability Engineer) Location: USAJob Type: Full-Time, no visa sponsorship available Coforge is seeking a Senior Application Support Engineer (SRE) to join our dynamic team of consulta…
Company Intro We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did I ever live without Coupang?” Born out of an obsession to make shopping, eating, and living easier…
Est. 115,000 USD
Delivery & Service Assurance Lead Customer Operations — SPEC-Ops (SRE & Platform Engineering) Role Summary The Delivery & Service Assurance Lead runs the day-to-day rhythm of the SPE Customer Operations team,…
Est. 165,000 USD
84.51° Overview: 84.51° is a retail data science, insights and media company. We help The Kroger Co., consumer packaged goods companies, agencies, publishers and affiliates create more personalized and valuable experienc…
We are representing a leading force in the decentralized exchange (DEX), and seeking a high-caliber technical leader to architect the backbone of a global financial ecosystem. In this role, you will bridge the gap betwee…
Est. 217,500 USD
About NscaleNscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-nativestartups and global enterprises, from bare metal up through the platform services teams actually build…
Job Title: Site Reliability Engineer (SRE)Key Skills: Kubernetes, AWS/Azure/GCP, Terraform, Python, Observability, CI/CDExperience: +6 YOE.Location: Costa Rica, Peru, Colombia, and Bolivia.Mode: Remote. We at Coforge are…
Summary: We are seeking an experienced Site Reliability Engineering (SRE) / DevOps Manager to lead our teams and drive reliability roadmaps. As a key player in our leading crypto tax and portfolio tracking platform, you…
Est. 180,000 USD
Summary: We are seeking an experienced Site Reliability Engineering (SRE) / DevOps Manager to lead our teams and drive reliability roadmaps. As a key player in our leading crypto tax and portfolio tracking platform, you…
Est. 70,000 USD
SummaryWe are seeking an experienced Site Reliability Engineering (SRE) / DevOps Manager to lead our teams and drive reliability roadmaps. As a key player in our leading crypto tax and portfolio tracking platform, you wi…
Company Introduction We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did we ever live without Coupang?” Born out of an obsession to make shopping, eating, and livin…
Est. 120,000 USD
SummaryWe are seeking an experienced Site Reliability Engineering (SRE) / DevOps Manager to lead our teams and drive reliability roadmaps. As a key player in our leading crypto tax and portfolio tracking platform, you wi…
Senior Staff Engineer – Gateway Services The Gateway Services team is responsible for building Coupang’s traffic infrastructure layer using Service Mesh. Key aspects of the infrastructure layer include service discovery,…
Please complete the attached Internal Transfer Request Form and submit. Please make sure to apply with your Coupang e-mail address. Company Introduction We exist to wow our customers. We know we’re doing the right thing…
Title: Staff Site Reliability Engineer, Product Area FocusLocation: Noida / Bangalore (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excel…
Who We Are 2K is headquartered in Novato, California and is a wholly owned label of Take-Two Interactive Software, Inc. (NASDAQ: TTWO). Founded in 2005, 2K Games is a global video game company, publishing titles develope…
Title: Senior Site Reliability Engineer - I, Product Area FocusLocation: Noida (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excellence o…
Est. 180,000 USD
Who We Are Cross River builds the infrastructure behind the world’s most innovative financial products. Our technology and capital solutions power payments, cards, lending, and digital asset capabilities that move money…
Developer Platform – Gateway Services – Job Description Staff Engineer – Gateway Services Level: L7-1 Location: Bangalore p
Est. 80,000 GBP
At NiCE, we don’t limit our challenges. We challenge our limits. Always. We’re ambitious. We’re game changers. And we play to win. We set the highest standards and execute beyond them. And if you’re like us, we can offer…
Est. 60,000 USD
SummaryWe are seeking an experienced Site Reliability Engineering (SRE) / DevOps Manager to lead our teams and drive reliability roadmaps. As a key player in our leading crypto tax and portfolio tracking platform, you wi…
Title: Staff Site Reliability Engineer, Product Area FocusLocation: Noida/ Bangalore (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excell…