Description
About SHEIN
SHEIN is a global online fashion and lifestyle retailer, offering SHEIN branded apparel and products from a global network of vendors, all at affordable prices. Headquartered in Singapore, with more than 15,000 employees operating from offices around the world, SHEIN is committed to making the beauty of fashion accessible to all, promoting its industry-leading, on-demand production methodology, for a smarter, future-ready industry.
Position Summary
We are seeking a Staff Site Reliability Engineer (Official Title: Staff Site Reliability Engineer I) with deep experience operating and evolving large-scale, mission-critical systems where availability and reliability are non-negotiable.
At SHEIN, Site Reliability Engineers are hybrid software and systems engineers responsible for keeping production services always on while enabling the platform to scale rapidly and safely. In this role, you will own and support complex services and infrastructure, ensuring they consistently meet reliability and performance expectations. At the Staff level, you will also provide technical leadership, influencing platform architecture, reliability strategy, and operational standards across the organization.
The SRE team owns and maintains critical open-source and in-house technologies that underpin the platform and serves as a core contributor to major engineering initiatives. We are accountable for driving platform operability forward by reducing incident frequency, minimizing MTTR, and improving system resilience, efficiency, and resource utilization.
You will work closely with global, cross-functional teams to design, build, and evolve observability and operational tooling—including metrics, logs, traces, alerting, and automation—providing deep visibility into system behavior. Through hands-on engineering and operational excellence, you will proactively identify risks and failure modes, help prevent incidents before they occur, and lead fast, effective responses when they do. To succeed in this role, you will combine strong software engineering skills, solid to deep expertise in Linux, networking, and distributed systems, and a passion for solving problems of scale, complexity, and reliability. Your work will directly contribute to delivering a stable, scalable, and high-performing experience for customers worldwide.
Job Responsibilities
- Keep SHEIN’s mission-critical production systems running 24/7/365, participating in on-call rotations and acting decisively during incidents.
- Triage and resolve production incidents, leveraging AI-assisted log analysis and anomaly detection to accelerate root cause identification; drive continuous improvements that reduce MTTR and prevent recurrence.
- Monitor and manage capacity planning and resource utilization, partnering with cross-functional teams to ensure systems scale safely while remaining cost-effective.
- Own and operate core open-source infrastructure such as APISIX, Nginx, Kubernetes, Kafka, Elasticsearch, Redis, Consul, Etcd, Zookeeper and other large-scale distributed systems.
- Design, build, and maintain observability solutions (metrics, logs, traces, alerting), incorporating AI-powered anomaly detection and intelligent alert correlation to surface actionable signals from high-volume telemetry, improving system visibility and resiliency.
- Automate operational workflows and eliminate manual toil through scripting, tooling, and process improvements, including the use of AI-assisted development tools (e.g., Claude Code) to accelerate the building and iteration of internal operational platforms.
- Develop and maintain technical documentation, including runbooks, architecture diagrams, operational procedures, and on-call playbooks.
- Work closely with global engineering teams to improve infrastructure reliability and performance through better system design and operational discipline.
- Mentor Senior and mid-level SREs, raising the overall technical bar and operational maturity of the team.
- Lead efforts to modernize the platform in alignment with industry best practices and evolving technology standards.
Job Requirements
- Bachelor’s degree in Computer Science, Information Systems, or a related technical discipline, or equivalent practical experience.
- 6+ years of experience owning and operating large-scale, high-traffic, 24/7 production systems, ideally in cloud or cloud-native environments.
- Experience applying AI/LLM-powered tools to reliability engineering, including designing and building automation or internal tools using AI-assisted development tools (e.g., Claude Code).
- Solid foundations in Linux, networking, and distributed systems, with the ability to debug complex production issues end to end.
- Hands-on experience with incident response, troubleshooting, and performance optimization in distributed systems.
- Strong software engineering skills with experience building automation, tooling, or platforms in languages such as Python or Go.
- Experience operating or supporting open-source infrastructure components such as APISIX, Nginx, Kubernetes, Kafka, Elasticsearch, Redis, Consul, Etcd, Zookeeper, etc.
- Experience with observability and monitoring systems (Prometheus, Grafana, Zabbix, etc.) and performance analysis.
- Familiarity with Git, CI/CD pipelines, and configuration management tools (e.g., Ansible).
- A strong sense of ownership, a systematic approach to problem-solving, and a passion for making systems more reliable.
- Strong communication skills and the ability to collaborate effectively with geographically distributed teams.
Nice to Have
- Bilingual fluency in Mandarin and English.
- Kubernetes Administrator certification or equivalent real-world experience.
- Experience operating big data platforms (Hadoop, Yarn, HBase, Hive, Spark).
- Experience applying AI/LLM-powered
Similar jobs
Job Overview We are seeking a self-driven Principal Site Reliability Engineer with a strong technical background and excellent communication skills. This individual will lead the development, construction, and management…
Est. 140,000 USD
K2 is building the largest and highest-power satellites ever flown, unlocking performance levels previously out of reach across every orbit. Backed by over $1 billion in total funding from leading investors including Alt…
Est. 155,000 USD
The future of aerospace and defense starts here. Ursa Major was founded to revolutionize how America and its allies access and apply high-performance propulsion, from hypersonics to solid rocket motors, satellite maneuve…
Est. 144,000 USD
Toshiba Global Commerce Solutions is seeking a hands-on Lead Software Engineer to drive end-to-end solution delivery for major retail platforms. In this role, you will own hands-on implementation (feature development, te…
Est. 124,000 USD
Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century…
Join us in bringing joy to customer experience. Five9 is a leading provider of cloud contact center software, bringing the power of cloud innovation to customers worldwide. Living our values everyday results in our team-…
APPLICATIONS FROM OUTSIDE COLOMBIA WILL NOT BE CONSIDERED FOR THIS ROLERobots & Pencils is an applied AI engineering firm building the next frontier of business architecture. We design and ship AI co-workers that int…
Est. 124,000 USD
Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century…
Est. 56,000 USD
Locations: South Jordan, UT Salary: $56,000 Launch Your Career in Technology Every app, website, payment, and digital service relies on technology running smoothly behind the scenes. When something goes wrong, Production…
About Agoda At Agoda, we bridge the world through travel. Our story began in 2005, when two lifelong friends and entrepreneurs, driven by their passion for travel, launched Agoda to make it easier for everyone to explore…
Est. 175,000 USD
Company Introduction We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did I ever live without Coupang?” Born out of an obsession to make shopping, eating, and living…
Est. 140,000 USD
Who We Are Flagship Pioneering is a biotechnology company that invents and builds platform companies that change the world. We bring together the greatest scientific minds with entrepreneurial company builders and assemb…
Est. 105,000 USD
Step into a career with ASM, where cutting edge technology meets collaborative culture. For over 55 years ASM has been ahead of what’s next, at the forefront of innovation and what’s technologically possible. With more…
Est. 120,000 PLN
EBS is a rapidly expanding, entrepreneurial technology company and part of Alarm.com. Alarm.com is the leading cloud-based platform for smart security and the Internet of Things. More than 7.6+ million home and business…
At Anaplan, we are a team of innovators focused on optimizing business decision-making through our leading AI-infused scenario planning and analysis platform so our customers can outpace their competition and the market.…
The Joblogic Story Established in 1998, Joblogic is the UK’s #1 Field Service Management (FSM) software platform. We are a global business with offices in the UK, Pakistan, and Vietnam. Since our management buy-out in 20…
Est. 95,000 EUR
Why Sony Interactive Entertainment? Sony Interactive Entertainment isn’t just the Best Place to Play — it’s also the Best Place to Work. Sony Interactive Entertainment (SIE) is the company behind the PlayStation brand. A…
Senior Staff Engineer – Gateway Services The Gateway Services team is responsible for building Coupang’s traffic infrastructure layer using Service Mesh. Key aspects of the infrastructure layer include service discovery,…
Robots & Pencils is seeking an AI Engineer to design, build, and deploy production-grade AI systems that deliver measurable business value. This is a hands-on engineering role focused on implementation and delivery.…
Here’s a summary of the role: You’ll be a Senior Software Engineer helping shape how AI is built, deployed, and scaled across Diligent’s platform. Working at the intersection of Applied Science, Product, and Engineering,…
Est. 60,000 GBP
THE POSITION Our roster has an opening with your name on it We're looking for a Software Engineer to join one of our Core Marketing Platforms teams, where you'll build and optimize high-performance microservices that pow…
Developer Platform – Gateway Services – Job Description Staff Engineer – Gateway Services Level: L7-1 Location: Bangalore p
Est. 124,000 USD
CentralReach is a leading provider of autism and IDD care software for Applied Behavior Analysis (ABA), multidisciplinary therapy, and special education. Trusted by more than 200,000 users, we enable therapy providers, e…
Est. 165,000 USD
WHO WE ARE Zeta Global (NYSE: ZETA) is the AI-Powered Marketing Cloud that leverages advanced artificial intelligence (AI) and trillions of consumer signals to make it easier for marketers to acquire, grow, and retain cu…
Est. 250,000 USD
We’re looking for a Principal Architect to lead the design and delivery of complex, multi-domain systems spanning cloud, data, and AI. This role is ideal for a deeply experienced engineer who owns the hardest architectur…
Est. 124,000 USD
Vulcan Elements is manufacturing American rare-earth permanent magnets for a secure, resilient future. With a focus on national security and economic resiliency, we serve critical industries such as defense, aerospace, a…
Est. 250,000 USD
We’re looking for a Principal Architect to lead the design and delivery of complex, multi-domain systems spanning cloud, data, and AI. This role is ideal for a deeply experienced engineer who owns the hardest architectur…
Est. 124,000 USD
Are you ready to make an impact?West Monroe is seeking a Software Engineer to join the AI Assets Foundry Team, working with the internal Labs product team to develop proprietary software assets and AI capabilities that h…
S-RM is recruiting a Senior DevOps Engineer to play a key role in the development and maintenance of products for our Corporate Intelligence team.
Est. 144,000 USD
The Nuclear Company is the fastest growing AI tech-enabled startup in the nuclear and energy space, pioneering a fleet-scale approach to building the next generation of nuclear reactors. Through our design-once, build-ma…