Description
Objective of the Role
The Senior SRE Engineer is a highly experienced role responsible for leading the enhancement and maintenance of the reliability, availability, and performance of the company's IT infrastructure and applications. This role focuses on advanced system stability, efficiency, and scalability through sophisticated monitoring, automation, and incident response. The Senior SRE Engineer plays a critical role in driving strategic initiatives and mentoring junior engineers, contributing significantly to the success of IT operations.
Main Responsibilities
- Advanced System Monitoring: Design, implement, and maintain sophisticated monitoring solutions to ensure optimal health and performance of infrastructure and applications.
- Incident Response: Lead complex incident response activities, diagnose and resolve critical system reliability issues, and conduct thorough post-incident reviews to prevent recurrence.
- Automation and Scripting: Develop and implement advanced automation scripts and tools to enhance system reliability, operational efficiency, and scalability.
- Performance Analysis: Collect, analyze, and interpret complex performance data to identify trends, anomalies, and potential issues, providing strategic insights and recommendations.
- Documentation: Ensure comprehensive and up-to-date documentation of system configurations, processes, and procedures, and contribute to knowledge sharing within the team.
- Collaboration: Collaborate closely with cross-functional teams and departments to support and lead reliability engineering projects and initiatives.
- Mentorship: Provide advanced guidance and support to junior and mid-level engineers, fostering their technical growth and development.
- Security Compliance: Implement and enforce robust security measures to protect systems and ensure compliance with security policies and industry standards.
- Continuous Improvement: Drive continuous improvement initiatives, exploring and integrating new technologies and methodologies to enhance system reliability and performance.
- Strategic Leadership: Actively contribute to strategic planning and decision-making processes, leveraging expertise to influence the direction of IT infrastructure and operations.
- Capacity Planning: Conduct capacity planning to ensure systems can handle future growth and demand.
- Disaster Recovery: Develop and maintain disaster recovery plans to ensure business continuity in case of system failures.
- Change Management: Participate in change management processes to ensure smooth implementation of system updates and changes.
- Root Cause Analysis: Lead root cause analysis for major incidents, identifying underlying issues and implementing long-term solutions to prevent recurrence.
- Autonomous Work Culture: Promote and embody an autonomous work culture by taking initiative, being self-motivated, and collaborating effectively in an agile and lean environment.
- Spin Culture Ambassador: Embody and promote Spin's values in every action, fostering a positive
and inclusive work environment
Required Knowledge and Experience
Education: Bachelor's degree in computer science, Information Technology, or a related field, or
equivalent work experience.
Experience: Minimum of 7+ years of experience in site reliability engineering or related fields.
Deep understanding of system reliability concepts, including advanced monitoring, automation,
and incident response.
Proficiency with multiple scripting languages and automation tools.
Strong problem-solving and troubleshooting skills.
Experience with cloud platforms and containerization technologies.
Excellent communication and teamwork skills.
Proven leadership and mentorship abilities.
Strong strategic thinking and decision-making skills.
Adaptability: Willingness to learn and adapt to new technologies and processes
En Spin estamos comprometidos con construir un lugar de trabajo diverso e inclusivo.
Creemos en la igualdad de oportunidades y promovemos un entorno libre de discriminación por motivos de raza, origen nacional, género, identidad de género, orientación sexual, discapacidad, edad o cualquier otra condición legalmente protegida.
Similar jobs
Objective of the RoleThe Staff Infrastructure Engineer – SRE is a senior technical leader responsible for architecting and scaling the reliability, availability, and performance of critical infrastructure platforms. This…
Objective of the RolePlays a critical role in leading the design, development, and implementation of scalable and innovative technology solutions. This role drives technical excellence, fosters collaboration across teams…
Objective of the RolePlays a critical role in leading the design, development, and implementation of scalable and innovative technology solutions. This role drives technical excellence, fosters collaboration across teams…
Objective of the RoleLeads the design, development, and optimization of advanced data solutions, ensuring scalability, reliability, and alignment with business objectives. This role plays a critical part in defining data…
Objective of the Role Responsible for leading the analysis, documentation, and design of technology solutions that translate business requirements into clear technical and functional specifications. This role takes a pro…
What is Cobre, and what do we do?Cobre is Latin America’s leading instant b2b payments platform. We solve the region’s most complex money movement challenges by building advanced financialinfrastructure that enables comp…
Est. 90,000 GBP
Who are we? Ensono is a global technology services provider dedicated to helping organizations navigate the complexity of digital transformation. Through Ensono Product, Consulting & Technology, our dedicated consult…
Est. 65,000 EUR
En EPAM NEORIS, creemos que la transformación empieza por las personas. Hoy, como parte de EPAM, ampliamos nuestro alcance global y nuestras capacidades, pero mantenemos lo más importante: una cultura donde cada persona…
About The Role The Site Reliability Engineering (SRE) team architects, builds, and maintains the rock-solid infrastructure that applications rely on. At the Senior Level, you own reliability, performance, and cost outcom…
Job Title: Site Reliability Engineer (SRE)Key Skills: Kubernetes, AWS/Azure/GCP, Terraform, Python, Observability, CI/CDExperience: +6 YOE.Location: Costa Rica, Peru, Colombia, and Bolivia.Mode: Remote. We at Coforge are…
Est. 124,000 USD
Application Support Engineer (Site Reliability Engineer) Location: USAJob Type: Full-Time, no visa sponsorship available Coforge is seeking a Senior Application Support Engineer (SRE) to join our dynamic team of consulta…
About impact.com impact.com is the world’s leading commerce partnership marketing platform, transforming the way businesses grow by enabling them to discover, manage, and scale partnerships across the entire customer jou…
NEORIS is a Digital accelerator that helps companies enter the future, having 20 years of experience as Digital Partners of some of the largest companies in the world. We have more than 4,000 professionals in 11 countrie…
Senior Infrastructure Engineer Nuestro equipo de Hosting entrega las soluciones de InterSystems como servicios alojados (hosted) o administrados (managed services) en cualquier parte del mundo. A medida que cada vez más…
Est. 120,000 USD
We are representing a leading force in the decentralized exchange (DEX), and seeking a high-caliber technical leader to architect the backbone of a global financial ecosystem. In this role, you will bridge the gap betwee…
Est. 165,000 USD
About us: Working at Tech Holding isn't just a job, it's an opportunity to be a part of something bigger. We are a full-service consulting firm that was founded on the premise of delivering predictable outcomes and high-…
At NiCE, we don’t limit our challenges. We challenge our limits. Always. We’re ambitious. We’re game changers. And we play to win. We set the highest standards and execute beyond them. And if you’re like us, we can offer…
At Reltio®, an SAP Company, we believe data should fuel your success in the enterprise AI era. Our Context Intelligence Platform turns fragmented data into a trusted, connected context so AI agents and systems can act wi…
SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime…
Est. 90,000 GBP
At NiCE, we don’t limit our challenges. We challenge our limits. Always. We’re ambitious. We’re game changers. And we play to win. We set the highest standards and execute beyond them. And if you’re like us, we can offer…
Job Title: Sr Manager, Engineering - DevOps About Reltio At Reltio®, an SAP Company, we believe data should fuel your success in the enterprise AI era. Our Context Intelligence Platform turns fragmented data into a trust…
At Optimove, we believe people are capable of more than a single job description. You’re not hired just to fill a position- you’re empowered to shape it, grow it, and make it your own.We call this being Positionless.And…
Sr. Director, Site Reliability Engineering Coupang operates one of the largest and most complex technology platforms in the world. We are seeking a Senior Director, Site Reliability Engineering (Head of SRE) to define an…
Sr. Director, Site Reliability Engineering Coupang operates one of the largest and most complex technology platforms in the world. We are seeking a Senior Director, Site Reliability Engineering (Head of SRE) to define an…
Our hosting team delivers InterSystems’ solutions as hosted or managed services, anywhere in the world. As more and more clients change to hosted solutions, we are looking for a senior systems engineers to join the Chile…
About the Role The Site Reliability Engineering (SRE) team architects, builds, and maintains the rock-solid infrastructure that applications rely on. We work closely with development teams to ensure scalability, reliabil…
About SecurityScorecard: SecurityScorecard is the global leader in cybersecurity ratings, with over 12 million companies continuously rated, operating in 64 countries. Founded in 2013 by security and risk experts Dr. Ale…
Orion Innovation is a premier, award-winning, global business and technology services firm. Orion delivers game-changing business transformation and product development rooted in digital strategy, experience design, and…
Est. 124,000 USD
Are you driven to be an innovative Site Reliability Engineer, and looking to join a team where open collaboration, customer focus, and a commitment to excellence are core values? At Ivanti, we work passionately and authe…
About SecurityScorecard: SecurityScorecard is the global leader in cybersecurity ratings, with over 12 million companies continuously rated, operating in 64 countries. Founded in 2013 by security and risk experts Dr. Ale…