Description
Objective of the Role
The Staff Infrastructure Engineer – SRE is a senior technical leader responsible for architecting and scaling the reliability, availability, and performance of critical infrastructure platforms. This role combines deep technical expertise, cross-team influence, and a strategic mindset to drive high-impact initiatives across engineering teams, ensuring operational excellence and long-term system resilience.
Main Responsibilities
• Design and evolve complex infrastructure systems to ensure scalability, reliability, and security of platform services.
• Define long-term architectural vision and influence infrastructure roadmap across multiple teams and services.
• Lead strategic SRE initiatives, collaborating with cross-functional teams to improve system resiliency and reduce operational toil.
• Identify systemic issues and lead root cause analysis efforts, establishing long-term corrective measures.
• Champion best practices in observability, including metrics, logging, and alerting frameworks for production systems.
• Design and implement advanced automation solutions to optimize infrastructure performance and operational workflows.
• Partner with security and compliance teams to ensure infrastructure meets regulatory and organizational standards.
• Mentor senior and mid-level engineers, fostering knowledge sharing and technical growth across the organization.
• Serve as a technical advisor to leadership, providing insight on system health, risk, and architectural tradeoffs.
• Collaborate with Product, Engineering, and Data teams to support the needs of autonomous squads while maintaining infrastructure cohesion.
• Lead capacity planning efforts across critical services to anticipate future growth and ensure infrastructure scalability.
• Own and continuously improve disaster recovery strategies and testing plans to maintain business continuity.
• Contribute to and enforce standards for infrastructure documentation, design reviews, and change management processes.
• Foster a culture of reliability engineering, automation-first mindset, and operational accountability.
• Embody and promote Spin’s cultural values, acting as a role model of collaboration, innovation, and ownership.
• Promote an autonomous work culture by encouraging self-management, accountability, and proactive problem-solving among team members.
• Serve as a Spin Culture Ambassador to foster and maintain a positive, inclusive, and dynamic work environment that aligns with the company's values and culture.
Required Knowledge and Experience
• Bachelor’s degree in Computer Science, Information Systems, or equivalent practical experience.
• 7+ years of experience in infrastructure, site reliability, or platform engineering roles.
• Proven experience designing and operating reliable systems at scale, including distributed systems and cloud-native platforms.
• Advanced knowledge of infrastructure-as-code tools, container orchestration systems (e.g., Kubernetes), and CI/CD practices.
• Strong programming and scripting skills (e.g., Python, Go, Bash).
• Deep understanding of monitoring, alerting, and incident management practices.
• Experience working with cloud platforms (AWS, GCP, or Azure) and hybrid environments.
• Demonstrated ability to drive cross-team collaboration and influence engineering practices across domains.
• Strong leadership presence, excellent communication skills, and ability to present complex ideas to technical and non-technical audiences.
• High level of autonomy, initiative, and problem-solving skills in dynamic and fast-paced environments
En Spin estamos comprometidos con construir un lugar de trabajo diverso e inclusivo.
Creemos en la igualdad de oportunidades y promovemos un entorno libre de discriminación por motivos de raza, origen nacional, género, identidad de género, orientación sexual, discapacidad, edad o cualquier otra condición legalmente protegida.
Similar jobs
Objective of the RoleThe Senior SRE Engineer is a highly experienced role responsible for leading the enhancement and maintenance of the reliability, availability, and performance of the company's IT infrastructure and a…
Objective of the RolePlays a critical role in leading the design, development, and implementation of scalable and innovative technology solutions. This role drives technical excellence, fosters collaboration across teams…
Objective of the RolePlays a critical role in leading the design, development, and implementation of scalable and innovative technology solutions. This role drives technical excellence, fosters collaboration across teams…
Objective of the RoleLeads the design, development, and optimization of advanced data solutions, ensuring scalability, reliability, and alignment with business objectives. This role plays a critical part in defining data…
Objective of the Role Responsible for leading the analysis, documentation, and design of technology solutions that translate business requirements into clear technical and functional specifications. This role takes a pro…
Est. 65,000 EUR
En EPAM NEORIS, creemos que la transformación empieza por las personas. Hoy, como parte de EPAM, ampliamos nuestro alcance global y nuestras capacidades, pero mantenemos lo más importante: una cultura donde cada persona…
Job Title: Site Reliability Engineer (SRE)Key Skills: Kubernetes, AWS/Azure/GCP, Terraform, Python, Observability, CI/CDExperience: +6 YOE.Location: Costa Rica, Peru, Colombia, and Bolivia.Mode: Remote. We at Coforge are…
Orion Innovation is a premier, award-winning, global business and technology services firm. Orion delivers game-changing business transformation and product development rooted in digital strategy, experience design, and…
About impact.com impact.com is the world’s leading commerce partnership marketing platform, transforming the way businesses grow by enabling them to discover, manage, and scale partnerships across the entire customer jou…
About SecurityScorecard: SecurityScorecard is the global leader in cybersecurity ratings, with over 12 million companies continuously rated, operating in 64 countries. Founded in 2013 by security and risk experts Dr. Ale…
About SecurityScorecard: SecurityScorecard is the global leader in cybersecurity ratings, with over 12 million companies continuously rated, operating in 64 countries. Founded in 2013 by security and risk experts Dr. Ale…
SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime…
Est. 173,500 USD
About SecurityScorecard: SecurityScorecard is the global leader in cybersecurity ratings, with over 12 million companies continuously rated, operating in 64 countries. Founded in 2013 by security and risk experts Dr. Ale…
About SecurityScorecard: SecurityScorecard is the global leader in cybersecurity ratings, with over 12 million companies continuously rated, operating in 64 countries. Founded in 2013 by security and risk experts Dr. Ale…
Est. 173,500 USD
About SecurityScorecard: SecurityScorecard is the global leader in cybersecurity ratings, with over 12 million companies continuously rated, operating in 64 countries. Founded in 2013 by security and risk experts Dr. Ale…
Senior Infrastructure Engineer Nuestro equipo de Hosting entrega las soluciones de InterSystems como servicios alojados (hosted) o administrados (managed services) en cualquier parte del mundo. A medida que cada vez más…
Est. 124,000 USD
Application Support Engineer (Site Reliability Engineer) Location: USAJob Type: Full-Time, no visa sponsorship available Coforge is seeking a Senior Application Support Engineer (SRE) to join our dynamic team of consulta…
Our hosting team delivers InterSystems’ solutions as hosted or managed services, anywhere in the world. As more and more clients change to hosted solutions, we are looking for a senior systems engineers to join the Chile…
Est. 90,000 GBP
Who are we? Ensono is a global technology services provider dedicated to helping organizations navigate the complexity of digital transformation. Through Ensono Product, Consulting & Technology, our dedicated consult…
Est. 120,000 USD
We are representing a leading force in the decentralized exchange (DEX), and seeking a high-caliber technical leader to architect the backbone of a global financial ecosystem. In this role, you will bridge the gap betwee…
We are seeking a skilled and passionate Engineer to join our team to build and operate a Whole-of-Government (WoG) runtime platform. As a Site Reliability Engineer, you will be responsible for designing and operating Git…
Sr. Director, Site Reliability Engineering Coupang operates one of the largest and most complex technology platforms in the world. We are seeking a Senior Director, Site Reliability Engineering (Head of SRE) to define an…
Sr. Director, Site Reliability Engineering Coupang operates one of the largest and most complex technology platforms in the world. We are seeking a Senior Director, Site Reliability Engineering (Head of SRE) to define an…
At NiCE, we don’t limit our challenges. We challenge our limits. Always. We’re ambitious. We’re game changers. And we play to win. We set the highest standards and execute beyond them. And if you’re like us, we can offer…
Job Title: DevOps EngineerKey Skills: DevOps, Cloud Operations, AWS, TerraformLocation: BrazilMode: RemoteWe at Coforge are hiring DevOps Engineer with the following skill set.Key Responsibilities: Participate in a bi-we…
Est. 140,000 USD
BeyondTrust is a place where you can bring your purpose to life through the work that you do, creating a safer world through our cybersecurity SaaS portfolio. Our culture of flexibility, trust, and continual learning mea…
About The Role The Site Reliability Engineering (SRE) team architects, builds, and maintains the rock-solid infrastructure that applications rely on. At the Senior Level, you own reliability, performance, and cost outcom…
Orion Innovation is a premier, award-winning, global business and technology services firm. Orion delivers game-changing business transformation and product development rooted in digital strategy, experience design, and…
At OneSpan, we specialize in digital identity and anti-fraud solutions that create exceptional and secure experiences.We are looking for a Site Reliability Engineer to join our growing platform team in Delhi NCR. You wil…
At Reltio®, an SAP Company, we believe data should fuel your success in the enterprise AI era. Our Context Intelligence Platform turns fragmented data into a trusted, connected context so AI agents and systems can act wi…