Member of Technical Staff — Reliability-CI Infrastructure
RadixArkPalo Alto, California, United States · Posted 4 months agoDescription
About the Role
RadixArk is hiring a Member of Technical Staff — CI / Infrastructure to own the infrastructure that keeps SGLang moving. Our CI system runs 300+ GPU tests across NVIDIA, AMD, Intel, and Ascend hardware pools, gating every commit to one of the fastest-growing open-source LLM inference engines. When CI is green and fast, 100+ contributors ship with confidence. When it isn't, the entire project stalls. That bottleneck is your problem to solve.
You won't just maintain pipelines — you'll architect them. You'll replace brittle static thresholds with regression-based detection, harden runners against supply-chain attacks from fork PRs, and cut cycle times so contributors get feedback in minutes, not hours. You'll work directly with core maintainers, hardware partners, and the open-source community to keep the system that gates every merge request trustworthy, fast, and secure.
This is not a role for someone who wants to write CI YAML and walk away. It's for an engineer who treats CI infrastructure the way we treat serving infrastructure — as a system worth designing well.
What You’ll Do
- Own CI reliability end-to-end — triage failures, distinguish real regressions from flaky tests and infra issues, keep main green
- Build regression-based CI — replace hardcoded static thresholds with automated baseline comparison (metrics pipeline, durable storage, detection logic)
- Harden runner infrastructure — ephemeral runners, container isolation, security hardening for fork PR execution
- Cut CI time — right-size eval suites, deduplicate server startups, separate PR smoke tests from nightly full runs
- Improve developer experience — faster feedback, clearer failure messages, workflow orchestration
Requirements
-
- 3+ years operating CI/CD at scale (GitHub Actions, Buildkite, Jenkins, GitLab CI, or similar)
- Deep Linux, Docker, GPU computing knowledge
- Self-hosted runner management experience
- Strong Bash and Python
- Security mindset — CI supply chain risks, fork PR attack vectors, runner hardening
- NVIDIA GPU drivers, CUDA, NCCL, InfiniBand/RDMA experience in CI contexts
- Familiarity with ML inference workloads (model loading, KV cache, quantization)
Nice to Have
- Large open-source project CI experience (100+ contributors)
- AMD ROCm or Intel XPU CI pipelines
What Success Looks Like
- Day 20 — Full CI landscape understood, daily triage taken over, top recurring flaky tests fixed, PR CI time reduced 30%+
- Day 40 — Regression-based checks live on nightly CI, ephemeral runner prototype deployed, runner isolation in place
- Day 60 — Zero flaky tests. Main CI 100% green when no real regression exists
How to Apply
Reach out via Slack or email. CI fix PRs to major open-source projects are worth more than a resume.
About RadixArk
RadixArk is an infrastructure-first company built by engineers who've shipped production AI systems, created SGLang (30K+ GitHub stars, the fastest open LLM serving engine), and developed Miles (our large-scale RL framework). Founded by AI infrastructure veterans from xAI and NVIDIA, we're on a mission to democratize frontier-level AI infrastructure by building world-class open systems for inference and training. Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs.
Compensation
Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
Equal Opportunity
RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
Similar jobs
Est. 300,000 USD
About the Role RadixArk is seeking a Member of Technical Staff — Kernel / Compiler / Communication to push the limits of performance for frontier AI systems. You will work at the lowest layers of the stack — kernels, run…
Est. 300,000 USD
About the Role RadixArk is looking for a Member of Technical Staff Cluster Infrastructure to architect and scale the core compute platform that powers frontier-level AI training and inference. You will design and operate…
Est. 300,000 USD
About the Role RadixArk is seeking a Developer Advocate to build and engage our technical community around SGLang, Miles, and our open source infrastructure. SGLang already has 30K+ GitHub stars and serves billions of to…
Est. 300,000 USD
About the Role RadixArk is seeking a Member of Technical Staff — Inference to push the limits of large-scale AI inference. You will work on the core systems that serve frontier models at scale, optimizing performance, la…
Est. 300,000 USD
About the Role RadixArk is seeking a Member of Technical Staff - Inference-Multi-Hardware to push the limits of performance for frontier AI systems. Most performance engineering assumes a single vendor's stack. This role…
Est. 300,000 USD
About the Role RadixArk is seeking a Member of Technical Staff, Developer Technology (DevTech) to make LLM inference and training dramatically faster, cheaper, and more accessible on modern GPU hardware. Our systems sit…
Est. 300,000 USD
About the Role RadixArk is looking for a Member of Technical Staff — Backend/API Platform Engineer to build the API layer, control plane, and platform services that power SGLang and Miles in production. You'll design and…
Est. 300,000 USD
About the Role RadixArk is hiring a Member of Technical Staff — Performance in Palo Alto, CA — someone who can push LLM inference and training systems to the limit across real production workloads. You’ll work on the per…
Est. 300,000 USD
About the Role RadixArk is seeking experienced product-focused engineers to join our team in building the developer-facing surfaces of our inference and training infrastructure. As a Member of Technical Staff — Product,…
Est. 140,000 USD
About The Role RadixArk is launching a full-time, paid, 1-year residency program for aspiring AI infrastructure engineers. You'll rotate across inference, training, kernels, compilers, and cluster infrastructure, working…
Est. 300,000 USD
About the Role RadixArk is seeking a Member of Technical Staff — Inference-Multimodal & Diffusion to advance the frontier of generative modeling. You will work on cutting-edge diffusion and flow-based models for imag…
Est. 165,000 USD
About the Role As a Technical Program Manager at RadixArk, you'll drive the execution of complex, cross-functional programs across our inference and training infrastructure. You'll partner closely with Product Management…
Est. 155,000 USD
About the Role RadixArk is seeking a Product Marketing Manager to own how SGLang, Miles, and our open source infrastructure are positioned and perceived across the market. SGLang already has 20K+ GitHub stars and serves…
Est. 144,000 USD
Key Responsibilities Product Strategy & Roadmap Define, prioritize, and drive the product roadmap for inference and training infrastructure. Stay ahead of AI trends, including new model architectures, hardware optimi…
Est. 150,000 USD
About the Role We're looking for a Head of Business Development to build the BD function at RadixArk from the ground up. The BD team is the institutional memory of this company — maintaining active relationships across e…
Est. 300,000 USD
About the Role As a Member of Technical Staff, Training, you will design, build, and operate the distributed systems behind large-scale model post-training — spanning training, inference, and orchestration, with a focus…
Est. 159,000 USD
About the Role We're looking for a hands-on Talent Operations Specialist to build and run the machinery behind talent and people ops as we scale. This isn't a traditional HR generalist role - it's for someone who treats…
Est. 300,000 USD
About the Role RadixArk is looking for a Member of Technical Staff — TPU Systems to build high-performance inference and training systems using JAX, XLA, and Pallas. You'll push model workloads to their limits on TPU har…
Est. 115,000 USD
About the Role RadixArk builds the open-source AI infrastructure behind SGLang and Miles, used by developers and enterprises around the world. We're looking for a visual designer to join our design team and to give our b…
Est. 140,000 USD
Who we are: At Arkana Laboratories, everyone has an important role to fill. Come join us and be a part of a team dedicated to making life better for those who need it most. This place is packed with super-smart people wh…
Est. 120,000 USD
Why Cast AI? Cast AI is an automation platform that operates cloud-native and AI infrastructure at scale. By embedding autonomous decision-making directly into Kubernetes and cloud environments, Cast AI continuously opti…
Est. 120,000 EUR
Location & work modality: EMEA (remote) Start: ASAP Type of Contract: Permanent, full-time About Radian Arc Radian Arc, now part of InferX, Submer's AI cloud and GPU infrastructure platform, provides an infrastructur…
Est. 310,000 USD
At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, custom training infrastructur…
Est. 120,000 EUR
Location & work modality: Europe (remote) Start: ASAP Type of Contract: Full-time (permanent or freeelance) About Radian Arc Radian Arc, now part of InferX, Submer's AI cloud and GPU infrastructure platform, provides…
Est. 286,000 USD
Biohub is the first large-scale initiative bringing frontier AI models, massive compute, and frontier experimental capabilities under one roof. We're building a general-purpose system to accelerate scientific discovery,…
Est. 315,000 USD
About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…
Est. 310,000 USD
At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, custom training infrastructur…
Est. 120,000 EUR
N-iX is a global software development company founded in 2002, connecting over 2,400+ tech professionals across 40+ countries. We deliver innovative technology solutions in cloud computing, data analytics, AI, embedded s…
Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with…
Est. 180,000 USD
Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with…