inhousefyi
← Back to listings

Member of Technical Staff — Inference-Kernel, Compiler & Communication

RadixArkPalo Alto, California, United States · Posted 6 months ago
Full-timeEst. 300,000 USD
Apply now

Description

About the Role

RadixArk is seeking a Member of Technical Staff — Kernel / Compiler / Communication to push the limits of performance for frontier AI systems.

You will work at the lowest layers of the stack — kernels, runtimes, compilers, and communication libraries — to unlock maximum efficiency from modern accelerators and interconnects.

This role is critical to scaling training and inference across thousands of GPUs, where microseconds and memory bandwidth matter. Your work will directly shape the performance envelope of next-generation AI systems.

This is a deeply technical role for engineers who enjoy working close to hardware and solving performance problems that most engineers never encounter.

Requirements

  • 5+ years of experience in systems, compiler, or performance engineering

  • Strong expertise in CUDA or accelerator programming

  • Deep understanding of GPU architecture and memory hierarchy

  • Experience writing or optimizing high-performance kernels

  • Strong background in compilers, runtimes, or code generation

  • Experience with distributed communication libraries (NCCL, MPI, RCCL, etc.)

  • Solid knowledge of networking and interconnect technologies

  • Proficiency in C++ and Python

  • Strong debugging and profiling skills at system level

Strong Plus

  • Experience with Triton, TVM, XLA, or MLIR

  • Experience building compiler passes or IR transformations

  • Familiarity with NVLink, InfiniBand, or RDMA

  • Experience optimizing collective communication at scale

  • Background in HPC or performance-critical systems

  • Contributions to kernel/compiler/ML systems open source

  • Experience scaling workloads to 1000+ GPUs

  • Experience with mixed-precision or quantized kernels

Responsibilities

  • Design and implement high-performance kernels for AI workloads

  • Optimize compiler and runtime stacks for ML systems

  • Improve communication efficiency across large GPU clusters

  • Reduce latency and increase throughput for distributed workloads

  • Profile and eliminate system bottlenecks across the stack

  • Collaborate with training and inference teams on performance optimization

  • Develop tooling for profiling and performance analysis

  • Contribute to long-term architecture for performance-critical systems

  • Push the limits of hardware–software co-design

About RadixArk

RadixArk is an infrastructure-first company built by engineers who've shipped production AI systems, created SGLang (30K+ GitHub stars, the fastest open LLM serving engine), and developed Miles (our large-scale RL framework). Founded by AI infrastructure veterans from xAI and NVIDIA, we're on a mission to democratize frontier-level AI infrastructure by building world-class open systems for inference and training. Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs.

Compensation

Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.

Equal Opportunity

RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Similar jobs

RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role RadixArk is seeking a Member of Technical Staff — Inference to push the limits of large-scale AI inference. You will work on the core systems that serve frontier models at scale, optimizing performance, la…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role RadixArk is looking for a Member of Technical Staff Cluster Infrastructure to architect and scale the core compute platform that powers frontier-level AI training and inference. You will design and operate…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role RadixArk is seeking a Member of Technical Staff - Inference-Multi-Hardware to push the limits of performance for frontier AI systems. Most performance engineering assumes a single vendor's stack. This role…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role RadixArk is seeking a Member of Technical Staff — Inference-Multimodal & Diffusion to advance the frontier of generative modeling. You will work on cutting-edge diffusion and flow-based models for imag…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role RadixArk is seeking a Member of Technical Staff, Developer Technology (DevTech) to make LLM inference and training dramatically faster, cheaper, and more accessible on modern GPU hardware. Our systems sit…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role RadixArk is hiring a Member of Technical Staff — Performance in Palo Alto, CA — someone who can push LLM inference and training systems to the limit across real production workloads. You’ll work on the per…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role RadixArk is seeking a Developer Advocate to build and engage our technical community around SGLang, Miles, and our open source infrastructure. SGLang already has 30K+ GitHub stars and serves billions of to…

Full-time
RadixArkPalo Alto, California, United States

Est. 165,000 USD

About the Role As a Technical Program Manager at RadixArk, you'll drive the execution of complex, cross-functional programs across our inference and training infrastructure. You'll partner closely with Product Management…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role RadixArk is seeking experienced product-focused engineers to join our team in building the developer-facing surfaces of our inference and training infrastructure. As a Member of Technical Staff — Product,…

Full-time
Product Manager7 months ago
RadixArkPalo Alto, California, United States

Est. 144,000 USD

Key Responsibilities Product Strategy & Roadmap Define, prioritize, and drive the product roadmap for inference and training infrastructure. Stay ahead of AI trends, including new model architectures, hardware optimi…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role RadixArk is looking for a Member of Technical Staff — Backend/API Platform Engineer to build the API layer, control plane, and platform services that power SGLang and Miles in production. You'll design and…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role RadixArk is hiring a Member of Technical Staff — CI / Infrastructure to own the infrastructure that keeps SGLang moving. Our CI system runs 300+ GPU tests across NVIDIA, AMD, Intel, and Ascend hardware poo…

Full-time
RadixArkPalo Alto, California, United States

Est. 140,000 USD

About The Role RadixArk is launching a full-time, paid, 1-year residency program for aspiring AI infrastructure engineers. You'll rotate across inference, training, kernels, compilers, and cluster infrastructure, working…

Full-time
RadixArkPalo Alto, California, United States

Est. 150,000 USD

About the Role We're looking for a Head of Business Development to build the BD function at RadixArk from the ground up. The BD team is the institutional memory of this company — maintaining active relationships across e…

Full-time
RadixArkPalo Alto, California, United States

Est. 155,000 USD

About the Role RadixArk is seeking a Product Marketing Manager to own how SGLang, Miles, and our open source infrastructure are positioned and perceived across the market. SGLang already has 20K+ GitHub stars and serves…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role RadixArk is looking for a Member of Technical Staff — TPU Systems to build high-performance inference and training systems using JAX, XLA, and Pallas. You'll push model workloads to their limits on TPU har…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role As a Member of Technical Staff, Training, you will design, build, and operate the distributed systems behind large-scale model post-training — spanning training, inference, and orchestration, with a focus…

Full-time
RadixArkPalo Alto, California, United States

Est. 159,000 USD

About the Role We're looking for a hands-on Talent Operations Specialist to build and run the machinery behind talent and people ops as we scale. This isn't a traditional HR generalist role - it's for someone who treats…

Full-time
Visual Designer2 months ago
RadixArkPalo Alto, California, United States

Est. 115,000 USD

About the Role RadixArk builds the open-source AI infrastructure behind SGLang and Miles, used by developers and enterprises around the world. We're looking for a visual designer to join our design team and to give our b…

Full-time
River AI Inc.Austin, Texas, United States

Est. 310,000 USD

At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, custom training infrastructur…

Full-time
SUBMERRemote

Est. 120,000 EUR

Location & work modality: EMEA (remote) Start: ASAP Type of Contract: Permanent, full-time About Radian Arc Radian Arc, now part of InferX, Submer's AI cloud and GPU infrastructure platform, provides an infrastructur…

Full-timeRemote
AnthropicSan Francisco, California, United States

Est. 600,000 USD

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of co…

Full-time
River AI Inc.Austin, Texas, United States

Est. 310,000 USD

At River, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, custom training infrastructure,…

Full-time
SUBMERRemote

Est. 120,000 EUR

Location & work modality: Europe (remote) Start: ASAP Type of Contract: Full-time (permanent or freeelance) About Radian Arc Radian Arc, now part of InferX, Submer's AI cloud and GPU infrastructure platform, provides…

Full-timeRemote
River AI Inc.Austin, Texas, United States

Est. 310,000 USD

At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, custom training infrastructur…

Full-time
AnthropicRemote

Est. 445,000 USD

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of co…

Full-timeRemote
River AI Inc.Palo Alto, California, United States

Est. 310,000 USD

At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, custom training infrastructur…

Full-time
xAIPalo Alto, California, United States

Est. 310,000 USD

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organi…

Full-time
Arkana LaboratoriesLittle Rock, Arkansas, United States

Est. 140,000 USD

Who we are: At Arkana Laboratories, everyone has an important role to fill. Come join us and be a part of a team dedicated to making life better for those who need it most. This place is packed with super-smart people wh…

Full-time
EnCharge AIGermany

Est. 80,000 EUR

Research Engineer, Applied AI Location: Germany About EnCharge AI: EnCharge AI is building the next generation AI platform. Our novel in-memory-computing architecture delivers a 10x step-function improvement in compute e…

Full-time