inhousefyi
← Back to listings

Technical Program Manager

RadixArkPalo Alto, California, United States · Posted 4 months ago
Full-timeEst. 165,000 USD
Apply now

Description

About the Role

As a Technical Program Manager at RadixArk, you'll drive the execution of complex, cross-functional programs across our inference and training infrastructure. You'll partner closely with Product Management, Research, and Engineering to turn ambitious technical roadmaps into shipped reality, coordinating across kernel teams, distributed systems engineers, and external partners to deliver infrastructure that serves billions of tokens daily and coordinates 10,000+ GPU training runs.
This role is for someone who thrives at the intersection of deep technical understanding and rigorous program execution. You'll own the "how" and "when" of our most critical initiatives.

Key Responsibilities

Program Execution & Delivery

  • Drive end-to-end execution of large-scale, cross-functional programs spanning inference engines (e.g., SGLang), training frameworks (e.g., Miles), and hardware integration efforts.
  • Define program structure, including milestones, dependencies, critical paths, risks, and success criteria. Maintain a clear source of truth for status across all stakeholders.
  • Run design reviews, sprint planning, release readiness reviews, and post-mortems. Ensure decisions are documented and follow-ups are closed out.
  • Identify and unblock cross-team dependencies across kernel, runtime, scheduler, networking, and model teams before they become release blockers.
  • Drive release management for major versions, including changelog ownership, compatibility validation, partner rollout sequencing, and rollback planning.

Technical Coordination

  • Partner with Product Management to translate roadmap priorities into executable program plans, with clear scope, staffing, and timelines.
  • Work shoulder-to-shoulder with engineering leads on technical trade-off decisions; understand the architecture deeply enough to ask the right questions and surface hidden risks.
  • Coordinate hardware enablement programs with partners like Nvidia, Google, and AWS, including new accelerator bring-up, kernel co-development, and benchmark validation.
  • Manage integration programs with frontier AI labs and early adopters, ensuring technical requirements, SLAs, and feedback loops are well-defined.

Operational Excellence

  • Build and improve the engineering operating cadence, including standups, planning rituals, OKR tracking, dashboards, and reporting to leadership.
  • Establish metrics and instrumentation for program health such as velocity, defect rates, benchmark regressions, and customer-reported issues, and drive accountability against them.
  • Lead incident response coordination for production issues affecting partners; own root-cause review and corrective-action tracking.
  • Improve developer productivity by identifying and removing systemic friction in our build, test, and release pipelines.

Stakeholder Communication

  • Serve as the connective tissue between engineering, product, GTM, and external partners, ensuring everyone has the right information at the right altitude.
  • Produce clear, concise written updates for leadership and partners. Translate engineering progress into business-relevant signals.
  • Represent program status honestly, including risks and slips, with concrete mitigation plans.

Qualifications

Minimum Requirements

  • Bachelor's or Master's degree in Computer Science, Engineering, or a related technical field.
  • 4+ years of direct experience in Technical Program Management, Engineering Management, or a senior engineering role with significant program ownership, in a software or infrastructure company.
  • Strong technical fluency in systems software, distributed systems, or AI/ML infrastructure; able to read code, follow architecture discussions, and challenge technical assumptions productively.
  • Demonstrated track record shipping complex, multi-team programs on time, including managing dependencies, risks, and scope changes.
  • Excellent written and verbal communication skills; able to drive alignment across engineers, executives, and external partners.

Preferred (Bonus) Qualifications

  • Direct experience shipping AI/ML infrastructure such as inference engines, training frameworks, GPU kernels, distributed schedulers, or model serving platforms.
  • Hands-on coding background (Python, C++, CUDA) and comfort working in engineering codebases, including reading PRs, running benchmarks, and reproducing issues.
  • Experience coordinating with hardware vendors (Nvidia, AMD, Google TPU, AWS Trainium/Inferentia) on enablement or co-engineering programs.
  • Experience driving open-source release programs or working in OSS communities, including issue triage, RFC processes, and contributor coordination.
  • Familiarity with release engineering, CI/CD systems, and observability tooling for large-scale distributed systems.
  • Experience supporting B2B or developer-facing products with enterprise SLAs.

About RadixArk

RadixArk is an infrastructure-first company built by engineers who've shipped production AI systems, created SGLang (30K+ GitHub stars, the fastest open LLM serving engine), and developed Miles (our large-scale RL framework). Founded by AI infrastructure veterans from xAI and NVIDIA, we're on a mission to democratize frontier-level AI infrastructure by building world-class open systems for inference and training. Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs.

Compensation

We offer competitive compensation with equity, comprehensive health benefits, and flexible work arrangements. Compensation is determined by location, level, and experience.

Equal Opportunity

RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Similar jobs

Product Manager7 months ago
RadixArkPalo Alto, California, United States

Est. 144,000 USD

Key Responsibilities Product Strategy & Roadmap Define, prioritize, and drive the product roadmap for inference and training infrastructure. Stay ahead of AI trends, including new model architectures, hardware optimi…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role RadixArk is seeking experienced product-focused engineers to join our team in building the developer-facing surfaces of our inference and training infrastructure. As a Member of Technical Staff — Product,…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role RadixArk is seeking a Developer Advocate to build and engage our technical community around SGLang, Miles, and our open source infrastructure. SGLang already has 30K+ GitHub stars and serves billions of to…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role RadixArk is seeking a Member of Technical Staff — Kernel / Compiler / Communication to push the limits of performance for frontier AI systems. You will work at the lowest layers of the stack — kernels, run…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role RadixArk is looking for a Member of Technical Staff Cluster Infrastructure to architect and scale the core compute platform that powers frontier-level AI training and inference. You will design and operate…

Full-time
RadixArkPalo Alto, California, United States

Est. 140,000 USD

About The Role RadixArk is launching a full-time, paid, 1-year residency program for aspiring AI infrastructure engineers. You'll rotate across inference, training, kernels, compilers, and cluster infrastructure, working…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role RadixArk is seeking a Member of Technical Staff, Developer Technology (DevTech) to make LLM inference and training dramatically faster, cheaper, and more accessible on modern GPU hardware. Our systems sit…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role RadixArk is seeking a Member of Technical Staff — Inference to push the limits of large-scale AI inference. You will work on the core systems that serve frontier models at scale, optimizing performance, la…

Full-time
RadixArkPalo Alto, California, United States

Est. 155,000 USD

About the Role RadixArk is seeking a Product Marketing Manager to own how SGLang, Miles, and our open source infrastructure are positioned and perceived across the market. SGLang already has 20K+ GitHub stars and serves…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role RadixArk is looking for a Member of Technical Staff — Backend/API Platform Engineer to build the API layer, control plane, and platform services that power SGLang and Miles in production. You'll design and…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role RadixArk is seeking a Member of Technical Staff — Inference-Multimodal & Diffusion to advance the frontier of generative modeling. You will work on cutting-edge diffusion and flow-based models for imag…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role RadixArk is seeking a Member of Technical Staff - Inference-Multi-Hardware to push the limits of performance for frontier AI systems. Most performance engineering assumes a single vendor's stack. This role…

Full-time
RadixArkPalo Alto, California, United States

Est. 150,000 USD

About the Role We're looking for a Head of Business Development to build the BD function at RadixArk from the ground up. The BD team is the institutional memory of this company — maintaining active relationships across e…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role RadixArk is hiring a Member of Technical Staff — CI / Infrastructure to own the infrastructure that keeps SGLang moving. Our CI system runs 300+ GPU tests across NVIDIA, AMD, Intel, and Ascend hardware poo…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role RadixArk is hiring a Member of Technical Staff — Performance in Palo Alto, CA — someone who can push LLM inference and training systems to the limit across real production workloads. You’ll work on the per…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role As a Member of Technical Staff, Training, you will design, build, and operate the distributed systems behind large-scale model post-training — spanning training, inference, and orchestration, with a focus…

Full-time
RadixArkPalo Alto, California, United States

Est. 300,000 USD

About the Role RadixArk is looking for a Member of Technical Staff — TPU Systems to build high-performance inference and training systems using JAX, XLA, and Pallas. You'll push model workloads to their limits on TPU har…

Full-time
RadixArkPalo Alto, California, United States

Est. 159,000 USD

About the Role We're looking for a hands-on Talent Operations Specialist to build and run the machinery behind talent and people ops as we scale. This isn't a traditional HR generalist role - it's for someone who treats…

Full-time
Visual Designer2 months ago
RadixArkPalo Alto, California, United States

Est. 115,000 USD

About the Role RadixArk builds the open-source AI infrastructure behind SGLang and Miles, used by developers and enterprises around the world. We're looking for a visual designer to join our design team and to give our b…

Full-time
Lightning AINew York, New York, United States

Est. 190,000 USD

Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with…

Full-time
XPENGSanta Clara, California, United States

Est. 289,800 USD

XPENG is a leading smart technology company at the forefront of innovation, integrating advanced AI and autonomous driving technologies into its vehicles, including electric vehicles (EVs), electric vertical take-off and…

Full-time
XPENGSanta Clara, California, United States

Est. 235,200 USD

XPENG is a leading smart technology company at the forefront of innovation, integrating advanced AI and autonomous driving technologies into its vehicles, including electric vehicles (EVs), electric vertical take-off and…

Full-time
Arkana LaboratoriesLittle Rock, Arkansas, United States

Est. 140,000 USD

Who we are: At Arkana Laboratories, everyone has an important role to fill. Come join us and be a part of a team dedicated to making life better for those who need it most. This place is packed with super-smart people wh…

Full-time
AnthropicSan Francisco, California, United States

Est. 400,000 USD

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of co…

Full-time
AnthropicSan Francisco, California, United States

Est. 600,000 USD

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of co…

Full-time
XPENGSanta Clara, California, United States

Est. 328,650 USD

XPENG is a leading smart technology company at the forefront of innovation, integrating advanced AI and autonomous driving technologies into its vehicles, including electric vehicles (EVs), electric vertical take-off and…

Full-time

Est. 175,000 USD

We’re looking for a Staff AI Engineer to lead the design and delivery of AI/ML systems. This role is ideal for an experienced engineer who thrives on architectural decisions, can confidently own systems end-to-end, and c…

Full-timeRemote
GradialSeattle, Washington, United States

Est. 205,000 USD

Gradial is the marketing operations system of work that helps marketers and creatives move from idea to execution faster. Our platform orchestrates across martech stacks, workflows, and people to automate marketing execu…

Full-time
AnthropicSan Francisco, California, United States

Est. 400,000 USD

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of co…

Full-time
AnthropicSan Francisco, California, United States

Est. 400,000 USD

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of co…

Full-time