inhousefyi
← Back to listings

Staff Storage Platform Engineer (AI Storage) - Radian Arc

SUBMERRemote · Posted 1 month ago
Full-timeRemoteEst. 120,000 EUR
Apply now

Description

About Radian Arc.
Radian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. Our teams across the USA, Australia, Central Europe, Malaysia, Singapore and Japan offer telecom operators a GPU-based edge computing platform without the need for capital expenditure, facilitating low latency and improved economics for value-added services and the monetization of 5G investments.

What impact you will have

Mission: Design, build, and operate the AI storage layer powering large-scale GPU infrastructure, enabling datasets, model artifacts, checkpoints, and inference state to be delivered to compute clusters with extremely high throughput and predictable latency.

You will play a key role in architecting and evolving the storage platform across edge and core deployments, supporting the full lifecycle of AI workloads including distributed inference, fine-tuning, and large-scale model training. The role spans multiple storage architectures used across the platform, including hyperconverged storage currently based on StorPool, local NVMe storage for latency-sensitive workloads and edge deployments, and disaggregated AI storage platforms such as VAST Data and Weka.

As the first dedicated storage platform role in the organization, this position combines Staff-level architectural ownership, technical direction, and cross-functional influence with hands-on execution across storage design, deployment, performance engineering, troubleshooting, platform integration, and operational improvement.

A key responsibility of this role is designing and optimizing the storage architecture underlying distributed inference stacks such as NVIDIA Dynamo, llm-d, or similar inference orchestration frameworks. This includes ensuring that storage systems efficiently support inference workloads through optimized dataset access, model artifact distribution, checkpoint handling, and KV-cache persistence. You will design scalable storage systems capable of feeding thousands of GPUs while balancing throughput, latency, resilience, and cost efficiency, and work closely with compute, networking, and platform engineering teams to ensure seamless integration with the platform orchestration layer.

Because this is currently the primary storage platform role in the company, the position is intentionally hybrid: you are expected to operate at L6 / Staff in terms of long-term design, standards, cross-team influence, and platform direction, while also directly executing critical storage work that, in a larger organization, would be distributed across multiple engineers.

What you'll need

Core Experience

  • Strong hands-on experience designing and operating distributed storage systems for high-performance compute environments.
  • Proven experience designing storage architectures for large-scale AI inference or training platforms, including dataset distribution, checkpointing, and KV-cache storage patterns.
  • Deep knowledge of the Linux storage and I/O stack.
  • Strong understanding of AI workload data access patterns.
  • Experience optimizing storage for GPU-accelerated workloads.
  • WEKA Data Platform (an enterprise high-performance storage system for AI and HPC)
  • Familiarity with Kubernetes storage integrations such as CSI.
  • Experience operating large-scale storage clusters.
  • Experience owning both architecture and direct implementation in lean or fast-scaling environments is strongly preferred.

Advanced AI Storage Expertise

The candidate should have deep expertise in designing and operating storage platforms optimized for GPU-heavy environments and distributed AI workloads.

This includes a strong understanding of how training, fine-tuning, and inference systems interact with storage, and how storage architecture affects throughput, latency, concurrency, checkpoint recovery, dataset distribution, and serving performance.

Relevant expertise includes:

  • Strong understanding of storage access patterns for distributed inference and training.
  • Experience designing storage platforms that support large dataset ingestion and model artifact distribution at scale.
  • Practical experience tuning storage architectures for checkpointing, distributed file access, object access, and high-concurrency inference.
  • Familiarity with storage patterns for KV-cache persistence and retrieval.
  • Experience optimizing data locality and reducing unnecessary network movement between storage and compute.
  • Understanding of how storage performance affects large-scale AI frameworks, model-serving systems, and inference orchestration layers.

Systems & Troubleshooting

  • Ability to debug complex cross-layer issues spanning:
    • Storage hardware,
    • Networking,
    • Linux kernel and I/O paths,
    • Filesystems,
    • Object and block storage layers,
    • Kubernetes integrations,
    • Distributed workload behavior.
  • Strong knowledge of storage hardware, NVMe devices, storage fabrics, and high-performance data paths.
  • Experience designing storage observability systems.
  • Strong ability to act as the senior escalation point for ambiguous, high-impact, and multi-domain technical issues.

Automation

  • Strong automation skills using Python and/or Bash.
  • Experience applying software engineering practices to storage automation and operational tooling.
  • Experience building reusable tooling, standards, validation patterns, or lifecycle automation that increase leverage across teams.

Leadership

  • Proven ability to lead complex technical initiatives across teams.
  • Comfortable collaborating across engineering, operations, deployment teams, vendors, and platform stakeholders.
  • Strong systems-level thinking balancing performance, reliability, scalability, operability, and cost efficiency.
  • Demonstrated ability to set architectural direction and drive adoption of engineering standards across an organization.
  • Proven ability to lead through technical influence across multiple teams and domains, without relying on formal people management authority.
  • Strong mentoring capability and ability to raise the technical level of adjacent engineering teams.
  • Able to balance short-term execution needs with long-term platform design, operational sustainability, and cost efficiency.

What you’ll do

Storage Architecture

  • Design scalable AI storage architectures supporting both edge and core deployments.
  • Define storage strategies for distributed inference, fine-tuning, and training workloads.
  • Architect solutions across multiple storage models:
    • Hyperconverged infrastructure such as StorPool,
    • Local NVMe storage,
    • Disaggregated storage systems such as VAST, Weka, and related architectures.
  • Define reference architectures, design principles, and reusable patterns for storage platforms so future deployments follow standards rather than one-off implementations.
  • Evaluate trade-offs across throughput, latency, resilience, data locality, cost, and operability, and make clear recommendations to engineering and leadership.
  • Influence the long-term storage roadmap, including architecture choices for edge, core, hyperconverged, and disaggregated environments.

AI Workload Optimization

  • Optimize storage throughput and latency for GPU-heavy clusters.
  • Design data locality strategies to minimize dataset movement across the network.
  • Benchmark storage performance under real AI workloads.
  • Optimize I/O patterns for large dataset ingestion, checkpointing, and model artifact distribution.
  • Work directly with compute teams to ensure storage architecture matches the access patterns of distributed training, fine-tuning, and inference frameworks.
  • Establish performance baselines and validation methods so storage platforms are tested against realistic AI workload behavior rather than only synthetic benchmarks.

Platform Integration

  • Implement and maintain CSI drivers.
  • Integrate storage platforms with Kubernetes and orchestration systems.
  • Integrate block, object, and shared file storage into the platform.
  • Design multi-tenant storage architectures supporting isolated workloads.
  • Ensure storage capabilities are correctly exposed into platform services, workload orchestration, and lifecycle automation.
  • Define standards for how storage should be integrated into Kubernetes-based and platform-managed environments across different deployment models.

Distributed Storage Systems

  • Co

Similar jobs

SUBMERRemote

Est. 80,000 EUR

About Radian Arc.Radian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. Our teams…

Full-timeRemote
Five9Bengaluru, India

Join us in bringing joy to customer experience. Five9 is a leading provider of cloud contact center software, bringing the power of cloud innovation to customers worldwide. Living our values everyday results in our team-…

Full-time
EqvilentRemote

Est. 120,000 USD

We build and operate core infrastructure across multiple data centers supporting a high-frequency trading environment. Our storage platform underpins data processing, analytics, and compute-heavy workloads across the com…

Full-timeRemote
Lightning AILondon, England, United Kingdom

Est. 200,000 USD

Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with…

Full-time
VectraAustin, Texas, United States

Est. 175,000 USD

Vectra® is the leader in AI-driven threat detection and response for hybrid and multi-cloud enterprises. The Vectra AI Platform delivers integrated signal across public cloud, SaaS, identity, and data center networks in…

Full-time
NscaleUnited Kingdom

Est. 120,000 GBP

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…

Full-time
NscaleUnited Kingdom

Est. 120,000 GBP

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…

Full-time
SUBMERRemote

Est. 120,000 EUR

Location & work modality: Europe/ Remote Start: Aug 2026 Type of Contract: Full time or Contract About Radian Arc Radian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificia…

Full-timeRemote
NebiusRemote

Est. 175,000 USD

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to…

Full-timeRemote
Together AIRemote

Est. 275,000 USD

About the Role In this role, you will operate, scale, and optimize multi-petabyte storage systems purpose-built for the world’s largest AI training and inference workloads. You’ll manage and scale high-performance parall…

Full-timeRemote
Astera LabsSan Jose, California, United States

Est. 141,000 USD

Astera Labs (NASDAQ: ALAB) provides rack-scale AI infrastructure through purpose-built connectivity solutions. By collaborating with hyperscalers and ecosystem partners, Astera Labs enables organizations to unlock the fu…

Full-time
Cloud Architect17 days ago
Robots and PencilsHouston, Texas, United States

Est. 141,000 USD

Cloud Architect Location: Houston, TX This is a 4 month assignment with potential to extend Houston, local OR be able to travel to Houston, TX to be at customer site every other week for a 4-5days/week Company Overview R…

Full-time
Squarepoint CapitalMontreal, Quebec, Canada

Job Summary Squarepoint is looking for a Platform Storage Specialist to join our growing global team. The candidate will work alongside our team to design, build, and maintain enterprise-grade storage services that are c…

Full-time
NscaleUnited States

Est. 225,000 USD

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…

Full-time
Coupang InternalSeattle, Washington, United States

Est. 165,000 USD

Please complete the attached the Internal Transfer Request Form and submit it.Please make sure you are applying with your Coupang e-mail address. Job Overview As the Director of Product Management for High Performance Co…

Full-time
NebiusRemote

Est. 140,000 USD

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to…

Full-timeRemote
Ursa MajorBerthoud, Colorado, United States

Est. 155,000 USD

The future of aerospace and defense starts here. Ursa Major was founded to revolutionize how America and its allies access and apply high-performance propulsion, from hypersonics to solid rocket motors, satellite maneuve…

Full-time

Est. 250,000 USD

We’re looking for a Principal Architect to lead the design and delivery of complex, multi-domain systems spanning cloud, data, and AI. This role is ideal for a deeply experienced engineer who owns the hardest architectur…

Full-timeRemote
CoupangMountain View, California, United States

Est. 165,000 USD

Company Introduction We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did we ever live without Coupang?” Born out of an obsession to make shopping, eating, and livin…

Full-time
CoupangSeattle, Washington, United States

Est. 165,000 USD

Company Introduction We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did we ever live without Coupang?” Born out of an obsession to make shopping, eating, and livin…

Full-time

Est. 250,000 USD

We’re looking for a Principal Architect to lead the design and delivery of complex, multi-domain systems spanning cloud, data, and AI. This role is ideal for a deeply experienced engineer who owns the hardest architectur…

Full-timeRemote
GlanceBangalore, India

Glance AI is an AI commerce platform shaping the next wave of e-commerce with inspiration-led shopping, less about searching for what you want and more about discovering who you could be. Operating in 140 countries, Glan…

Full-time
CoupangMountain View, California, United States

Est. 175,000 USD

Company Introduction We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did I ever live without Coupang?” Born out of an obsession to make shopping, eating, and living…

Full-time
GlanceBangalore, India

Glance AI is an AI commerce platform shaping the next wave of e-commerce with inspiration-led shopping, less about searching for what you want and more about discovering who you could be. Operating in 140 countries, Glan…

Full-time
CongaBoston, Massachusetts, United States

Est. 140,000 USD

A career that’s the whole package! At Conga, we’ve built a community where our colleagues can thrive. Here you’ll find opportunities to innovate and support growth through individual and team development, all within an e…

Full-time
SHEINSan Diego, California, United States

Est. 175,000 USD

About SHEIN SHEIN is a global online fashion and lifestyle retailer, offering SHEIN branded apparel and products from a global network of vendors, all at affordable prices. Headquartered in Singapore, with more than 15,0…

Full-time

APPLICATIONS FROM OUTSIDE COLOMBIA WILL NOT BE CONSIDERED FOR THIS ROLERobots & Pencils is an applied AI engineering firm building the next frontier of business architecture. We design and ship AI co-workers that int…

Full-timeRemote
CoupangMountain View, California, United States

Est. 140,000 USD

We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did we ever live without Coupang?” Born out of an obsession to make shopping, eating, and living easier than ever, w…

Full-time
NebiusRemote

Est. 140,000 USD

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to…

Full-timeRemote
AxleRockville, Maryland, United States

Est. 165,000 USD

(ID: 2026-1572) Axle is a bioscience and information technology company that offers advancements in translational research, biomedical informatics, and data science applications to research centers and healthcare organiz…

Full-time