inhousefyi
← Back to listings

Performance Engineer

Astera LabsSan Jose, California, United States · Posted 1 month ago
Full-timeEst. 141,000 USD
Apply now

Description

Astera Labs (NASDAQ: ALAB) provides rack-scale AI infrastructure through purpose-built connectivity solutions. By collaborating with hyperscalers and ecosystem partners, Astera Labs enables organizations to unlock the full potential of modern AI. Astera Labs’ Intelligent Connectivity Platform integrates CXL®, Ethernet, NVLink, PCIe®, and UALink™ semiconductor-based technologies with the company’s COSMOS software suite to unify diverse components into cohesive, flexible systems that deliver end-to-end scale-up, and scale-out connectivity. The company’s custom connectivity solutions business complements its standards-based portfolio, enabling customers to deploy tailored architectures to meet their unique infrastructure requirements. Discover more at www.asteralabs.com.

Senior Performance Engineer

Location: San Jose, CA (On-site)

Role Overview

Astera Labs is a hyper-growth connectivity company enabling the rack-scale AI infrastructure powering the world's most advanced GPU clusters. Our Scorpio scale-up fabric switches are purpose-built to unlock the performance of next-generation AI workloads, and we're looking for a Senior Performance Engineer to demonstrate the real-world value of our silicon where it matters most: on real inference and training workloads running on GPUs at scale.

In this role, you will define how the world measures scale-up fabric performance. You'll build the roofline models, benchmarks, and end-to-end workload studies that quantify our performance leadership, expose bottlenecks, and drive performance fine-tuning of real AI workloads on our fabric to inform product direction. Your data will directly shape architecture, firmware, and product decisions — and fuel the marketing narrative that positions Astera Labs at the center of AI connectivity.

Key Responsibilities

  • Performance Characterization & Benchmarking
  • Establish theoretical and measured roofline models for Astera Labs' scale-up fabric across key performance metrics, defining the reference for all comparative testing.
  • Build and maintain baseline performance benchmarks using industry-standard tools such as NVBandwidth and NCCL across a range of GPU configurations and switch topologies.
  • Quantify the impact of differentiated Astera Labs AI fabric features (e.g., Hypercast, In-Network Computing) against baselines using both synthetic benchmarks and real inference workloads.
  • Real Workload Analysis & Fabric Scalability
  • Run end-to-end inference model workloads on target hardware to capture real-world performance beyond synthetic benchmarks, supporting architecture decisions and customer-facing demonstrations.
  • Evaluate fabric performance as inference cluster size scales from 16 to 32 GPUs and beyond, identifying bottlenecks and building performance scaling models for state-of-the-art AI workloads.
  • Design and execute head-to-head performance comparisons against competing fabric switch solutions to produce data-driven differentiation evidence.
  • Test Infrastructure & Automation
  • Design, build, and maintain automated lab infrastructure including test execution pipelines, traffic generation tooling, and data collection and reporting systems.
  • Enable repeatable, high-quality, and scalable performance measurements across all hardware configurations, reducing manual effort and accelerating the test cycle.
  • Share infrastructure and playbooks with the Product Applications team to accelerate customer application development and issue resolution.
  • Cross-Functional Impact & Innovation
  • Partner closely with ASIC architecture, firmware, software, Product Definition, Product Applications, and Product Marketing teams to communicate findings, influence design decisions, and resolve performance-impacting issues.
  • Serve as a key technical resource in the early evaluation of new fabric architectures, interconnect technologies (UALink, PCIe Gen 6/Gen 7, Ethernet, UEC), and AI/ML communication paradigms.
  • Provide performance data, analysis, and live benchmark support for key customer engagements and industry events; produce clear, audience-appropriate performance reports, technical briefs, and marketing collateral, and maintain living documentation in Confluence.

Basic Qualifications

  • Bachelor's degree in Computer Engineering, Computer Science, Electrical Engineering, or a related technical field. We welcome both recent graduates with strong, directly relevant project, research, or internship experience and candidates with 2–5 years of industry experience in performance or systems engineering.
  • Hands-on experience running AI/ML workloads on GPU clusters — including benchmarking, performance analysis, and fine-tuning of workloads across clusters of GPUs or accelerators. This can come from industry, research, or substantial academic projects.