LLM Pre-training & Distributed Engineer (AI Infrastructure)
Hyphen Connect LimitedHong Kong, Hong Kong · Posted 4 months agoDescription
We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The ideal candidate will have a deep understanding of GPU clusters and extensive experience in system engineering to ensure efficient and reliable training processes.
Responsibilities:
- Orchestrate distributed training runs across 1,000+ GPUs using PyTorch, DeepSpeed, or Megatron-LM.
- Optimize networking (InfiniBand/RDMA) and memory management to prevent out-of-memory errors.
- Automate checkpointing and failure recovery during month-long training runs.
Required Skills:
- Deep expertise in 3D parallelism (Data, Tensor, Pipeline).
- Experience managing SLURM or Kubernetes-based GPU clusters.
- Strong systems engineering background (C++, CUDA, Python).
Similar jobs
We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The idea…
We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The idea…
Est. 140,000 USD
We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The idea…
Est. 140,000 USD
We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The idea…
We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The idea…
Est. 141,000 USD
We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The idea…
Est. 141,000 USD
We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The idea…
We are seeking a skilled MLOps & Agentic Platform Engineer. This role involves managing model registries, developing continuous training loops, and implementing A/B testing infrastructure. The ideal candidate will ha…
Est. 250,000 USD
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify inn…
We are seeking a skilled MLOps & Agentic Platform Engineer. This role involves managing model registries, developing continuous training loops, and implementing A/B testing infrastructure. The ideal candidate will ha…
Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with…
We are seeking a skilled MLOps & Agentic Platform Engineer. This role involves managing model registries, developing continuous training loops, and implementing A/B testing infrastructure. The ideal candidate will ha…
We are hiring for one of our ecosystem projects in the financial services/ HFT space, we are looking for a Linux Trading System Engineer to build, improve, and maintain a complex Linux environment. This position is based…
Est. 237,500 USD
Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with…
Est. 315,000 USD
About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…
Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with…
Est. 200,000 GBP
Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with…
Est. 80,000 USD
We are looking for an AI Specialist Engineer to enhance the performance of large language and vision models for on-device inference. Your expertise will be crucial in developing and deploying cutting-edge AI solutions, e…
We are seeking a skilled MLOps & Agentic Platform Engineer. This role involves managing model registries, developing continuous training loops, and implementing A/B testing infrastructure. The ideal candidate will ha…
Est. 127,500 USD
Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with…
We are seeking a talented and innovative Synthetic Data Engineer. In this role, you will design and implement domain-specific synthetic data generation pipelines, ensuring high-quality data management for training loops.…
Position Title: Machine Learning Specialist (Research & Engineering) Work Location: BGC, Taguig City. (2 x onsite per week hybrid set up) We are seeking a versatile Machine Learning Specialist to own the end-to-end l…
Est. 85,000 GBP
Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with…
We are hiring for one of our ecosystem projects in the financial services/ HFT space, we are looking for a Linux Trading System Engineer to build, improve, and maintain a complex Linux environment. This position is based…
We are hiring for one of our ecosystem projects in the financial services/ HFT space, we are looking for a Linux Trading System Engineer to build, improve, and maintain a complex Linux environment. This position is based…
Est. 215,000 USD
Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with…
Est. 60,000 USD
We are looking for an AI Engineer to join a Shipping SaaS Platform. This isn’t just a "prompt engineering" role—you will work at the intersection of deep mathematics and production-grade engineering to build, optimize, a…
Est. 200,000 USD
Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with…
Est. 45,000 USD
We are looking for an AI Engineer to join a Shipping SaaS Platform. This isn’t just a "prompt engineering" role—you will work at the intersection of deep mathematics and production-grade engineering to build, optimize, a…
Est. 141,000 USD
We are seeking a skilled MLOps & Agentic Platform Engineer. This role involves managing model registries, developing continuous training loops, and implementing A/B testing infrastructure. The ideal candidate will ha…