Description
About the Role
Together AI runs one of the largest GPU fleets in the world. The Infra Agent Systems team builds the software systems that power and automate that infrastructure.
We develop production AI agents that diagnose hardware failures, investigate incidents, correlate signals across the fleet, and automate operational workflows. Alongside these agents, we build the platform they run on, including knowledge graphs, retrieval systems, orchestration frameworks, and developer tooling.
You’ll work across two areas:
Infrastructure Agent Systems — Build production AI agents that help operate our GPU fleet by diagnosing failures, investigating incidents, gathering evidence from live systems, and assisting with remediation. These agents are used every day by our infrastructure and datacenter teams through APIs, CLI, dashboards, and Slack.
Core Agent Platform — Build the platform that powers these agents, including knowledge graphs, search and retrieval, orchestration, evaluation, and the tooling that enables agents to reason, act, and continuously improve.
We’re working on something that hasn’t really been done before: building knowledge graphs and self-improving AI agents that understand, operate, and continuously improve large-scale AI infrastructure.
This is an opportunity to work at the intersection of AI agents, distributed systems, infrastructure, and automation, solving challenging engineering problems with real production impact. There’s an enormous amount to build, learn, and shape as we define the future of autonomous infrastructure.
responsible for delivering the software but also for operating and supporting it in production.
Why this Role
You’ll work on two hard problems at the same time: making AI agents trustworthy enough to operate production infrastructure, and building the knowledge, retrieval, and distributed systems that make those agents effective.
You’ll have the opportunity to build foundational systems from the ground up, work on infrastructure at massive scale, and help define how self-improving AI agents operate real-world AI infrastructure.
Remote based in India
Responsibilities
- Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world’s largest GPU fleets.
- Build the distributed services, orchestration framework, knowledge graph, and retrieval systems that power infrastructure agents.
- Develop fleet intelligence systems that combine telemetry, infrastructure state, operational knowledge, and historical incidents to help agents make better decisions.
- Integrate with observability, incident management, ticketing, fleet inventory, source control, chat, and internal infrastructure systems through well-designed APIs.
- Own services end to end, including architecture, implementation, testing, deployment, observability, and production operations.
- Improve agent performance through evaluations, retrieval improvements, better tools, and production feedback loops.
- Turn what agents learn in production into reliable, reviewed software and automation.
Requirements
- 5+ years of experience building production backend systems, distributed systems, or infrastructure platforms.
- Strong systems design skills and experience owning significant systems from design through production.
- Depth in at least one of the following:
- AI agent systems, orchestration, tool use, evaluation, or grounding
- Knowledge graphs or graph data modeling
- Search, retrieval, ranking, RAG, or semantic search systems
- Strong backend engineering experience, including API design, service boundaries, data modeling, and integrations across complex systems.
- Experience with Kubernetes, GitOps such as ArgoCD, infrastructure-as-code, and cloud platforms.
- Comfortable working across languages such as Go, TypeScript, Python, or Rust.
Experience in the following is a plus:
- GPU infrastructure, datacenters, bare-metal systems, hardware failure modes, BMC/IPMI, or cluster schedulers
- Graph databases
- Event-driven systems and messaging platforms such as NATS or Kafka
- Observability platforms such as Prometheus and Grafana
- Building evaluation frameworks or improving the quality and reliability of LLM-powered systems
About Together AI
Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers and engineers in our journey in building the next generation AI infrastructure.
Equal Opportunity
Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
Similar jobs
About the Role We're looking for a Software Engineer to build the systems that treat infrastructure as software. This role owns the software state machines that provision hardware, bring it into service, and manage its f…
Est. 120,000 EUR
AI Infrastructure Systems Engineer Hybrid at our office in Amsterdam or Remote in the UK. Build the infrastructure powering the next generation of AI. At Together AI, you’ll build and operate one of the world’s largest G…
AI Infrastructure Systems Engineer Build the infrastructure powering the next generation of AI. At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inferenc…
Est. 255,000 USD
About the Role Together AI is building the AI Native Cloud, an end-to-end platform for the fullgenerative AI lifecycle, combining the fastest LLM inference engine with state-of-the-artAI cloud infrastructure. The Togethe…
Est. 120,000 GBP
About the Role We're looking for a Software Engineer to build the systems that treat infrastructure as software. This role owns the software state machines that provision hardware, bring it into service, and manage its f…
Est. 230,000 USD
Build the infrastructure powering the next generation of AI. At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrast…
Est. 90,000 EUR
About the Role We're looking for a Software Engineer to build the systems that treat infrastructure as software. This role owns the software state machines that provision hardware, bring it into service, and manage its f…
Est. 120,000 EUR
About the Role Together AI is building the AI Acceleration Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure. As…
Est. 195,000 USD
About the Role Together AI is building the AI Acceleration Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure. As…
Est. 260,000 USD
About the Role We're looking for a Software Engineer to build the systems that treat infrastructure as software. This role owns the software state machines that provision hardware, bring it into service, and manage its f…
About the Role As a Customer Support Engineer at a pioneering AI company, you'll be the first line of defense to support customers as they build out training, fine tuning, and inference solutions with Together AI. You'll…
About the role As a Technical Support Engineer at a pioneering AI company, you'll be the first line of defense to support customers as they build out training, fine tuning, and inference solutions with Together AI. You'l…
Est. 195,000 USD
About the role As a Technical Support Engineer at a pioneering AI company, you'll be the first line of defense to support customers as they build out training, fine tuning, and inference solutions with Together AI. You'l…
Est. 195,000 USD
About the role As a Technical Support Engineer at a pioneering AI company, you'll be the first line of defense to support customers as they build out training, fine tuning, and inference solutions with Together AI. You'l…
Est. 275,000 USD
About the Role In this role, you will operate, scale, and optimize multi-petabyte storage systems purpose-built for the world’s largest AI training and inference workloads. You’ll manage and scale high-performance parall…
Est. 155,000 USD
About The Role Together AI is rapidly scaling its compute infrastructure across multiple sites and deployment types. The Associate, Infrastructure Strategy & Operations will be the analytical backbone of the Infrastr…
Est. 220,000 USD
About the Role As a Solutions Architect at Together AI, you will work with customers and prospects to create business value through Generative AI applications. Solutions Architects at Together are trusted advisors to our…
Est. 205,000 USD
About the Role Together AI is building the Inference Platform that brings the most advanced generative AI models to the world. Our platform powers multi-tenant serverless workloads and dedicated endpoints, enabling devel…
Est. 275,000 USD
About the role As a Customer Success Engineer at Together AI, you will serve as the named technical owner for one of our most strategic customer relationships. You will be the primary technical point of contact across al…
Est. 120,000 EUR
About the Role Together AI is building the AI Acceleration Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure. As…
Est. 240,000 USD
About the Role Together AI is building the AI Acceleration Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure. The…
Est. 197,500 USD
About the Role Our product surface is expanding fast - GPU clusters, managed storage, networking, and observability - and we're adding a Product Manager to the Together Cloud team to own the day-to-day product work that…
Est. 240,000 USD
About the Role The Turbo team sits at the intersection of efficient inference (algorithms, architectures, engines) and post‑training / RL systems. We build and operate the systems behind Together’s API, including high‑pe…
Est. 230,000 USD
About The Role Together AI is building its infrastructure footprint at scale, and this role is central to making that happen. As an Infrastructure Design Engineer, you will own the design, planning, and technical executi…
Est. 215,000 USD
About The Role Together AI is scaling its own data center infrastructure to power the next generation of AI workloads. We are looking for a Program Manager, Data Center Delivery to serve as our representative across our…
Est. 275,000 USD
About the Role Together AI is scaling its physical AI infrastructure rapidly — and we're looking for a Director of Data Center Operations to help us build it right. This is a ground-floor opportunity to own the operation…
Est. 240,000 USD
About the Role This is a research engineering role with direct production impact. You won’t be publishing ideas in isolation—you will translate new RL algorithms, scheduling methods, and inference optimizations into prod…
Est. 235,000 USD
About the Role Together AI is looking for a Senior Network Engineer to design, deploy, and operate the global network infrastructure supporting our production services and high-performance AI compute environments. This i…
Est. 120,000 GBP
About the Role As a Solutions Architect (Inference) at Together AI, you will work with customers and prospects to create business value through Generative AI applications. Solutions Architects at Together are trusted adv…
Est. 260,000 USD
About the Role Together AI is hiring a Staff Platform Engineer to join the Product Foundations engineering organization and drive its service infrastructure strategy. Product Foundations builds and operates Together’s mi…