T

[Remote] LLM Inference Frameworks and Optimization Engineer

Together AI · Anywhere

Full-timePythonKubernetesPyTorch

About the Role

Note: The job is a remote job and is open to candidates in USA. Together AI is a research-driven artificial intelligence company focused on optimizing AI systems. They are seeking an Inference Frameworks and Optimization Engineer to design and optimize distributed inference engines for large language models, ensuring high performance and scalability. Responsibilities Design and develop fault-tolerant, high-concurrency distributed inference engine for text, image, and multimodal generation modelsImplement and optimize distributed inference strategies, including Mixture of Experts (MoE) parallelism, tensor parallelism, pipeline parallelism for high-performance servingApply CUDA graph optimizations, TensorRT/TRT-LLM graph optimizations, and PyTorch-based compilation (torch.compile), and speculative decoding to enhance efficiency and scalabilityCollaborate with hardware teams on performance bottleneck analysis, co-optimize inference performance for GPUs, TPUs, or custom acceleratorsWork closely with AI researchers and infrastructure engineers to develop efficient model execution plans and optimize E2E model serving pipelines Skills 3+ years of experience in deep learning inference frameworks, distributed systems, or high-performance computingFamiliar with at least one LLM inference frameworks (e.g., TensorRT-LLM, vLLM, SGLang, TGI(Text Generation Inference))Background knowledge and experience in at least one of the following: GPU programming (CUDA/Triton/TensorRT), compiler, model quantization, and GPU cluster schedulingDeep understanding of KV cache systems like Mooncake, PagedAttention, or custom in-house variantsProficient in Python and C++/CUDA for high-performance deep learning inferenceDeep understanding of Transformer architectures and LLM/VLM/Diffusion model optimizationKnowledge of inference optimization, such as workload scheduling, CUDA graph, compiled, efficient kernelsStrong analytical problem-solving skills with a performance-driven mindsetExcellent collaboration and communication skills across teamsExperience in developing software systems for large-scale data center networks with RDMA/RoCEFamiliar with distributed filesystem(e.g., 3FS, HDFS, Ceph)Familiar with open source distributed scheduling/orchestration frameworks, such as Kubernetes (K8S)Contributions to open-source deep learning inference projects Benefits Startup equityHealth insuranceOther competitive benefits Company Overview Together AI provides a cloud platform for developing, training, fine-tuning, and deploying generative AI models. It was founded in 2022, and is headquartered in San Francisco, California, USA, with a workforce of 201-500 employees. Its website is https://www.together.ai. Company H1B Sponsorship Together AI has a track record of offering H1B sponsorships, with 14 in 2026, 19 in 2025, 6 in 2024, 3 in 2023. Please note that this does not guarantee sponsorship for this specific role.

💬 Developer Questions

Ask the team a question — answers show up here

🎯

What does the interview process look like?

🤖

What AI/vibe coding tools does the team use daily?

👥

How big is the engineering team?

Is the team fully async or are there required meetings?

🚀

What does onboarding look like for remote hires?

🔧

Can you share more about the tech stack and architecture?

📈

What does career growth look like in this role?

📅

What does a typical day look like?

💰

Is there a salary range you can share?

📊

Is equity or stock options part of the package?

🌍

Are there timezone requirements or preferences?

🛂

Do you sponsor work visas?

🏢 Is this your listing? Claim it to answer questions

Similar Jobs

Helpful resources

Hiring for a similar role? Post your job here — it's free →