About the Role
Note: The job is a remote job and is open to candidates in USA. Together AI is a research-driven artificial intelligence company focused on optimizing AI systems. They are seeking an Inference Frameworks and Optimization Engineer to design and optimize distributed inference engines for large language models, ensuring high performance and scalability.
Responsibilities
Design and develop fault-tolerant, high-concurrency distributed inference engine for text, image, and multimodal generation modelsImplement and optimize distributed inference strategies, including Mixture of Experts (MoE) parallelism, tensor parallelism, pipeline parallelism for high-performance servingApply CUDA graph optimizations, TensorRT/TRT-LLM graph optimizations, and PyTorch-based compilation (torch.compile), and speculative decoding to enhance efficiency and scalabilityCollaborate with hardware teams on performance bottleneck analysis, co-optimize inference performance for GPUs, TPUs, or custom acceleratorsWork closely with AI researchers and infrastructure engineers to develop efficient model execution plans and optimize E2E model serving pipelines
Skills
3+ years of experience in deep learning inference frameworks, distributed systems, or high-performance computingFamiliar with at least one LLM inference frameworks (e.g., TensorRT-LLM, vLLM, SGLang, TGI(Text Generation Inference))Background knowledge and experience in at least one of the following: GPU programming (CUDA/Triton/TensorRT), compiler, model quantization, and GPU cluster schedulingDeep understanding of KV cache systems like Mooncake, PagedAttention, or custom in-house variantsProficient in Python and C++/CUDA for high-performance deep learning inferenceDeep understanding of Transformer architectures and LLM/VLM/Diffusion model optimizationKnowledge of inference optimization, such as workload scheduling, CUDA graph, compiled, efficient kernelsStrong analytical problem-solving skills with a performance-driven mindsetExcellent collaboration and communication skills across teamsExperience in developing software systems for large-scale data center networks with RDMA/RoCEFamiliar with distributed filesystem(e.g., 3FS, HDFS, Ceph)Familiar with open source distributed scheduling/orchestration frameworks, such as Kubernetes (K8S)Contributions to open-source deep learning inference projects
Benefits
Startup equityHealth insuranceOther competitive benefits
Company Overview
Together AI provides a cloud platform for developing, training, fine-tuning, and deploying generative AI models. It was founded in 2022, and is headquartered in San Francisco, California, USA, with a workforce of 201-500 employees. Its website is https://www.together.ai.
Company H1B Sponsorship
Together AI has a track record of offering H1B sponsorships, with 14 in 2026, 19 in 2025, 6 in 2024, 3 in 2023. Please note that this does not guarantee sponsorship for this specific role.