C

Senior AI Infrastructure Engineer, LLM/AI Platforms

CrowdStrike Holdings, Inc. · Anywhere

Full-timeLeadPythonAWSGCPDockerKubernetesPyTorchLangChain

🔥10 people viewed this job

About the Role

About the Role: CrowdStrike is looking for a Senior AI Infrastructure Engineer with expertise in Large Language Models (LLMs) Infrastructure and data platforms to join our growing AI Infrastructure Team. You will be a key leader, helping to design, build, and deploy cutting-edge AI infrastructure that powers our next generation of AI-driven security products. This role requires hands-on experience in LLM infrastructure to support multiple large scale training pipelines and scalable AI-powered systems. You will champion engineering best practices, write high-quality code, and actively mentor and strengthen the team's technical knowledge and capabilities. CrowdStrike is a computer security company, but we do not require candidates for this role to have prior security industry experience. We will mentor and train in security topics as needed. We do expect a strong interest in CrowdStrike's mission and a willingness to engage with the needs of our product teams. The scale of our systems and data are approaching Exabytes in size. Experience with extremely large-scale systems, including DevSecOps patterns, practices, and standards are important for this work. What You'll Do: Provision and configure large GPU clusters and compute resources for LLM training, finetuning, and inference workloads.Develop and optimize LLM model-serving infrastructure, including deployment and optimization of various inference frameworks.Lead model lifecycle management including versioning, checkpointing and reproducibility across training and inference deployments.Design and champion robust evaluation frameworks to assess model performance, accuracy, and reliability, ensuring AI systems are consistently at production-ready standards.Identify and address GPU utilization and GPU memory efficiency bottlenecks and apply techniques like quantization, batching, and caching.Architect and maintain data platforms and pipelines specifically designed to support LLMs, Retrieval-Augmented Generation (RAG), and AI Agentic Systems at scale.Deliver production-ready code with a focus on performance, maintainability, and testing rigor, ensuring the ability to ship fast without compromising quality.Apply expertise in data modeling, normalization, and semantic cataloging for AI/ML workloads.Define and enforce best practices for MLOps/DataOps surrounding LLMs, including monitoring, observability, and zero-touch recovery mechanisms for AI services.Document architectural designs thoroughly and communicate technical decisions clearly to stakeholdersCollaborate across the organization with Data Scientists, Product Managers, and other engineering teams to transform research prototypes into robust, production-grade services. Tech Stack (Experience in several areas is expected): Hands-on experience with MLOps Tools (MLflow, Sagemaker, Vertex AI).Strong understanding of CUDA, NVIDIA drivers, GPU, and TPU compute fundamentals.Experience with inference serving frameworks such as vLLM and Triton Inference Server.Proficiency with distributed training frameworks including Pytorch, Ray, Megatron, and JAX.Expert-level proficiency in a high-level coding language (Python).Deep knowledge of containerization and orchestration (Docker, Kubernetes, Slurm, Airflow).Proficiency with Infrastructure as Code tooling like Terraform and Ansible.Experience with cloud platforms (AWS, GCP, or OCI) and related data services. What You'll Need: Bachelor's degree in Computer Science, Data Engineering, or a related STEM field; Master's degree preferred6+ years of experience in Infrastructure/Data Engineering, with at least 2 years focused on building and maintaining platforms/pipelines that support LLM-based systems and applicationsDemonstrable hands-on experience in LLM infrastructure engineering including cluster provisioning, optimizing training workloads, and maintaining inference pipelinesExceptional ability to write clean, elegant, performant, and well-tested code, coupled with a strong focus on action and delivering results quickly.Thorough understanding of engineering practices including effective peer code reviews and resilient architecture designDemonstrates technical leadership and mentorship capabilitiesProven experience utilizing AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency and drive business outcomes.Bonus Points: Prior experience in the cybersecurity, intelligence, or high-compliance industries.Direct experience building, deploying, and managing LLMs in a production environment.Experience with common agentic workflow frameworks (e.g., LangChain, LlamaIndex).Experience with distributed data processing frameworks (e.g., Spark, Dask, Flink). #LI-RC1 #LI-Remote Benefits of Working at CrowdStrike: Market leader in compensation and equity awardsComprehensive physical and mental wellness programsCompetitive vacation and holidays for rechargePaid parental and adoption leavesProfessional development opportunities for all empl

CrowdStrike Holdings, Inc. has 1 open position on Remote Vibe Coding Jobs.

💬 Developer Questions

Ask the team a question — answers show up here

🎯

What does the interview process look like?

🤖

What AI/vibe coding tools does the team use daily?

👥

How big is the engineering team?

Is the team fully async or are there required meetings?

🚀

What does onboarding look like for remote hires?

🔧

Can you share more about the tech stack and architecture?

📈

What does career growth look like in this role?

📅

What does a typical day look like?

💰

Is there a salary range you can share?

📊

Is equity or stock options part of the package?

🌍

Are there timezone requirements or preferences?

🛂

Do you sponsor work visas?

🏢 Is this your listing? Claim it to answer questions

Similar Jobs

Helpful resources

Hiring for a similar role? Post your job here — it's free →