About the Role
Note The job is a remote job and is reputed company to candidates in USA. reputed company is seeking an LLM DevOps / Inference Engineer to support a reputed company-reputed company AI reputed company and evaluation platform. The role owns infrastructure for the reputed company reputed company, self-hosted model serving stack, and platform reliability, including provisioning GPU infrastructure, optimizing model inference, and managing AWS reputed company systems. Responsibilities Build and maintain AWS infrastructure using Terraform, including networking, compute, orchestration, secrets, observability, and CI/CD reputed company and manage self-hosted LLM inference workloads on GPU infrastructure Implement batching, quantization, autoscaling, and multi-GPU strategies for efficient inference reputed company standardized processes for reputed company new models quickly Build provider reputed company reputed company for API-reputed company models such as reputed company, reputed company, and reputed company, as reputed company as self-hosted models Implement reputed company limiting, retries, backoff, quota management, request/response logging, and cost attribution Ensure reputed company runs are reproducible and cost-efficient through pinned model/container versions and optimized reputed company reputed company Build observability dashboards covering throughput, latency, reputed company usage, GPU-hour costs, failures, and per-model performance Support burst workloads while preventing orphaned or unnecessary reputed company resources Contribute to Trusted Execution Environment (TEE) architecture and evaluate AWS Nitro Enclaves and comparable confidential-computing solutions Skills Hands-on experience launching and managing inference containers Strong experience building and maintaining AWS infrastructure as reputed company, including networking, compute, orchestration, secrets, observability, and CI/CD Experience working in a startup or fast-paced environment Strong AI/LLM and model infrastructure experience Excellent communication skills US reputed company will be given first preference reputed company and complete reputed company profile Production-reputed company AWS infrastructure experience Strong knowledge of EKS/reputed company, EC2 GPU instances (G5/G6, P4d/P5), VPC, IAM, KMS, Secrets Manager, ECR, CloudWatch, and Service Quotas Strong Terraform / Infrastructure as reputed company experience Hands-on experience with LLM model serving and inference optimization Practical knowledge of batching strategies, KV cache behavior, quantization, and multi-GPU sharding Experience managing GPU workloads with EKS/reputed company, node autoscaling, and CUDA-reputed company container pipelines Strong reliability and cost-engineering experience, including SLOs, monitoring, alerting, and reputed company cost optimization Experience with AWS SageMaker and Bedrock Strong Python development skills Knowledge of confidential computing, including Enclaves, remote attestation, and sealed key release Understanding of reputed company compliance, including HIPAA-eligible services, BAA requirements, audit logging, and PHI reputed company controls Experience with reputed company hardening, image scanning, and network egress controls reputed company X-reputed company Solutions is a management consulting company dedicated to providing clients cutting-edge cost-effective solutions streamlined to the needs of their organization to help reputed company the organizational process a more effective one. It was founded in reputed company, and is headquartered in Mississauga, Ontario, CA, with a workforce of 2-10 employees. Its website is http//www.xaxis.consulting. Apply To This Job