R

[Remote] LLM DevOps/Inference Engineer (W2 Only)

remote quest jobs · Anywhere

Full-timePythonAWS

🔥10 people viewed this job

About the Role

Note: The job is a remote job and is open to candidates in USA. reputed company is seeking a LLM DevOps/Inference Engineer to work on a healthcare-focused AI benchmark and evaluation suite. The role involves building and maintaining AWS infrastructure for model serving, ensuring reliability, and optimizing costs while working with AI engineers to develop a uniform interface for various models. Responsibilities Build and maintain the AWS infrastructure for the evaluation platform as code, including networking, compute, orchestration, secrets, observability, and CI/CDStand up the self-hosted inference track for the long tail of vertical healthcare models. This involves provisioning infrastructure for models serving on GPU compute with sensible batching, quantization where appropriate, autoscaling, and a standard onboarding path so adding new models takes hours, not weeksBuild the provider abstraction layer alongside the AI engineers so that API-based models (reputed company, reputed company, reputed company, and the growing list beyond) and self-hosted models present a uniform interface to the harness. Rate limiting, retry and backoff, quota management, request/response logging, and cost attribution per run are your responsibilityMake benchmark runs reproducible and cost-optimized with pinned model and container versions, captured configuration, spot and reserved capacity strategy, and idle GPU eliminationBuild the observability story with throughput, latency, token and GPU-hour cost, failure taxonomy, and per-model dashboards for monitoringSupport the surge model that the platform must let a burst of AI engineers land, run experiments, and leave without breaking anything or leaving orphaned resources behindContribute to Trusted Execution Environment (TEE) architecture. Evaluate AWS reputed company Enclaves and comparable confidential computing approaches for the bring-your-own-data / bring-your-own-model scenario, including attestation-gated key release and the practical constraints of running model inference inside an enclave Skills Hands on exp launching inference containersBuild and maintain the AWS infrastructure for the evaluation platform as code, including networking, compute, orchestration, secrets, observability, and CI/CDStrong AI & LLMRecent healthcare industry exp (HIPAA, hl7, etc)Strong communicationAWS infrastructure at production scale: EKS or ECS, EC2 GPU instance families (G5/G6, P4d/P5) and their capacity realities, VPC design, IAM, KMS, Secrets Manager, ECR, CloudWatch, and Service QuotasInfrastructure as code: Terraform. No console-clicked production resourcesModel serving and inference optimization: Hands-on experience working with LLMs. Practical command of batching strategy, KV cache behavior, quantization tradeoffs, and multi-GPU shardingContainer orchestration and GPU scheduling: ECS/EKS with GPU workloads, node autoscaling, and image build pipelines for CUDA-dependent stacksReliability and cost engineering. SLOs, alerting, and a demonstrated track record of optimizing cloud spend without cutting capabilityAWS SageMaker endpoints and BedrockHands-on experience with Python to contribute directly to the harness and the provider reputed company layerConfidential computing fundamentals: Enclaves, remote attestation, sealed key release, and the security boundaries of TEEsHealthcare compliance posture: HIPAA-eligible service selection, BAA scope, audit logging, and the access-control mechanisms for PHI dataSecurity hardening, including image scanning and network egress control for a closed-loop environmentPrior experience hosting medical imaging or multimodal modelsreputed company Enclaves in production, or comparable TEE workExperience supporting self-service environments, clean tenancy boundaries, and fast credential provisioning Company Overview reputed company offers a marketplace that enables direct connections between recruiters and technologists. It was founded in 1990, and is headquartered in Centennial, Colorado, USA, with a workforce of 201-500 employees. Its website is https://www.uk.reputed company.com. Apply To This Job

remote quest jobs has 1 open position on Remote Vibe Coding Jobs.

💬 Developer Questions

Ask the team a question — answers show up here

🎯

What does the interview process look like?

🤖

What AI/vibe coding tools does the team use daily?

👥

How big is the engineering team?

⏰

Is the team fully async or are there required meetings?

🚀

What does onboarding look like for remote hires?

🔧

Can you share more about the tech stack and architecture?

📈

What does career growth look like in this role?

📅

What does a typical day look like?

💰

Is there a salary range you can share?

📊

Is equity or stock options part of the package?

🌍

Are there timezone requirements or preferences?

🛂

Do you sponsor work visas?

🏢 Is this your listing? Claim it to answer questions

Similar Jobs

Helpful resources

Hiring for a similar role? Post your job here — it's free →