Z

LLM DevOps/Inference Engineer (W2 Only)

Zuplon · Anywhere

Full-timePythonAWSOpenAI

About the Role

Role: LLM DevOps/Inference Engineer (W2 Only) Location: REMOTE Must Have: Hands on exp launching inference containers Build and maintain the AWS infrastructure for the evaluation platform as code, including networking, compute, orchestration, secrets, observability, and CI/CD. Strong AI & LLM Recent healthcare industry exp (HIPAA, hl7, etc) Strong communication Context: We are looking for a DevOps/Inference Engineer for one of our clients building a healthcare-focused AI benchmark and evaluation suite. The initial target is for clinical prediction tasks including sepsis onset, days-to-death, and lab value trend forecasting, evaluated across multiple frontier and vertical-specific models. Role Summary You will be responsible for the infrastructure the benchmark harness runs on, the selfhosted model serving stack, and the reliability of the platform. This is a role for someone who can provision a GPU cluster in the morning and tune the inferencing model engine in the afternoon. What You will Own * Build and maintain the AWS infrastructure for the evaluation platform as code, including networking, compute, orchestration, secrets, observability, and CI/CD. * Stand up the self-hosted inference track for the long tail of vertical healthcare models. This involves provisioning infrastructure for models serving on GPU compute with sensible batching, quantization where appropriate, autoscaling, and a standard onboarding path so adding new models takes hours, not weeks. * Build the provider abstraction layer alongside the AI engineers so that APIbased models (OpenAI, Anthropic, Gemini, and the growing list beyond) and self-hosted models present a uniform interface to the harness. Rate limiting, retry and backoff, quota management, request/response logging, and cost attribution per run are your responsibility. * Make benchmark runs reproducible and cost-optimized with pinned model and container versions, captured configuration, spot and reserved capacity strategy, and idle GPU elimination. * Build the observability story with throughput, latency, token and GPU-hour cost, failure taxonomy, and per-model dashboards for monitoring. * Support the surge model that the platform must let a burst of AI engineers land, run experiments, and leave without breaking anything or leaving orphaned resources behind. * Contribute to Trusted Execution Environment (TEE) architecture. Evaluate AWS Nitro Enclaves and comparable confidential computing approaches for the bring-your-own-data / bring-your-own-model scenario, including attestation- gated key release and the practical constraints of running model inference inside an enclave. Required Skills * AWS infrastructure at production scale: EKS or ECS, EC2 GPU instance families (G5/G6, P4d/P5) and their capacity realities, VPC design, IAM, KMS, Secrets Manager, ECR, CloudWatch, and Service Quotas. * Infrastructure as code: Terraform. No console-clicked production resources. * Model serving and inference optimization: Hands-on experience working with LLMs. Practical command of batching strategy, KV cache behavior, quantization tradeoffs, and multi-GPU sharding. * Container orchestration and GPU scheduling: ECS/EKS with GPU workloads, node autoscaling, and image build pipelines for CUDA-dependent stacks. * Reliability and cost engineering. SLOs, alerting, and a demonstrated track record of optimizing cloud spend without cutting capability. Desirable Skills * AWS SageMaker endpoints and Bedrock. * Hands-on experience with Python to contribute directly to the harness and the provider adapter layer. * Confidential computing fundamentals: Enclaves, remote attestation, sealed key release, and the security boundaries of TEEs. * Healthcare compliance posture: HIPAA-eligible service selection, BAA scope, audit logging, and the access-control mechanisms for PHI data. * Security hardening, including image scanning and network egress control for a closed-loop environment. Nice to Have * Prior experience hosting medical imaging or multimodal models. * Nitro Enclaves in production, or comparable TEE work. * Experience supporting self-service environments, clean tenancy boundaries, and fast credential provisioning.

Zuplon has 1 open position on Remote Vibe Coding Jobs.

💬 Developer Questions

Ask the team a question — answers show up here

🎯

What does the interview process look like?

🤖

What AI/vibe coding tools does the team use daily?

👥

How big is the engineering team?

Is the team fully async or are there required meetings?

🚀

What does onboarding look like for remote hires?

🔧

Can you share more about the tech stack and architecture?

📈

What does career growth look like in this role?

📅

What does a typical day look like?

💰

Is there a salary range you can share?

📊

Is equity or stock options part of the package?

🌍

Are there timezone requirements or preferences?

🛂

Do you sponsor work visas?

🏢 Is this your listing? Claim it to answer questions

Similar Jobs

Helpful resources

Hiring for a similar role? Post your job here — it's free →