About the Role
Machine Learning Engineers, MLOps & LLM Engineers
We are seeking experienced Machine Learning Engineers, AI Engineers, MLOps Engineers, LLM Engineers, and Applied AI Developers to evaluate and improve frontier AI coding agents through realistic technical tasks.
You will use advanced coding agents to solve and review machine learning engineering workflows involving model training, inference, deployment, MLOps, LLM applications, and production AI systems.
What You'll Do
Use frontier AI coding agents to complete complex machine learning engineering tasks
Review AI-generated implementations for correctness, scalability, reliability, and performance
Evaluate model training, inference, deployment, and production ML workflows
Identify bugs, edge cases, architectural weaknesses, performance bottlenecks, and failure modes
Compare outputs from multiple AI coding models
Assess technical tradeoffs and determine which implementation is stronger
Debug AI-generated Python, ML, data, and infrastructure code
Apply real-world engineering judgment to production-style AI and ML scenarios
Provide structured technical evaluations and clear written reasoning
Test whether generated solutions actually work in realistic environments
Who Can Apply
Relevant backgrounds include:
Machine Learning Engineers, Senior Machine Learning Engineers, ML Engineers, AI Engineers, Artificial Intelligence Engineers, Applied AI Engineers, LLM Engineers, Generative AI Engineers, Deep Learning Engineers, Applied Machine Learning Engineers, and AI Software Engineers.
We also welcome:
MLOps Engineers, ML Platform Engineers, Machine Learning Infrastructure Engineers, AI Infrastructure Engineers, Model Deployment Engineers, Model Serving Engineers, ML Systems Engineers, ML Reliability Engineers, ML Production Engineers, AI Platform Engineers, and Model Operations Engineers.
LLM and GenAI backgrounds may include:
LLM Application Engineers, LLMOps Engineers, AI Agent Engineers, Agentic AI Engineers, RAG Engineers, Prompt Engineers with strong coding experience, AI Product Engineers, AI Backend Engineers, Conversational AI Engineers, NLP Engineers, and Foundation Model Engineers.
Additional relevant roles include:
Data Scientists, Applied Scientists, Research Engineers, Research Scientists, Computer Vision Engineers, NLP Engineers, Speech ML Engineers, Recommendation Engineers, Ranking Engineers, Search Engineers, Data Engineers, Backend Engineers, Software Engineers, Platform Engineers, and Distributed Systems Engineers with strong production machine learning experience.
Relevant Machine Learning Experience
Experience with one or more of the following is valuable:
Model training and fine-tuning
Deep learning
Supervised and unsupervised learning
Transformer architectures
Large language models
Retrieval-augmented generation
AI agents and tool use
Model inference and serving
Batch and real-time prediction systems
Feature engineering and feature stores
Model evaluation and benchmarking
Experiment tracking
Hyperparameter optimization
Data preprocessing and training pipelines
Distributed training
GPU-based workloads
Model monitoring and observability
Model versioning
Production ML pipelines
ML API development
Scalability and latency optimization
Failure analysis and debugging
AI safety and model evaluation
AI Coding Agent Experience
Regular use of AI coding tools is strongly preferred, including:
Cursor, Claude Code, Codex, Windsurf, Gemini CLI, GitHub Copilot, Cline, Roo Code, Aider, Replit, or similar AI coding agents.
Requirements
2+ years of professional machine learning engineering or closely related experience
Hands-on experience building real ML, AI, or data-driven software systems
Experience with model training, production inference, ML infrastructure, LLM applications, or AI-powered products
Strong Python programming skills
Ability to understand and debug unfamiliar machine learning codebases
Familiarity with modern AI coding agents
Ability to evaluate AI-generated implementations and technical tradeoffs
Strong understanding of software engineering fundamentals
Strong debugging, analytical, and technical reasoning skills
Clear written communication and attention to detail
Preferred Background
Experience deploying machine learning systems to production
Experience operating high-scale or latency-sensitive inference systems
Experience with MLOps and ML platform infrastructure
Experience building LLM, RAG, or AI agent applications
Experience with distributed training or GPU workloads
Experience reviewing code written by other ML engineers
Experience designing ML benchmarks or evaluation frameworks
Experience with model observability, monitoring, and production debugging
Prior work evaluating AI-generated code or frontier coding agents
Experience with research-to-production machine learning workflows
This opportunity is ideal for experienced ML engineers who already use AI coding agents heavily and can quickly determine whether an AI-genera