About the Role
About OpenTrainOpenTrain is the hiring and contracting organization for this role and the #1 platform for finding and building careers in AI training and data labeling. Contributors use OpenTrain to discover projects, build a credible AI training profile, and grow their technical experience into a lasting portfolio.
About AI Training WorkAI training is the human side of building modern artificial intelligence. Technical evaluators help improve AI systems by creating challenging tasks, reviewing model-generated code, checking solutions against objective criteria, and documenting why implementations succeed or fail.
This work offers a direct way to contribute to cutting-edge AI development while working remotely and choosing a flexible schedule. OpenTrain brings together opportunities across the industry so specialists can build experience through meaningful projects.
The RoleOpenTrain is seeking a Machine Learning Engineering Evaluator to create, solve, review, and validate demanding machine-learning engineering tasks. The work covers model development, training and inference systems, numerical computing, performance optimization, Python workflows, and technical evaluation.
You will assess implementations for correctness, reproducibility, efficiency, and performance while documenting technical decisions, trade-offs, limitations, and failure modes. This is a global, fully remote contractor role requiring approximately 15 hours per week, with flexible scheduling including weekends.
The role is listed as entry level, but it requires substantial professional or research experience in machine learning and advanced technical capability.
Work type: Part-time contractorSchedule: Approximately 15 hours per week with flexible schedulingWork arrangement: Fully remoteCompensation: Listed at $100-$150 per hourLanguage: EnglishEligibility: Applicants in the countries listed for this roleWhat You'll DoYou will work across the full technical lifecycle of machine-learning systems, from implementation and data preparation through evaluation, optimization, and verification. You will also review AI-generated code and technical solutions for quality and practical correctness.
Develop and validate machine-learning models, training pipelines, inference systems, and supporting infrastructure.Implement model components, data pipelines, evaluation systems, and numerical methods.Build reproducible workflows using Python and command-line tools.Work with tensor operations, automatic differentiation, model architectures, tokenization, batching, and generation.Optimize latency, throughput, memory usage, and hardware utilization.Diagnose numerical instability, tensor errors, memory bottlenecks, distributed-system failures, and performance regressions.Review AI-generated code and technical solutions.Design objective tests, benchmarks, and verification criteria.Explain implementation choices, performance trade-offs, limitations, and failure modes clearly.RequirementsYou should have advanced machine-learning knowledge and the ability to debug systems beyond surface-level API usage. Equivalent tools, substantial open-source work, or academic experience may qualify where relevant.
Master's degree or PhD in computer science, machine learning, artificial intelligence, applied mathematics, statistics, engineering, or a closely related quantitative discipline.Strong professional or research experience in machine learning.Practical proficiency with Python and reproducible technical workflows.Meaningful experience with at least two relevant machine-learning frameworks, libraries, or inference tools.Strong understanding of model training, evaluation, numerical computation, or inference.Ability to debug ML systems and explain implementation decisions, performance trade-offs, and failure modes.Relevant tools may include PyTorch, JAX, NumPy, SciPy, SGLang, vLLM, llama.cpp, Hugging Face Transformers, and Hugging Face Tokenizers.Who Should ApplyThis opportunity is suited to machine-learning practitioners, researchers, and engineers who enjoy turning technical requirements into rigorous tests and reliable implementations. It may be a strong fit if you can investigate system behavior deeply, identify performance or correctness issues, and communicate technical conclusions precisely.
The work is flexible and remote, making it possible to contribute part time while building a visible portfolio of AI training and evaluation experience through OpenTrain.
Machine-learning engineers who understand training and inference systems.Researchers with experience implementing and validating models or numerical methods.Python developers comfortable with reproducible experiments and command-line workflows.Technical reviewers who can evaluate AI-generated code objectively.Engineers experienced with model performance, memory use, hardware utilization, or distributed systems.How It WorksCreate a free OpenTrain account, build your technical profile, and apply in minut