About the Role
We are sharing a specialised part-time consulting opportunity for experienced machine learning engineers with hands-on experience using AI coding agents and building production ML systems, model deployment infrastructure, LLM applications, or AI-powered products.
This sprint-based role supports an advanced AI research initiative focused on evaluating frontier coding models through realistic machine learning engineering workflows. Selected professionals will use AI coding agents to complete technical tasks, review model-generated implementations, identify bugs and failure modes, and compare how different models perform across practical ML engineering scenarios.
Key Responsibilities
Machine Learning Engineering Evaluation
Review complex machine learning and AI engineering tasks completed with frontier coding agentsEvaluate implementations involving model training, inference systems, MLOps, and LLM applicationsAssess technical correctness, architecture choices, implementation quality, and engineering trade-offsApply professional ML engineering judgment to realistic production-oriented scenarios
AI Coding Agent Testing
Use AI coding agents as part of hands-on technical workflowsEvaluate how effectively coding models interpret requirements and implement solutionsIdentify bugs, incomplete implementations, edge cases, and unexpected behaviourAssess where models require additional prompting, correction, or manual engineering intervention
Technical Quality & Failure Analysis
Identify performance issues, reliability problems, and model failure modesReview generated code for maintainability, correctness, and practical usabilityEvaluate whether implementations would function appropriately in realistic ML environmentsDocument technical strengths, weaknesses, and important implementation risks
Model Comparison & Technical Judgment
Compare outputs produced by multiple frontier coding modelsAssess differences in implementation strategy, code quality, technical reasoning, and reliabilityDetermine which approaches best satisfy task requirementsProvide clear written assessments explaining relevant engineering trade-offs
Ideal Profile
Strong candidates may have:
At least 2 years of professional machine learning engineering experienceExperience building production ML systems, AI-powered applications, or model-serving infrastructureHands-on experience with model training, inference, deployment, or MLOpsExperience developing LLM applications or integrating foundation models into production systemsRegular use of AI coding agents within software or machine learning development workflowsStrong ability to evaluate model-generated code and technical implementation decisionsExcellent debugging, analytical reasoning, and written communication skillsAbility to work efficiently within short, intensive project sprints
Educational Background
A degree in computer science, machine learning, artificial intelligence, software engineering, or a related technical discipline may be helpfulAdvanced study in machine learning or computer science may strengthen an applicationEquivalent professional experience building and deploying production ML systems may also be consideredPractical engineering depth is particularly important for this engagement
Nice to Have
Experience with Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or comparable AI coding toolsProduction experience deploying machine learning modelsFamiliarity with model-serving architectures and inference optimisationExperience building LLM-powered applications or agentic systemsKnowledge of MLOps, deployment pipelines, monitoring, or model infrastructureExperience evaluating generated code across multiple AI coding systemsPrevious exposure to AI evaluation, benchmark development, or structured technical review
Why This Opportunity
Work directly with frontier AI coding agents on realistic ML engineering problemsEvaluate advanced models across production-oriented machine learning workflowsApply practical engineering experience to identify subtle technical failure modesCompare multiple coding systems and help improve their reliabilityParticipate in intensive technical sprints with task-based compensation
Contract Details
Independent contractor roleFully remote with flexible schedulingSprint-based project with task windows typically spanning approximately 12–24 hoursCompensation is $400 per accepted taskTypical tasks require approximately 2–3 hours after ramp-upCompensation is tied to successfully accepted workWork may include ML implementation review, coding-agent evaluation, debugging, model comparison, and technical analysisWeekly payments via Stripe or WiseProjects may be extended, shortened, or adjusted depending on scope and performanceWork will not involve access to confidential or proprietary information from any employer, client, or institution
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remo