About the Role
About OpenTrainOpenTrain AI is the hiring and contracting organization for this opportunity and the #1 platform for finding and building careers in AI training and data labeling. OpenTrain helps people discover projects, build their AI-work profile, and apply in minutes. Creating an account is free.
Fully remote contract workPart-time schedule of 20+ hours per weekOpen to candidates in Bangladesh, Georgia, India, Indonesia, Malaysia, Pakistan, the Philippines, Sri Lanka, Thailand, and VietnamAbout AI Training And Agent EvaluationAI training is the human side of building modern artificial intelligence. Engineers and contributors create the tools, environments, test cases, and evaluation systems that help researchers measure how reliably AI models and agents perform.
This fast-growing field supports cutting-edge systems through work such as coding evaluation, model-output review, scoring pipelines, and developer tooling. Your infrastructure work will help researchers run repeatable, secure experiments.
Contribute to infrastructure used in LLM training workflowsSupport evaluation of AI agent performanceWork remotely with flexible project participationThe RoleOpenTrain AI is seeking senior-minded Python engineers to design and maintain infrastructure for LLM training workflows and agent tooling. You will architect secure sandboxes and task environments, build reusable repositories and scoring pipelines, maintain CI/CD systems, and support researchers who depend on these tools.
The listing identifies the experience level as intermediate while requiring at least five years of professional Python experience. The work is fully remote, available to eligible Asia-Low candidates, and conducted in English.
Role type: Part-time contractorTime requirement: 20+ hours per weekData type: Computer code programmingSubject matter: Python AI infrastructure, LLM tooling, and automationWhat You'll DoYou will create reliable, reusable infrastructure for researchers and developers working on AI agent evaluation. The role combines production Python engineering, environment automation, security hardening, testing, and hands-on technical support.
Design and maintain secure sandboxes and task environmentsBuild reusable repositories, scoring pipelines, and developer environmentsWrite clean, production-grade, test-driven Python codeCreate unit, integration, and functional tests with pytestContainerize services with Docker and debug multi-stage DockerfilesDesign CI/CD pipelines that lint, test, build, and deploySupport researchers who rely on the infrastructureExplain tools, write concise documentation, and pair-program when neededUse AI coding assistants safely, including tools such as Cursor, Claude Code, or CopilotRequired QualificationsApplicants should bring strong professional Python experience and the ability to own infrastructure that must be secure, testable, and dependable. The project specifically calls for at least five years of professional Python work and practical experience across development environments, services, containers, and automation.
5+ years of professional Python experience, including production code, packaging, async I/O, and legacy-module refactoringStrong testing mindset with pytest, coverage, and reliability practicesAdvanced Linux command-line skills using tools such as bash, grep, curl, jq, and systemdUnderstanding of basic networking, permissions, and Linux operationsHands-on Docker expertise, including multi-stage Dockerfiles and docker-composeExperience designing GitHub Actions or similar CI/CD workflowsKnowledge of CI/CD secrets management and cachingFastAPI or Flask proficiency for modular REST or asynchronous servicesExperience with authentication, validation using Pydantic, and service loggingAbility to create devcontainer.json files, Makefiles, .env workflows, and pre-commit hooksExposure to LLM or agent infrastructure, such as sandboxes, scoring pipelines, or evaluation frameworksSecurity awareness, including least-privilege design, image hardening, and CI scanners such as Trivy or SnykVersion-control discipline, including semantic commits, PR templates, and code-review practicesPreferred Experience And CollaborationThe strongest candidates will be comfortable supporting technical researchers and adopting modern engineering workflows. Experience training or evaluating AI systems, particularly on coding-focused tasks, is a bonus rather than a stated prerequisite.
Comfort mentoring teammates and explaining infrastructure clearlyAbility to write concise technical documentation and pair-programHands-on experience with Cursor, Claude Code, Copilot, or similar AI coding toolsPrior experience training or evaluating AI systems is a bonusCoding-focused AI work on platforms such as OpenTrain or Alignerr is a bonusCompensation And ScreeningThe structured listing shows a payment rate of $12.50 USD per hour. The project description also specifies experience-based hourly tiers of $9 for Junior, $12 for