O

LLM Evaluation and Repository Validation Engineer

OpenTrain AI · Remote

🔥15 people viewed this job

About the Role

About OpenTrainOpenTrain is the hiring and contracting organization for this role and the #1 platform for finding and building careers in AI training and data labeling. We help contributors discover specialized projects, build a credible AI training profile, and grow their experience in a rapidly developing industry. About AI Training and Software EvaluationAI training is the human side of building modern artificial intelligence. For software-focused projects, experienced engineers create and review coding tasks, test model-generated changes, and assess whether an AI system can understand, debug, and improve real code. Your technical judgment will help evaluate large language models against realistic repository histories and bug-fixing challenges. This work puts software engineers close to the development of cutting-edge AI systems while offering flexible, remote project work. The RoleOpenTrain is seeking an LLM Evaluation and Repository Validation Engineer to build verifiable software engineering tasks from public repository histories. You will analyze GitHub issues, configure complex codebases, establish development environments, assess test quality, and run projects locally to evaluate LLM bug-fixing performance. The assignment is a remote contractor engagement scheduled for 20 hours per week over three months. The role is listed as entry level, while the responsibilities call for strong hands-on software engineering judgment and experience working with complex repositories. Contractor assignmentPart-time schedule of 20 hours per weekThree-month engagementRemote and worldwideEnglish-language workWhat You'll DoYou will work with open-source software repositories and collaborate with researchers to select meaningful evaluation examples. The work combines repository setup, software testing, debugging, issue analysis, and technical assessment of AI-generated solutions. Analyze and triage GitHub issues across open-source libraries.Set up and configure repositories, including Dockerization and development environment setup.Evaluate unit test coverage and overall code quality.Run, modify, and test real-world projects locally.Assess LLM performance on bug-fixing and software engineering tasks.Identify repositories and issues that are especially challenging for LLMs.Collaborate with researchers on repository selection and task design.Apply senior software engineering judgment to technical evaluations.Required Skills and ExperienceYou should be comfortable navigating complex codebases, debugging locally, and setting up software projects from public repositories. Strong experience with at least one supported programming language is required, along with practical knowledge of Git, Docker, and basic software pipeline setup. Strong experience with at least one of Python, JavaScript, Java, Go, Rust, C, C++, C#, or Ruby.Proficiency with Git, Docker, and basic software pipeline setup.Ability to navigate complex codebases and debug locally.Ability to run, modify, and test projects on a local machine.Experience evaluating unit tests, debugging code, and assessing bug-fixing performance.Strong experience working with public GitHub repositories and complex codebases.Helpful BackgroundThe following experience is preferred or helpful for this assignment. It can support effective collaboration with researchers and stronger evaluation of challenging software tasks. Experience contributing to or evaluating open-source projects.Previous work in LLM research or evaluation.Experience with developer tools and automation agents.Tech lead-level software engineering experience.Familiarity with high-quality public GitHub repositories, including widely used repositories with 500 or more stars.Comfort collaborating with researchers on AI evaluation tasks.Remote collaboration experience with overlapping Pacific Time hours.Why Build Your AI Training Career with OpenTrainAI training and data-labeling work is expanding beyond image and text annotation into specialized software evaluation. Contributors can use their technical expertise to shape how AI systems code, reason about repositories, and respond to real engineering problems. OpenTrain gives freelancers a place to build a durable record of AI training experience. A stronger profile can help you demonstrate relevant work, find projects aligned with your skills, and develop a long-term career in AI training. Work remotely with a flexible part-time schedule.Apply software engineering expertise to cutting-edge AI evaluation.Build experience in LLM testing and developer automation.Develop a portfolio of specialized AI training work through OpenTrain.How to ApplyCreate a free OpenTrain account, build your profile around your software engineering experience, and apply in minutes. Highlight your programming language experience, Git and Docker proficiency, repository work, debugging background, and any LLM evaluation or open-source contributions.

OpenTrain AI has 1 open position on Remote Vibe Coding Jobs.

💬 Developer Questions

Ask the team a question — answers show up here

🎯

What does the interview process look like?

🤖

What AI/vibe coding tools does the team use daily?

👥

How big is the engineering team?

Is the team fully async or are there required meetings?

🚀

What does onboarding look like for remote hires?

🔧

Can you share more about the tech stack and architecture?

📈

What does career growth look like in this role?

📅

What does a typical day look like?

💰

Is there a salary range you can share?

📊

Is equity or stock options part of the package?

🌍

Are there timezone requirements or preferences?

🛂

Do you sponsor work visas?

🏢 Is this your listing? Claim it to answer questions

Similar Jobs

Helpful resources

Hiring for a similar role? Post your job here — it's free →