V

LLM Engineering Expert – AI Evaluation

vraify · Anywhere

Full-timeLeadPython

🔥8 people viewed this job

About the Role

About the role Vraify is hiring experienced LLM Engineering Experts to create and validate challenging, simulation-based engineering design problems for evaluating advanced AI agents. This remote contractor opportunity is open only to candidates residing in the United States or Canada. You will design multi-constraint tasks, configure open-source simulation tools, analyze agent execution logs, diagnose reasoning and tool-use failures, and build objective automated graders across electrical, mechanical, aerospace, control systems, systems engineering, and robotics. Important experience note This is a senior specialist role requiring at least 8 years of directly relevant engineering experience. To help candidates focus on opportunities aligned with their background, applications that do not meet this minimum will be automatically screened out. Freshers and early-career applicants are therefore not eligible for this role, and we warmly encourage them to consider opportunities better suited to their current experience level. Responsibilities Author original, self-contained engineering design tasks with competing constraints, explicit optimization targets, validated reference solutions, and objective autograders. Build, run, and validate problem environments using open-source simulation tools and custom Python test benches. Evaluate coding-agent outputs and execution logs across repeated trials to identify systemic failure modes. Refine task difficulty using empirical model-performance data without introducing ambiguity or missing information. Collaborate with AI researchers, pod leads, and domain experts to integrate rigorous benchmarks into the model-evaluation pipeline. Required qualifications Master's degree or PhD in Electrical Engineering, Mechanical Engineering, Aerospace Engineering, or a closely related field. 8+ years of hands-on engineering design experience. Proficiency with at least one relevant open-source simulation package, such as ngspice, PySpice, OpenFOAM, FEniCSx, CalculiX, python-control, CadQuery, build123d, OpenModelica, Cantera, or Gmsh. Strong Python scripting skills. Hands-on experience with modern LLMs or coding agents and evaluation concepts such as pass@k, failure-mode analysis, and nondeterministic behavior. Ability to audit trajectory logs and isolate core reasoning or tool-use failures. Strong attention to physical plausibility, unit consistency, boundary conditions, convergence criteria, and technical documentation. Stable high-speed internet and a personal desktop or laptop suitable for remote work. Engagement details Remote contractor engagement for up to 24 weeks, starting immediately. Expected commitment: 40 hours per week with at least 4 hours of overlap with Pacific Time. Weekend on-call availability required; part-time engagement may be considered. Compensation: USD $500 per completed task; each task is expected to require approximately 30–40 hours. Shortlisted candidates will complete a calibration and delivery review process. Eligible locations United States and Canada only. Work Location: Remote

💬 Developer Questions

Ask the team a question — answers show up here

🎯

What does the interview process look like?

🤖

What AI/vibe coding tools does the team use daily?

👥

How big is the engineering team?

⏰

Is the team fully async or are there required meetings?

🚀

What does onboarding look like for remote hires?

🔧

Can you share more about the tech stack and architecture?

📈

What does career growth look like in this role?

📅

What does a typical day look like?

💰

Is there a salary range you can share?

📊

Is equity or stock options part of the package?

🌍

Are there timezone requirements or preferences?

🛂

Do you sponsor work visas?

🏢 Is this your listing? Claim it to answer questions

Similar Jobs

Helpful resources

Hiring for a similar role? Post your job here — it's free →