AI Evaluator & QA Specialist Needed
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted2 hours ago
We are seeking experienced AI Evaluators, QA Specialists, and Technical Pod Leads to evaluate and benchmark frontier LLMs and agentic AI systems for client pilot projects at JudgeMyAI.
Key Responsibilities:
• Evaluate AI model outputs across multi-turn reasoning, agentic tool-use, and code generation (SWE-bench).
• Stress-test models via adversarial prompts, edge-case analysis, and red-teaming.
• Apply and calibrate evaluation rubrics (RLHF, SFT, accuracy, safety, and hallucination checks).
• Review and annotate dataset batches with high inter-annotator agreement.
Requirements:
• Experience with LLM evaluation, prompt engineering, or AI benchmarking (prior work on platforms like Turing, Outlier, micro1 or enterprise AI labs is a strong plus).
• Background in Python, Software Engineering, Data Science, or Advanced STEM.
• Strong analytical reasoning and strict attention to detail.
Engagement Details:
• Remote and flexible hours (asynchronous execution).
• Fixed milestone or hourly payouts in USD.
Key Responsibilities:
• Evaluate AI model outputs across multi-turn reasoning, agentic tool-use, and code generation (SWE-bench).
• Stress-test models via adversarial prompts, edge-case analysis, and red-teaming.
• Apply and calibrate evaluation rubrics (RLHF, SFT, accuracy, safety, and hallucination checks).
• Review and annotate dataset batches with high inter-annotator agreement.
Requirements:
• Experience with LLM evaluation, prompt engineering, or AI benchmarking (prior work on platforms like Turing, Outlier, micro1 or enterprise AI labs is a strong plus).
• Background in Python, Software Engineering, Data Science, or Advanced STEM.
• Strong analytical reasoning and strict attention to detail.
Engagement Details:
• Remote and flexible hours (asynchronous execution).
• Fixed milestone or hourly payouts in USD.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.