Frontier-Style Task Creator
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted1 hour ago
Software Engineering Task Author (Contract)
About the Role
We're looking for an experienced software engineer to design and build Frontier-style software engineering tasks — long-horizon, repository-based coding challenges used to train and evaluate AI coding agents. You'll take real engineering problems from your own area of expertise and turn them into rigorous, verifiable tasks: a seeded repository, a clear public specification, and an automated grader that checks observable behavior (builds, hidden test suites, protocol conformance, concurrency/recovery scenarios, deterministic replay, numeric budgets, etc.).
This is deep, research-grade work — closer to writing a hard systems/algorithms exam with an autograder than to typical QA or content work.
What You'll Do
Design non-trivial engineering problems grounded in real-world systems (e.g., concurrency bugs, recovery/consistency issues, legacy modernization, protocol implementation, performance-constrained rewrites).
Build a starter repository, write a complete public specification (instruction.md), and implement a reference solution that solves the task correctly.
Write automated grading code that measures exactly what the instructions ask for — no hidden requirements, no LLM judges, no shortcuts.
Produce a no-op/naive baseline to confirm the task has real difficulty, and adversarial "attack" submissions to confirm the grader can't be gamed.
Validate, iterate, and submit tasks via pull request through a structured review pipeline.
Use AI tools to accelerate parts of your workflow (explaining code, debugging validation errors, reviewing your own logic) — but the engineering judgment and task design must be your own.
What We're Looking For
Strong professional software engineering background (backend systems, distributed systems, compilers, low-level/systems programming, or similar) — ideally in an area where you have real depth and war stories, not just familiarity.
Comfort with concurrency, correctness, and testing — the ability to reason precisely about what a grader should and shouldn't check.
Experience writing clear technical specifications and/or authoring hard technical interview problems, coding challenges, or test suites.
Git/GitHub workflow fluency (branches, PRs, CI).
Self-directed: comfortable working from written guidelines, validating your own work, and iterating based on automated feedback and review.
Bonus: experience with reinforcement learning, model evaluation, or "reward hacking" adversarial thinking (i.e., trying to break your own grader before someone else does).
Nice to Have
A specific technical niche you know cold (e.g., database internals, network protocols, embedded systems, compilers/language runtimes, numerical computing) that would make for a compelling, hard-to-fake task.
Prior experience creating technical assessments, benchmarks, or eval datasets.
Engagement Details
Contract / freelance, task-based.
Deliverables are reviewed against defined structure, CI, and quality checks before acceptance.
Ongoing support and clarification provided via a shared Discord community.
To apply, please share examples of complex systems you've built or debugged, any prior experience authoring technical assessments or benchmarks, and a brief note on what kind of engineering problem you'd want to turn into a task.
About the Role
We're looking for an experienced software engineer to design and build Frontier-style software engineering tasks — long-horizon, repository-based coding challenges used to train and evaluate AI coding agents. You'll take real engineering problems from your own area of expertise and turn them into rigorous, verifiable tasks: a seeded repository, a clear public specification, and an automated grader that checks observable behavior (builds, hidden test suites, protocol conformance, concurrency/recovery scenarios, deterministic replay, numeric budgets, etc.).
This is deep, research-grade work — closer to writing a hard systems/algorithms exam with an autograder than to typical QA or content work.
What You'll Do
Design non-trivial engineering problems grounded in real-world systems (e.g., concurrency bugs, recovery/consistency issues, legacy modernization, protocol implementation, performance-constrained rewrites).
Build a starter repository, write a complete public specification (instruction.md), and implement a reference solution that solves the task correctly.
Write automated grading code that measures exactly what the instructions ask for — no hidden requirements, no LLM judges, no shortcuts.
Produce a no-op/naive baseline to confirm the task has real difficulty, and adversarial "attack" submissions to confirm the grader can't be gamed.
Validate, iterate, and submit tasks via pull request through a structured review pipeline.
Use AI tools to accelerate parts of your workflow (explaining code, debugging validation errors, reviewing your own logic) — but the engineering judgment and task design must be your own.
What We're Looking For
Strong professional software engineering background (backend systems, distributed systems, compilers, low-level/systems programming, or similar) — ideally in an area where you have real depth and war stories, not just familiarity.
Comfort with concurrency, correctness, and testing — the ability to reason precisely about what a grader should and shouldn't check.
Experience writing clear technical specifications and/or authoring hard technical interview problems, coding challenges, or test suites.
Git/GitHub workflow fluency (branches, PRs, CI).
Self-directed: comfortable working from written guidelines, validating your own work, and iterating based on automated feedback and review.
Bonus: experience with reinforcement learning, model evaluation, or "reward hacking" adversarial thinking (i.e., trying to break your own grader before someone else does).
Nice to Have
A specific technical niche you know cold (e.g., database internals, network protocols, embedded systems, compilers/language runtimes, numerical computing) that would make for a compelling, hard-to-fake task.
Prior experience creating technical assessments, benchmarks, or eval datasets.
Engagement Details
Contract / freelance, task-based.
Deliverables are reviewed against defined structure, CI, and quality checks before acceptance.
Ongoing support and clarification provided via a shared Discord community.
To apply, please share examples of complex systems you've built or debugged, any prior experience authoring technical assessments or benchmarks, and a brief note on what kind of engineering problem you'd want to turn into a task.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.