Urgent LLM Benchmark Testing
Budget / Salary₹1,500–12,500
TypeFreelance project
LocationRemote
Posted1 hour ago
We are conducting a research project on benchmarking LLM-powered/autonomous penetration-testing tools and urgently need an experienced cybersecurity/AI-security expert to assist us.
Deadline: Wednesday, August 12, 2026
This is not a conventional VAPT project. We need someone who understands both offensive security and LLM/AI-agent evaluation.
What we are benchmarking
We are evaluating applicable AI cybersecurity/pentesting tools and agents, including tools such as:
PentestGPT
VulnBot
HackingBuddyGPT
Cochise
AutoPentest
AutoAttacker
RapidPlan
BridgeSeek
PenHeal
AutoSec Agent
XBOW
CyBench
ZeroDayBench
CVE-related benchmarks
Other relevant autonomous pentesting tools
Controlled targets
The primary web targets include:
bWAPP
OWASP Juice Shop
DVWA
OWASP WebGoat
What we need from the expert
We need help with:
Designing/validating a fair and reproducible benchmarking methodology
Understanding the correct configuration of each tool
Determining its full supported execution capacity
Running/evaluating applicable tools against controlled targets
Establishing vulnerability ground truth
Validating reported vulnerabilities
Measuring true positives, false positives and false negatives
Measuring vulnerability coverage and exploitation success
Evaluating evidence and generated reports
Recording execution time, failures, crashes and hallucinations
Ensuring the results are honest, original and scientifically defensible
We do not want arbitrary limits or modifications to the tools. The goal is to evaluate each system in its intended configuration and at its supported full capacity.
For tools supporting local models, Ornith 1 9B may be used in a separate controlled experiment. Tools requiring their own supported LLM/provider will be evaluated accordingly.
Ideal candidate
Strong preference for someone with experience in:
Penetration testing / VAPT
Web application security
OWASP
AI/LLM security
Autonomous penetration testing
Cybersecurity research
LLM-agent evaluation
Python/Linux/Kali
CTFs
Experience with PentestGPT, XBOW, CyBench, ZeroDayBench, VulnBot or similar autonomous cyber agents is highly desirable.
Urgency
We need to start immediately and have the required benchmarking guidance/work completed or substantially progressed by Wednesday, August 12, 2026.
Please apply only if you can start immediately.
In your proposal, please answer:
How many years of offensive-security experience do you have?
Have you worked with LLM-powered/autonomous pentesting agents?
Which tools from the above list have you personally used?
Have you designed or evaluated a cybersecurity benchmark before?
How would you establish ground truth and distinguish AI hallucinations from real vulnerabilities?
Can you start immediately?
What is your hourly/project rate?
How much work can you complete before Wednesday?
We are looking for an expert who can contribute immediately, not a generic VAPT tester.
Deadline: Wednesday, August 12, 2026
This is not a conventional VAPT project. We need someone who understands both offensive security and LLM/AI-agent evaluation.
What we are benchmarking
We are evaluating applicable AI cybersecurity/pentesting tools and agents, including tools such as:
PentestGPT
VulnBot
HackingBuddyGPT
Cochise
AutoPentest
AutoAttacker
RapidPlan
BridgeSeek
PenHeal
AutoSec Agent
XBOW
CyBench
ZeroDayBench
CVE-related benchmarks
Other relevant autonomous pentesting tools
Controlled targets
The primary web targets include:
bWAPP
OWASP Juice Shop
DVWA
OWASP WebGoat
What we need from the expert
We need help with:
Designing/validating a fair and reproducible benchmarking methodology
Understanding the correct configuration of each tool
Determining its full supported execution capacity
Running/evaluating applicable tools against controlled targets
Establishing vulnerability ground truth
Validating reported vulnerabilities
Measuring true positives, false positives and false negatives
Measuring vulnerability coverage and exploitation success
Evaluating evidence and generated reports
Recording execution time, failures, crashes and hallucinations
Ensuring the results are honest, original and scientifically defensible
We do not want arbitrary limits or modifications to the tools. The goal is to evaluate each system in its intended configuration and at its supported full capacity.
For tools supporting local models, Ornith 1 9B may be used in a separate controlled experiment. Tools requiring their own supported LLM/provider will be evaluated accordingly.
Ideal candidate
Strong preference for someone with experience in:
Penetration testing / VAPT
Web application security
OWASP
AI/LLM security
Autonomous penetration testing
Cybersecurity research
LLM-agent evaluation
Python/Linux/Kali
CTFs
Experience with PentestGPT, XBOW, CyBench, ZeroDayBench, VulnBot or similar autonomous cyber agents is highly desirable.
Urgency
We need to start immediately and have the required benchmarking guidance/work completed or substantially progressed by Wednesday, August 12, 2026.
Please apply only if you can start immediately.
In your proposal, please answer:
How many years of offensive-security experience do you have?
Have you worked with LLM-powered/autonomous pentesting agents?
Which tools from the above list have you personally used?
Have you designed or evaluated a cybersecurity benchmark before?
How would you establish ground truth and distinguish AI hallucinations from real vulnerabilities?
Can you start immediately?
What is your hourly/project rate?
How much work can you complete before Wednesday?
We are looking for an expert who can contribute immediately, not a generic VAPT tester.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.