Urgent LLM Benchmark Testing

via Freelancer ·

Budget / Salary₹1,500–12,500
TypeFreelance project
LocationRemote
Posted1 hour ago
We are conducting a research project on benchmarking LLM-powered/autonomous penetration-testing tools and urgently need an experienced cybersecurity/AI-security expert to assist us.

Deadline: Wednesday, August 12, 2026

This is not a conventional VAPT project. We need someone who understands both offensive security and LLM/AI-agent evaluation.

What we are benchmarking

We are evaluating applicable AI cybersecurity/pentesting tools and agents, including tools such as:

PentestGPT
VulnBot
HackingBuddyGPT
Cochise
AutoPentest
AutoAttacker
RapidPlan
BridgeSeek
PenHeal
AutoSec Agent
XBOW
CyBench
ZeroDayBench
CVE-related benchmarks
Other relevant autonomous pentesting tools
Controlled targets

The primary web targets include:

bWAPP
OWASP Juice Shop
DVWA
OWASP WebGoat
What we need from the expert

We need help with:

Designing/validating a fair and reproducible benchmarking methodology
Understanding the correct configuration of each tool
Determining its full supported execution capacity
Running/evaluating applicable tools against controlled targets
Establishing vulnerability ground truth
Validating reported vulnerabilities
Measuring true positives, false positives and false negatives
Measuring vulnerability coverage and exploitation success
Evaluating evidence and generated reports
Recording execution time, failures, crashes and hallucinations
Ensuring the results are honest, original and scientifically defensible

We do not want arbitrary limits or modifications to the tools. The goal is to evaluate each system in its intended configuration and at its supported full capacity.

For tools supporting local models, Ornith 1 9B may be used in a separate controlled experiment. Tools requiring their own supported LLM/provider will be evaluated accordingly.

Ideal candidate

Strong preference for someone with experience in:

Penetration testing / VAPT
Web application security
OWASP
AI/LLM security
Autonomous penetration testing
Cybersecurity research
LLM-agent evaluation
Python/Linux/Kali
CTFs

Experience with PentestGPT, XBOW, CyBench, ZeroDayBench, VulnBot or similar autonomous cyber agents is highly desirable.

Urgency

We need to start immediately and have the required benchmarking guidance/work completed or substantially progressed by Wednesday, August 12, 2026.

Please apply only if you can start immediately.

In your proposal, please answer:
How many years of offensive-security experience do you have?
Have you worked with LLM-powered/autonomous pentesting agents?
Which tools from the above list have you personally used?
Have you designed or evaluated a cybersecurity benchmark before?
How would you establish ground truth and distinguish AI hallucinations from real vulnerabilities?
Can you start immediately?
What is your hourly/project rate?
How much work can you complete before Wednesday?

We are looking for an expert who can contribute immediately, not a generic VAPT tester.
ai chatbot development ai model development ai development ai agents ai integration ai compliance ai ethics ai governance ai quality assurance ai red teaming
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.