NLP Anonymisation R&D Lead

via Freelancer ·

Budget / Salary$10,000–25,000
TypeFreelance project
LocationRemote
Posted1 hour ago
Project Title

Senior Privacy Research Engineer — NLP Anonymisation & De-Identification (R&D Contract)

Category / Skills tags

Python, Natural Language Processing, Machine Learning, Data Privacy, Algorithm Development, Artificial Intelligence, Research, Named Entity Recognition (NER)

Project Type

Full-time or heavy part-time contractor engagement. Open to individuals, teams, or agencies — please make clear in your bid who would actually be doing the work and how continuity is maintained if more than one person is involved. Bidders should propose their own rate and estimated timeline based on the scope below.

Description

We're building a privacy-preserving enterprise AI governance platform (UK patent-pending) and are hiring for the research-critical slice of the build — not integration work, but genuinely novel algorithm development that sits at the core of the platform's patentable value.
You'll work directly with the Client's technical lead, who holds final technical authority and personally reviews every deliverable. All communication, before and after award, takes place through the platform.

What you'll be working on:

- A text scrubbing/obfuscation engine that transforms sensitive text into safe representations, across several distinct operating modes
- A multi-pass de-identification protocol (package construction → surrogate evaluation → final acceptance/rehydration), designed so re-identification risk is measurably and provably reduced
- Local inference architecture for a privacy-isolated processing tier
- Optimisation work for surrogate model nodes
- An adversarial red-team test harness to stress-test the above

This is real R&D, not a fixed-scope integration job. Several components have already required a full re-estimate once it became clear they were novel algorithms rather than integration tasks — so please size your bid based on the scope description, not an assumed hour count. Deliverables are structured as acceptance-tested “packets,” each with explicit pass/fail criteria agreed up front, so timelines and pricing should reflect genuine research uncertainty rather than a fixed ticket-based estimate.

Work packages (scope, not hours):
- Research + deep study — architecture annexes, prior art review, threat modeling
- Scrubbing/Obfuscation Engine — 5 distinct operating modes
- De-Identification Protocol, Pass 1 — package construction
- De-Identification Protocol, Pass 2 — surrogate evaluation (expected to be the hardest research segment)
- De-Identification Protocol, Pass 3 — acceptance, lock, rehydration, evidence
- Local inference — three-tier architecture for a privacy-isolated processing tier
- Surrogate optimisation — three sub-nodes
- Red-team harness — adversarial testing for the above components

Required skills — please only bid if you have genuine, demonstrable experience in most of the following:

- Hands-on NLP/text-processing work — named entity recognition, tokenisation, anonymisation/pseudonymisation, or closely adjacent fields (not just general ML or LLM-application experience)
- Proven ability to read and implement from dense technical specifications and formal acceptance criteria, rather than working from loosely-defined tickets
- Research maturity — comfortable reporting “this didn't work and here's why” rather than forcing a result to look complete
- Strong, production-grade Python — including working close to model internals/behaviour, not solely calling third-party APIs
- A track record of turning ambiguous, open-ended research problems into shippable, tested code (please reference specific examples in your proposal)
- Prior direct exposure to privacy-preserving systems, PII detection, or de-identification techniques is a strong plus, though not a hard requirement

Working model:
- Async-first; UK/EU working-hours overlap is helpful for review cycles but not mandatory
- Direct reporting to and review by the Client's technical lead
- AI coding assistance is provided as an acceleration tool but doesn't replace your judgement on research-critical passes
- IP: all work delivered under this engagement is owned outright by the Client on completion and payment — not licensed. Please do not propose ongoing licence fees or retained-ownership structures.

Process:
There is no fixed bidding deadline — we will review bids on a rolling basis as they arrive, and may respond, ask clarifying questions, or close the listing at any point once we have a clear picture of bid quality. Bidders who clearly demonstrate the required skills above may be invited to a short take-home screening task, followed by a technical discussion.

The budget shown is indicative only. We will consider proposals above this range where the technical approach and experience justify it — price will be judged on the strength of the proposal relative to the scope, not simply ranked by amount. Please quote what you genuinely believe the work requires, not a figure anchored to the displayed range.

To apply:
Please include a brief note on directly relevant NLP/privacy-engineering experience, who would actually be doing the work, your estimate of total effort/timeline based on the scope above, and your proposed price (fixed-price or per-work-package preferred).
python research machine learning (ml) data science artificial intelligence data extraction data analysis deep learning data protection natural language processing
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.