Montreal Realty Ownership Tracing Pipeline
Budget / SalaryC$1,500–3,000
TypeFreelance project
LocationRemote
Posted1 hour ago
Build a property-ownership data pipeline — Montreal real estate (contract, milestone-paid, remote)
I run a commercial real estate brokerage in Montreal. I'm building a private database that maps every multi-residential building in a neighbourhood to the people and companies that actually own it — including owners hidden behind numbered companies. You'll build the data pipeline and database from scratch. It has to be clean enough to grow for years.
What you'll build
Ingest Quebec open data (Montreal assessment roll CSV, Registre des entreprises bulk export) into Supabase / Postgres
Schema: Properties · Entities (people + companies) · Links · Sources — keyed on matricule and NEQ, not address strings
Monthly refresh that's idempotent — re-run it ten times, same result, no duplicates
Link engine: connect entities by shared director / shared address / same name+address. Confident matches auto-link; anything fuzzy goes to a human review queue, never auto-merged
Flag portfolio owners (3+ properties)
Push scored targets to Monday.com via API
Monthly diff report: new sales, ownership changes, new buildings
Phase 2 (separate quote): contact-enrichment waterfall across Apollo, RocketReach and Lusha APIs, with source + confidence logged per contact
Stack: Python, PostgreSQL/Supabase, REST APIs. You should be comfortable using AI coding tools (Cursor, Claude Code, Copilot) and using an LLM API to parse messy text fields when it makes sense.
Who I'm looking for
Software Engineering, Computer Science or AI student/recent grad — McGill or Concordia preferred, Montreal-based a plus (you'll understand what REQ and the rôle d'évaluation are)
You've shipped at least one data pipeline that ran unattended for months
You care about the boring parts: keys, dedup, provenance, tests, a runbook someone else can follow
You push back when a spec is wrong
How it works
Milestone 0 — paid test, $100: download the Montreal open-data roll, filter one borough to 5+ unit residential buildings, load it into Supabase with a clean schema. Two to three hours. This is real work, not a quiz.
Milestones 1–3: schema + ingestion → link engine + review queue → Monday push + monthly diff. Phase 1 budget: $2,000–2,500, paid per milestone on delivery.
Phase 2 quoted separately after Phase 1 ships.
NDA + IP assignment signed before Milestone 0.
To apply, answer one question: describe a data pipeline you built that ran on its own for 3+ months. What broke, and what did you do about it? Skip the cover letter.
Phase 2 — Contact enrichment bot (quoted separately, after Phase 1 ships)
Build a bot that finds the best way to reach each scored owner in the database and writes every result into the Contacts table.
What it does
Runs only on owners flagged as targets — never the whole database
Strict waterfall, in this order, stopping when a verified mobile or direct line is found:
Apollo.io API — corporate email + phone on company-name and director-name matches
RocketReach API — personal mobile + email on director name + city
Lusha API — mobile fallback
Web search — company site, LinkedIn, Canada411, Google — only when the three APIs return nothing
Every hit becomes a new row in Contacts: entity_id · type · value · source · confidence score · date found. Nothing is overwritten
Cross-checks: if two sources return the same number, confidence goes up; if a number belongs to a different person with the same name, it's flagged, not stored as the owner's
Owners with no result after the full waterfall are tagged "field / direct mail" and pushed to Monday.com with the mailing address from the rôle
Logs API credit usage per run so cost per found contact is visible
Respects rate limits and each provider's terms
I run a commercial real estate brokerage in Montreal. I'm building a private database that maps every multi-residential building in a neighbourhood to the people and companies that actually own it — including owners hidden behind numbered companies. You'll build the data pipeline and database from scratch. It has to be clean enough to grow for years.
What you'll build
Ingest Quebec open data (Montreal assessment roll CSV, Registre des entreprises bulk export) into Supabase / Postgres
Schema: Properties · Entities (people + companies) · Links · Sources — keyed on matricule and NEQ, not address strings
Monthly refresh that's idempotent — re-run it ten times, same result, no duplicates
Link engine: connect entities by shared director / shared address / same name+address. Confident matches auto-link; anything fuzzy goes to a human review queue, never auto-merged
Flag portfolio owners (3+ properties)
Push scored targets to Monday.com via API
Monthly diff report: new sales, ownership changes, new buildings
Phase 2 (separate quote): contact-enrichment waterfall across Apollo, RocketReach and Lusha APIs, with source + confidence logged per contact
Stack: Python, PostgreSQL/Supabase, REST APIs. You should be comfortable using AI coding tools (Cursor, Claude Code, Copilot) and using an LLM API to parse messy text fields when it makes sense.
Who I'm looking for
Software Engineering, Computer Science or AI student/recent grad — McGill or Concordia preferred, Montreal-based a plus (you'll understand what REQ and the rôle d'évaluation are)
You've shipped at least one data pipeline that ran unattended for months
You care about the boring parts: keys, dedup, provenance, tests, a runbook someone else can follow
You push back when a spec is wrong
How it works
Milestone 0 — paid test, $100: download the Montreal open-data roll, filter one borough to 5+ unit residential buildings, load it into Supabase with a clean schema. Two to three hours. This is real work, not a quiz.
Milestones 1–3: schema + ingestion → link engine + review queue → Monday push + monthly diff. Phase 1 budget: $2,000–2,500, paid per milestone on delivery.
Phase 2 quoted separately after Phase 1 ships.
NDA + IP assignment signed before Milestone 0.
To apply, answer one question: describe a data pipeline you built that ran on its own for 3+ months. What broke, and what did you do about it? Skip the cover letter.
Phase 2 — Contact enrichment bot (quoted separately, after Phase 1 ships)
Build a bot that finds the best way to reach each scored owner in the database and writes every result into the Contacts table.
What it does
Runs only on owners flagged as targets — never the whole database
Strict waterfall, in this order, stopping when a verified mobile or direct line is found:
Apollo.io API — corporate email + phone on company-name and director-name matches
RocketReach API — personal mobile + email on director name + city
Lusha API — mobile fallback
Web search — company site, LinkedIn, Canada411, Google — only when the three APIs return nothing
Every hit becomes a new row in Contacts: entity_id · type · value · source · confidence score · date found. Nothing is overwritten
Cross-checks: if two sources return the same number, confidence goes up; if a number belongs to a different person with the same name, it's flagged, not stored as the owner's
Owners with no result after the full waterfall are tagged "field / direct mail" and pushed to Monday.com with the mailing address from the rôle
Logs API credit usage per run so cost per found contact is visible
Respects rate limits and each provider's terms
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.