Applied Process-Mining

via Freelancer ·

Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted3 hours ago
We are seeking an applied process-mining research engineer to implement a time-bounded, preregistered comparative analysis of a historical loan-application event log.

This is not a conventional dashboarding or process-optimization consulting assignment. The primary objective is to build a reproducible analytical pipeline that applies three frozen diagnostic selection methods to the same discovery data, evaluates their resulting decisions on a concealed chronological holdout, and clearly distinguishes what the event log establishes from what remains unknown or merely hypothesized.

The work will support a practitioner research presentation in early October. The selected contractor must be able to execute a written protocol faithfully, identify ambiguities before analysis, and resist changing analytical rules after seeing favorable or unfavorable results.

Responsibilities

You will:

Inspect and validate the a Loan Application event log and its lifecycle semantics.
Construct a canonical, case-level analytical representation from the raw event data.
Document mappings among cases, applications, offers, activities, resources, lifecycle events and timestamps.
Implement preregistered discovery and chronological holdout partitions.
Calculate specified transition intervals, waiting measures, case durations, loops, rework, handoffs and variant features.
Distinguish inter-event gaps from supported measures of active work or organizational waiting.
Implement the frozen wait-first and automation-first comparison methods.
Implement the preregistered calculations required by the Flow-First arm without silently adding analyst discretion.
Support outcome-linked candidate analysis using frozen exposure definitions, adjusted models, uncertainty estimates and materiality rules.
Run specified robustness and sensitivity tests, including business-calendar and lifecycle-definition checks.
Preserve separation between discovery and holdout evidence.
Produce audit-ready tables and exhibits for technical review and presentation.
Maintain automated data-quality and calculation tests.
Document analytical limitations and identify claims that the data cannot support.

The methodology owner will perform the interpretive Flow-First diagnosis. The contractor will support its computational execution but will not silently redefine its target, mechanism or selection rule.

Required qualifications
Strong Python skills, particularly pandas or Polars and reproducible analytical workflows.
Demonstrated experience with event logs, process mining, transaction histories or longitudinal operational data.
Familiarity with XES and process-mining libraries such as PM4Py.
Experience reconstructing process lifecycles and validating timestamp semantics.
Ability to calculate case-, event- and interval-level measures without conflating them.
Working knowledge of regression or time-to-event methods, bootstrapping and uncertainty estimation.
Experience with chronological validation, holdouts or other leakage-sensitive evaluation designs.
Proficiency with Git, automated tests and documented analytical environments.
Ability to explain the difference among descriptive evidence, observational association, causal inference and simulated effects.
Strong written documentation and attention to preregistered rules.

Helpful but not required
Experience with financial-services application or case-management workflows.
Experience evaluating competing analytical or diagnostic methods.
Familiarity with Celonis, Disco, UiPath Process Mining or comparable platforms.
Experience preparing reproducibility packages or technical appendices for research.
Familiarity with calendar-aware duration calculations and process-variant analysis.

Commercial process-mining platforms may be used for exploration or visual validation, but the authoritative analysis must be reproducible through documented code without requiring a proprietary license.

Deliverables
Validated analytical dataset
documented entity and event mappings;
lifecycle reconstruction;
discovery and holdout assignments;
data-quality findings and exclusions;
case-, interval- and candidate-level analytical tables.
Reproducible analysis pipeline
version-controlled source code;
environment and dependency specification;
deterministic configuration;
automated data-quality and calculation tests;
one-command or clearly documented rerun procedure.
Frozen-method implementation
wait-first calculation and selection output;
automation-first calculation and selection output;
Flow-First computational measures defined by the protocol;
audit trail showing the input, calculation and selected candidate for each method.
Holdout and robustness results
frozen holdout evaluation;
specified robustness and sensitivity tests;
uncertainty estimates;
explicit treatment of calendar effects, lifecycle ambiguity, missingness and segment instability.
Technical findings package
presentation-ready tables and charts;
process and variant views where analytically useful;
concise technical appendix;
limitations and claim-boundary register;
reproducibility handoff documentation.

Optional exploratory interface

A compact dashboard or notebook permitting inspection of individual cases, variants and key distributions may be included if it does not displace the authoritative analytical deliverables.

Acceptance criteria

The work will be accepted when:

all frozen analytical rules are implemented without undocumented discretion;
discovery and holdout data remain appropriately separated;
every reported number can be traced to source events and executable code;
the same inputs and configuration reproduce the same outputs;
automated tests cover critical mappings, durations, partitions and selection calculations;
exclusions, transformations and missing-data treatments are documented;
observed associations are not presented as proven causal effects;
simulations, if included, expose their assumptions and uncertainty and are not presented as observed gains;
another qualified analyst can rerun and review the work without the contractor’s intervention;
presentation exhibits agree with the authoritative computational outputs.

Expected engagement
Approximately 15–20 hours per week.
Roughly 90–120 total hours, subject to final scope.
Start as soon as possible.
Core analytical work completed by late September to allow independent review before the early-October presentation.
Short scheduled reviews at protocol clarification, pipeline validation, pre-holdout freeze and final-results stages.

Application materials

Please provide:

one relevant example involving event-sequence or operational-process data;
a short description of how you validated event and lifecycle semantics;
your experience with Python, PM4Py or comparable tools;
your approach to preventing holdout leakage and undocumented analytical discretion;
your availability, hourly rate and anticipated hours through late September.
python git data visualization data analysis pandas regression analysis
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.