Finalization, Refactoring and Validation of an Automation, Ranking and Data Analysis System with Machine Learning

via Freelancer ·

Budget / Salary€250–750
TypeFreelance project
LocationRemote
Posted1 hour ago
Finalization, Refactoring and Validation of an Automation, Ranking and Data Analysis System with Machine Learning

Project Summary
This project consists of a partially developed system with a functional architecture that requires finalization, refactoring, validation, and deployment to operate in a stable, reliable manner with correct metrics.

The system collects historical data, processes information, generates rankings, exports files by ID range, and provides an interface for visualization and analysis. The current code is functional but concentrated in a single monolithic file of approximately 15,600 lines, making maintenance and evolution unfeasible.

What Already Exists (Confirmed Codebase)
Backend: FastAPI (Python) with approximately 180 endpoints.

Frontend: React + TypeScript + Vite, with functional dashboard.

Database: MongoDB with modeled collections and validated historical data (over 7,000 records).

Machine Learning: Pipeline with implemented models (Gradient Boosting, Random Forest, LSTM).

Automation: Selenium for data collection (scraping).

Corrections Reported as Completed but Not Verified in Production
The following corrections have been documented as completed, but are not deployed on the VPS and therefore cannot be considered delivered until verified in production:

Future-data leakage fix with permanent guard

Deterministic and reproducible system (fixed seeds, stable ranking)

Immutable IDs (combo_id) for each combination

Protection against master file overwriting

Hash validation (checksum) between generated and exported files

Secure learning reset with archiving

Diversity optimization (M6)

Expanded historical data (7,268 draws)

Rebuilt historical features (respecting format changes)

What Needs to Be Done (Mandatory Scope)
1. Architecture Refactoring
Split main.py (15,608 lines) into modular components:

routers/ – HTTP endpoints

services/ – business logic

repositories/ – data access

models/ – schemas and validation

ml/ – Machine Learning and ranking pipeline

2. VPS Deployment
Deploy all corrections, improvements, and expanded data to production environment.

Configure the system to run stably and continuously.

3. Verification and Validation of Reported Corrections
Validate that all listed corrections actually work in the production environment.

Fix anything that is not working as expected.

4. Continuous Learning from the First Draw
Configure the system to process and learn from the first available draw of each lottery.

Respect chronological order and format/rule changes over time.

Ensure learning is incremental and cumulative.

5. Learning Reset Button
Implement (or verify and finalize) a button in the interface that:

Deletes all previous learning

Restarts processing from the first available draw

Preserves immutable IDs and already generated master files

Archives old data before deletion (safety)

6. Learning Evolution Progress Bar
Implement (or verify and finalize) a visual indicator that shows:

Historical processing progress (draws processed vs. total)

Evolution of the stability metric between consecutive executions

Charts and indicators of continuous learning

7. Master File Generation by Range
Implement export of files by ID range.

Preserve immutable IDs and original order (no renumbering).

Allow download through the user interface.

8. Master File Maintenance with Hash
Store each generated master file with its respective hash (checksum).

Ensure the downloadable file is identical to the internally generated one.

Provide file history for auditing.

9. Stability Metric Between Executions
Calculate, store, and display the position difference of the first prize between one draw and the next.

Display evolution on the dashboard with charts, alerts, and stability indicators.

10. Feedback Loop with Exponential Penalty
Adjust the continuous learning mechanism to penalize large variations between consecutive executions.

Apply exponential penalty when the difference exceeds the expected limit.

11. Incremental Reordering
Replace full ranking reordering with local incremental adjustments.

Preserve the relative position of the prize between executions, avoiding abrupt fluctuations.

12. Dashboard Refactoring
Replace generic charts with actionable metrics:

Evolution of the difference between consecutive executions

Percentage of executions within expected limit

Automatic alerts for critical variations

Historical averages, medians, best and worst results

Learning evolution progress bar

13. Pre-2005 Data Format Fix (El Gordo)
Correctly handle draws prior to 2005 (format 6/49 vs 5/54).

Ensure the system does not ignore or corrupt this data.

14. Daily Processing Automation
Configure the system to run automatically after each new draw.

Update rankings, metrics, and master files without manual intervention.

Required Technical Skills
Backend: Advanced Python (FastAPI, Pydantic, asyncio), modular code structuring.

Database: MongoDB (pymongo), modeling and optimized queries.

Frontend: React, TypeScript, Vite, REST API integration.

Automation: Selenium, scraping, authenticated website navigation.

Machine Learning: scikit-learn, PyTorch (existing models – no need to create new ones).

Infrastructure: Linux, VPS, systemd, Git/GitHub.

Plus: Docker, CI/CD, ranking optimization, immutable files, hash validation, stability metrics.

Estimated Timeline
The system has most of the code written, but nothing has been validated in production.

Estimated timeline: 3 to 5 weeks, depending on the professional's experience.

Payment Terms
Payment per milestone, with validation of each stage before release. Milestones will be defined based on the scope above.

Important Notes
Source code already exists and is available.

Many corrections have been reported as completed, but are not deployed on the VPS – therefore, they need to be verified, validated, and, if necessary, redone.

No new Machine Learning models need to be developed.

The focus is on finalization, organization, correct metrics, continuous learning, master file generation by range, verification of reported corrections, and deployment.

The system must learn from the first available draw.

Master files must be generated with immutable IDs, exported by range, and validated by hash.

The reset button and learning evolution progress bar must be implemented or finalized and displayed on the dashboard.

How to Apply
Submit a proposal with:

Brief presentation of your experience with the listed technical requirements.

Suggested approach for refactoring, verification of reported corrections, continuous learning, master file generation by range, stability metrics, reset button, and learning evolution progress bar.

Estimated timeline and detailed cost breakdown by milestones.

Examples of previous work with similar systems (automation, ranking, dashboards, immutable files, stability metrics).
python linux machine learning (ml) git react.js mongodb process validation data analysis selenium automation
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.