Arabic Multimodal LLM Fine-Tuning
Budget / Salary$30–250
TypeFreelance project
LocationRemote
Posted2 hours ago
I have Arabic Multimodal Dataset (Text + Image) designed for Multimodal Aspect-Based Sentiment Analysis (MABSA). Each data point consists of an Arabic review text paired with one or more relevant images.
I am seeking an expert LLM Specialist to build a comprehensive, state-of-the-art experimental framework to validate the dataset's learnability and establish standard benchmarks for the scientific community. The project involves task-specific fine-tuning using modern text and vision-language architectures
The dataset must be evaluated across two distinct tasks following standard evaluation benchmarks (e.g., SemEval standards):
Task 1: Aspect Category Detection (Identifying specific predefined aspect categories mentioned).
Task 2: Aspect-Based Sentiment Polarity Classification (Classifying sentiment as Positive, Negative, Neutral, or Irrelevant toward each aspect).
Scope of Work & Technical Strategy:
Data Stratification & Splitting:
Implement a strict Train / Validation / Test split. Ensure full reproducibility by fixing random seeds and documenting text/image preprocessing pipelines.
Experiment 1: Text-Only Fine-Tuned Baseline
Fine-tune Arabic-native encoder models (e.g., AraBERTv2, MARBERT, CAMeLBERT-DA).
Experiment 2: Multimodal Fine-Tuned Baseline (Text + Image)
Implement task-specific fine-tuning using architectures that natively support Arabic textual inputs paired with visual.
Recommended approaches include Multilingual CLIP (mCLIP) or modern open-source Vision-Language Models (VLMs) like Qwen2-VL or specialized LLaVA variants.
Note: each review is linked to one or multiple images.
Baseline Benchmarking:
Compute and report a Majority Class Baseline and a Random Baseline to provide context for the models' performance given any class imbalances.
Project Deliverables:
Reproducible Codebase: Fully documented, clean Python code (Jupyter Notebooks or structured scripts) utilizing PyTorch and Hugging Face libraries.
Separated Evaluation Matrices: Structured results tables reporting Accuracy, Precision, Recall, and Macro-F1 calculated separately for Task 1 (Aspect Detection) and Task 2 (Sentiment Classification). This must include a per-class/per-aspect F1 breakdown to account for data imbalances.
Training Dynamics & Diagnostic Charts: High-quality visual assets including Loss Convergence Curves (Training vs. Validation loss over epochs) and Confusion Matrices for both experiments to visually demonstrate the dataset's learnability.
Technical Framework Documentation: A clear summary detailing the exact model checkpoints, hyperparameters (learning rates, batch sizes, epochs), prompting templates (if any), and data fusion methods applied.
document: build the technical validation section with results and discussion that should be written in an academic way.
the deadline is 16-9-2026
I am seeking an expert LLM Specialist to build a comprehensive, state-of-the-art experimental framework to validate the dataset's learnability and establish standard benchmarks for the scientific community. The project involves task-specific fine-tuning using modern text and vision-language architectures
The dataset must be evaluated across two distinct tasks following standard evaluation benchmarks (e.g., SemEval standards):
Task 1: Aspect Category Detection (Identifying specific predefined aspect categories mentioned).
Task 2: Aspect-Based Sentiment Polarity Classification (Classifying sentiment as Positive, Negative, Neutral, or Irrelevant toward each aspect).
Scope of Work & Technical Strategy:
Data Stratification & Splitting:
Implement a strict Train / Validation / Test split. Ensure full reproducibility by fixing random seeds and documenting text/image preprocessing pipelines.
Experiment 1: Text-Only Fine-Tuned Baseline
Fine-tune Arabic-native encoder models (e.g., AraBERTv2, MARBERT, CAMeLBERT-DA).
Experiment 2: Multimodal Fine-Tuned Baseline (Text + Image)
Implement task-specific fine-tuning using architectures that natively support Arabic textual inputs paired with visual.
Recommended approaches include Multilingual CLIP (mCLIP) or modern open-source Vision-Language Models (VLMs) like Qwen2-VL or specialized LLaVA variants.
Note: each review is linked to one or multiple images.
Baseline Benchmarking:
Compute and report a Majority Class Baseline and a Random Baseline to provide context for the models' performance given any class imbalances.
Project Deliverables:
Reproducible Codebase: Fully documented, clean Python code (Jupyter Notebooks or structured scripts) utilizing PyTorch and Hugging Face libraries.
Separated Evaluation Matrices: Structured results tables reporting Accuracy, Precision, Recall, and Macro-F1 calculated separately for Task 1 (Aspect Detection) and Task 2 (Sentiment Classification). This must include a per-class/per-aspect F1 breakdown to account for data imbalances.
Training Dynamics & Diagnostic Charts: High-quality visual assets including Loss Convergence Curves (Training vs. Validation loss over epochs) and Confusion Matrices for both experiments to visually demonstrate the dataset's learnability.
Technical Framework Documentation: A clear summary detailing the exact model checkpoints, hyperparameters (learning rates, batch sizes, epochs), prompting templates (if any), and data fusion methods applied.
document: build the technical validation section with results and discussion that should be written in an academic way.
the deadline is 16-9-2026
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.