Prompt Design Regression Study
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted1 hour ago
I’m running a systematic investigation into how prompt phrasing influences in-context regression performance. The work spans six separate axes, and for this engagement I want you to concentrate on two of them—Data preprocessing and Evaluation metrics—while keeping the other axes in mind so our findings remain extensible.
The core set-up centres on text-based prompts only; no visual or multimodal inputs will appear in this round. All experiments will target polynomial regression tasks, so the prompts, data splits, and metric choices should reflect the non-linear nature of the underlying relationships.
Here is what I need from you:
• Curate or generate a clean, well-documented dataset suitable for polynomial regression, then outline the preprocessing steps you apply (normalisation, tokenisation strategy, train/validation/test partitioning, and any feature engineering).
• Design several prompt templates that systematically vary along the preprocessing and evaluation dimensions we’ve highlighted.
• Implement an evaluation pipeline that reports standard regression scores—MSE, MAE, R-squared—as well as any prompt-specific diagnostic statistics you find insightful.
• Run the experiments across at least two popular transformer-based language models so we can compare cross-model behaviour without making architecture a primary axis.
• Deliver a concise technical report (Jupyter notebook or Markdown) explaining methodology, code snippets, results tables, plots, and your interpretation of the trends observed across our two focal axes.
Acceptance criteria
1. Reproducible code (Python, PyTorch or JAX) with clear README.
2. All metrics reproducible on my machine using seed values you provide.
3. Discussion connects empirical results back to Data preprocessing and Evaluation metrics axes in a way that can scale to the remaining four axes later.
If you’re comfortable navigating prompt engineering for regression tasks and can translate results into clear, actionable insights, I’m eager to review your proposal.
The core set-up centres on text-based prompts only; no visual or multimodal inputs will appear in this round. All experiments will target polynomial regression tasks, so the prompts, data splits, and metric choices should reflect the non-linear nature of the underlying relationships.
Here is what I need from you:
• Curate or generate a clean, well-documented dataset suitable for polynomial regression, then outline the preprocessing steps you apply (normalisation, tokenisation strategy, train/validation/test partitioning, and any feature engineering).
• Design several prompt templates that systematically vary along the preprocessing and evaluation dimensions we’ve highlighted.
• Implement an evaluation pipeline that reports standard regression scores—MSE, MAE, R-squared—as well as any prompt-specific diagnostic statistics you find insightful.
• Run the experiments across at least two popular transformer-based language models so we can compare cross-model behaviour without making architecture a primary axis.
• Deliver a concise technical report (Jupyter notebook or Markdown) explaining methodology, code snippets, results tables, plots, and your interpretation of the trends observed across our two focal axes.
Acceptance criteria
1. Reproducible code (Python, PyTorch or JAX) with clear README.
2. All metrics reproducible on my machine using seed values you provide.
3. Discussion connects empirical results back to Data preprocessing and Evaluation metrics axes in a way that can scale to the remaining four axes later.
If you’re comfortable navigating prompt engineering for regression tasks and can translate results into clear, actionable insights, I’m eager to review your proposal.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.