Fine-Tune a BERT Model for Text Classification, Optimized for Small/Edge Devices
Budget / Salary₹3,000–4,000
TypeFreelance project
LocationRemote
Posted1 hour ago
I need a DistilBERT-based text-classification model that not only reaches strong accuracy but can also run smoothly on memory-limited edge hardware. You will help me build the entire pipeline: first by working with me to assemble and label a solid training corpus, then by fine-tuning DistilBERT on that data.
After training, I want the network aggressively compressed—quantisation down to INT8, pruning where it makes sense, and any additional knowledge-distillation tricks you feel will preserve accuracy. The compact model must export cleanly to a deployment-friendly format (ONNX is my first choice, though TensorFlow Lite or TorchScript are fine if they give better performance).
Before and after optimisation, benchmark accuracy, size on disk, inference latency and peak RAM so we can see exactly what we gained. Finally, package a lightweight demo: a Python script or micro-API that loads the exported model and runs live inference on the target device.
Please send a detailed project proposal describing the approach, toolchain (e.g. PyTorch + Hugging Face Transformers, ONNX Runtime, TensorFlow Lite), and the timeline you foresee for data creation, training, compression and final validation.
After training, I want the network aggressively compressed—quantisation down to INT8, pruning where it makes sense, and any additional knowledge-distillation tricks you feel will preserve accuracy. The compact model must export cleanly to a deployment-friendly format (ONNX is my first choice, though TensorFlow Lite or TorchScript are fine if they give better performance).
Before and after optimisation, benchmark accuracy, size on disk, inference latency and peak RAM so we can see exactly what we gained. Finally, package a lightweight demo: a Python script or micro-API that loads the exported model and runs live inference on the target device.
Please send a detailed project proposal describing the approach, toolchain (e.g. PyTorch + Hugging Face Transformers, ONNX Runtime, TensorFlow Lite), and the timeline you foresee for data creation, training, compression and final validation.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.