Jobiglo

Aucun resultat.

RLHF Specialist

Odixcity Consulting

Remote
Remote Mid 🇬🇧 English
Python PyTorch JAX TensorFlow LoRA QLoRA Llama 2 Mistral Gemma LabelBox Scale AI Snorkel AWS SageMaker GCP Vertex AI

Description du poste

About the role

We are looking for an RLHF Specialist to lead the design and execution of human‑feedback pipelines that improve large language models. The role is fully remote and focuses on aligning model behavior with safety, factual accuracy and human values through reinforcement learning techniques.

Key responsibilities

  • Generate high‑quality preference data by ranking model responses on helpfulness, honesty and harmlessness.
  • Design multi‑turn prompts to stress‑test reasoning and safety, and write chain‑of‑thought explanations for reward‑model training.
  • Collaborate with ML engineers to analyse failure modes, identify data gaps and propose interventions.
  • Develop and iterate annotation strategies, ensuring consistent preference scoring across a global team.
  • Probe models for biases, hallucinations or vulnerabilities, document findings and suggest corrective data.
  • Analyse edge cases where reward models behave unexpectedly and provide detailed feedback.
  • Create templated instruction sets for large annotation teams and translate complex RL concepts into repeatable tasks.
  • Maintain a personal benchmark set, regularly re‑evaluating new model versions against historical performance.

Required profile

  • Minimum 2 years experience in data annotation, model evaluation, computational linguistics or AI trust & safety.
  • Strong proficiency in Python and deep‑learning frameworks such as PyTorch, JAX or TensorFlow.
  • Solid understanding of reinforcement‑learning concepts (PPO, trust‑region methods, reward hacking) applied to language generation.
  • Hands‑on experience fine‑tuning open‑source models (e.g., Llama 2/3, Mistral, Gemma) using LoRA/QLoRA.
  • Experience with annotation platforms (LabelBox, Scale AI, Snorkel) and human‑in‑the‑loop workflows.
  • Familiarity with cloud ML services (AWS SageMaker, GCP Vertex AI) and open‑source alignment libraries.

Required skills

  • Python
  • PyTorch
  • JAX
  • TensorFlow
  • LoRA
  • QLoRA
  • Llama 2 / Llama 3
  • Mistral
  • Gemma
  • LabelBox
  • Scale AI
  • Snorkel
  • AWS SageMaker
  • GCP Vertex AI

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Odixcity Consulting.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Pourquoi signalez-vous cette offre ?

Merci pour votre signalement. Nous allons examiner cette offre.

Postulez en 30 secondes

Entrez votre email pour postuler. Un compte sera cree automatiquement.

En continuant, vous acceptez nos conditions d'utilisation.

Deja un compte ? Connexion

💬 Contactez-nous sur Telegram Discuter sur WhatsApp

Publie il y a 1 mois

Expire dans 1 semaine

49 vues · 0 interesses

Boostez vos chances

Importez votre CV : nous vous proposons les offres qui matchent votre profil.

Analyse de votre CV en cours...

Odixcity Consulting