← Tillbaka till jobb

Post-Training Engineer (RL Environments)

  • På plats
  • Finland
  • Engelska
  • Publicerad 07.09.26 08:21

Would you like to operate at the frontier of AI evaluation, post-training, and model improvements? We are now expanding our core Hashlist AI research team to support the creation of domain-specific RL environments for constraint-based embedded programming & complex enterprise engineering workflows. What you will do:Build and automate our platform for creating RL environmentsConstruct simulated worlds and explore data shapes that expose meaningful model failure modes across the embedded coding domain & related enterprise workflowsTurn AI training objectives into concrete data and evaluation specificationsBuild the reward layer + reward-hacking mitigationRun rollouts at scale. Hundreds of sandboxed attempts per task in parallelFine-tune open-source models: Before-and-after fine-tunes, failure reports, and dataset exportsCommunication with our clients (OEMs & AI Labs) on specific research or fine-tuning questions they have Skills needed:Production machine learning depth. Python and PyTorch, reinforcement learning training with TRL, verl, SkyRL or OpenRLHF, PPO, GRPO or DPO, LoRA fine-tuning, and inference with vLLM or SGLang.Training or evaluation systems, including at least one RL environment you built end-to-end and trained a model against. Comfortable with the infrastructure around it, e.g Docker, or similar containerization tools, to design and monitor systems at scaleExperience with harness & agentic optimisation for evalsStrong familiarity with common reinforcement learning algorithms and methods, especially with respect to post-training LLMsHigh level of personal drive, motivation, and good communication skills. Bonus:You have published an environment, benchmark or evaluation harness we can look at.You have worked with embedded, safety-critical or other physical engineering software. Company benefitsCompetitive compensation + meaningful equityCentral office in HelsinkiLunch benefitBe a part of a quickly scaling tech company working directly with model providers