About this role
Job title: ML Systems Engineer - Model Training and Infrastructure (SWE-focused LLMs)
Location: London; full in-office working as default
Start date: ASAP
Compensation: £80,000 - £110,000 Base Salary & £80,000 - £110,000 Share options.
___________________________________________________________________________
Help build the software engineers of the future
Cosine is building autonomous AI engineers that plan, write and ship code inside real development workflows.
Our agents work across complex software systems, and our Lumen models are trained to do more than produce code that looks correct. They are built to understand existing architectures, follow established patterns and produce software that engineers can actually maintain.
We develop our agent tooling entirely in-house and post-train open-source models for reliable, enterprise-grade coding performance. Our products are designed for on-premise, VPC and fully air-gapped environments, including security-critical settings where control, privacy and robustness are non-negotiable.
In 2024, Cosine achieved a 72% score on OpenAI’s SWE-Lancer benchmark, placing us among the strongest real-world software-engineering AI systems evaluated.
We’re now looking for an ML Systems Engineer to help train the next generation of Lumen models.
This is a highly hands-on role at the intersection of machine learning, software engineering, data and infrastructure. You’ll build the environments in which models learn to write software, develop the pipelines that generate and curate training data, and run the fine-tuning and reinforcement-learning workloads that shape model behaviour.
If you’re excited by the idea that the future quality of coding agents will be determined not just by model architecture, but by the quality of their data, environments and reward functions, this is an opportunity to work directly on that problem.
___________________________________________________________________________
The role
You’ll work closely with ML researchers, infrastructure engineers and product teams to decide:
- What Lumen should learn next.
- How to generate the right training data.
- How to design environments that reflect real software-engineering work.
- How to reward models for producing useful, maintainable code.
- How to measure whether a new training run genuinely improves the experience for engineers.
The systems you build will sit directly inside our model-training loop. Models will write code, use tools, run tests and interact with real repositories. Your work will determine how those interactions are generated, evaluated and fed back into future training.
This is not a narrow research role and it is not traditional MLOps. You’ll move between custom PyTorch code, distributed data pipelines, Dockerised services, RL environments, evaluation infrastructure and production-quality software.
___________________________________________________________________________
What you’ll do
Build the training systems behind Lumen
- Contribute to the end-to-end training of software-engineering models.
- Implement supervised fine-tuning pipelines using curated code and conversation datasets.
- Build reinforcement-learning loops in which models write code, run tests and use development tools.
- Develop custom PyTorch dataloaders, training objectives and evaluation workflows.
- Run and analyse fine-tuning and RL experiments across large, modern open-source models.
Create the data that teaches models to engineer
- Develop synthetic data-generation pipelines for future RL and fine-tuning runs.
- Design systems for producing, filtering, transforming and sampling large-scale datasets.
- Work with object storage, dataset sharding and data-quality checks.
- Investigate which examples, tasks and sampling strategies lead to better model behaviour.
- Turn model failures into concrete improvements to training data and future experiments.
Build reliable RL infrastructure
- Design, build and deploy containerised services that support model training and evaluation.
- Use Docker and orchestration platforms such as Kubernetes to operate RL infrastructure.
- Build environments where agents can modify repositories, run tests, use tools and receive meaningful feedback.
- Improve the reliability, reproducibility and observability of large-scale training runs.
- Work across Python, Go and the surrounding infrastructure required to run these systems.
Improve how we evaluate SWE models
- Help maintain and extend evaluation suites for code models.
- Build evaluations around unit tests, benchmark suites, repository-level tasks and real engineering workflows.
- Analyse model failure modes, reward-hacking risks and regressions.
- Develop better ways to measure code quality, maintainability, correctness and architectural fit.
- Feed evaluation results back into model, data and infrastructure decisions.
Shape the next training direction
- Work with research, infrastructure and product teams to identify the highest-value training problems.
- Turn broad goals such as “make Lumen better at X” into clear experiments and measurable outcomes.
- Develop opinionated reward functions for software-engineering agents.
- Document decisions, communicate trade-offs and help the wider team understand what the results mean.
- Take ownership of projects from initial idea through to deployment and iteration.
___________________________________________________________________________
What we’re looking for
You may come from software engineering, ML infrastructure, data engineering, applied machine learning or a closely related field.
You should be comfortable with:
Strong software engineering fundamentals
- Typically 3–5 years of experience, or equivalent evidence of strong technical ability.
- Reading, debugging and writing non-trivial production code.
- Working primarily in Python and Go.
- Caring about correctness, maintainability and code quality as much as model metrics.
- Taking ownership of ambiguous technical problems and turning them into working systems.
Training frameworks and ML systems
- Practical experience with at least one of PyTorch, TensorFlow or JAX.
- Implementing custom training loops, losses or dataloaders.
- Understanding the relationship between data, objectives, evaluation and model behaviour.
- A willingness to work close to the code rather than treating training infrastructure as a black box.
Containers and cloud infrastructure
- Experience with Docker and container-management or orchestration platforms such as Kubernetes.
- Experience with at least one major cloud platform, such as GCP, AWS or Azure.
- An understanding of how to build services that are reliable, observable and straightforward to operate.
- Comfort working across application code, infrastructure and deployment systems.
Data engineering instincts
- Experience working with large-scale datasets and object storage.
- Understanding of sharding, filtering, sampling and dataset versioning.
- The judgement to recognise that data quality can matter as much as model architecture.
- An interest in building data pipelines that are reproducible and useful for experimentation.
Clear communication and ownership
- The ability to explain technical decisions and experimental results clearly.
- Comfort documenting trade-offs and walking others through your reasoning.
- A practical approach to deciding what to build, what to measure and what to leave out.
- The curiosity to investigate failures rather than dismissing them as noise.
You do not need to have worked on every part of the stack. We care more about strong engineering fundamentals, technical judgement and the ability to learn quickly than about matching every keyword.
Nice to have
You don’t need all of these, but experience in the following areas would help you get up to speed quickly:
- Synthetic data generation for code or language models.
- Training LLMs in distributed environments.
- Reinforcement learning, preference optimisation or reward modelling.
- Data tooling such as SQL, Apache Iceberg or DuckDB.
- LLM-as-a-judge systems or automated code evaluation.
- Reward-hacking detection, robustness evaluation or safety-focused training.
- Open-source contributions to LLM tooling, ML infrastructure or RL libraries.
- Experience working with repository-level code tasks or software-engineering benchmarks.
___________________________________________________________________________
What success looks like
In your first few months, you will:
- Become a trusted engineering partner to the research team.
- Contribute to reliable data-generation, training and evaluation workflows.
- Ship improvements to the infrastructure used in Lumen fine-tuning and RL runs.
- Help identify why models fail on real software-engineering tasks.
- Turn those failures into concrete changes to data, environments, reward functions or evaluation.
- Build a strong understanding of the systems required to train capable SWE agents.
Longer term, you’ll help define how Cosine trains software-engineering models: what they learn, how they practise, what they are rewarded for and how we decide whether they are genuinely getting better.
___________________________________________________________________________
Why this role matters
Coding agents are moving quickly, but producing plausible code is not the same as being a good software engineer.
The models we build need to work within existing codebases. They need to understand context, respect architectural decisions, write maintainable code, use tools effectively and recover when their first attempt fails.
That behaviour will not emerge from model scale alone. It will come from better environments, better data, better evaluations and better incentives.
You’ll work directly on all four.
Your contributions will influence the Lumen models used in Cosine’s self-serve and enterprise products, including deployments in organisations with demanding security and infrastructure requirements. You’ll have close proximity to the research, infrastructure and product decisions that determine where the system goes next.
This is a role for someone who wants to work on the full stack of modern model training, while staying grounded in the standards of production software engineering.
___________________________________________________________________________
Why join Cosine
- Direct impact: Your work will shape the next generations of Lumen software-engineering models.
- Real scale: Work with large open-source models, long context lengths and multi-node training runs.
- Full-stack ML engineering: Move between PyTorch, distributed systems, data curation, RL infrastructure and MLOps.
- Frontier problems: Help define how AI systems learn to perform real software-engineering work.
- Meaningful ownership: Take ideas from vague training objective to deployed system and measurable result.
- Close collaboration: Work alongside research, infrastructure, product and engineering in our Hoxton office.
We’re an in-office team, five days a week, by design. The problems we’re solving benefit from close collaboration, fast feedback and shared context.
If you want to help build AI systems that write software engineers can trust, this is an opportunity to work on one of the most important problems in the field.
Come help us build the future of software engineering.
___________________________________________________________________________
Cosine is an equal opportunity employer.
We value diverse backgrounds, perspectives, and ways of thinking, and we’re committed to creating an inclusive and respectful workplace.
We encourage applications from anyone who meets the role requirements, even if you don’t meet every single qualification. If you need reasonable adjustments at any stage of the hiring process, we’re happy to discuss them.
___________________________________________________________________________
Compensation, Benefits & Ways of Working
We’re an in-office team, five days a week, by design. We believe the work we’re doing benefits from being together, collaborating closely, and building shared context.
What you can expect:
- Competitive salary , benchmarked to the market
- Equity / share options , so you share in the upside you help create
- 30 days’ holiday + bank holidays
- Genuine 9–5 working hours - we don’t expect late nights or weekend work
- Work hard in the office, collaborate closely, and switch off properly
- Dog-friendly office - bring your dog to work
- Weekly team breakfast & lunch
- Monthly socials
- Pension
- High-quality equipment to do your best work
We care about focus, sustainability, and doing great work — not performative overwork. We value people who show up, contribute thoughtfully, collaborate well with their colleagues, and then go home.
This role won’t suit everyone. But if you want structure, clarity, strong collaboration, and a team that takes both the work and work-life balance seriously, it’s a great place to be.
___________________________________________________________________________
Agency & Data Protection Notice
To comply with UK GDPR and our internal data-protection and equal-opportunity obligations, we only accept candidate applications and agency submissions via our Applicant Tracking System (ATS). This ensures appropriate privacy notices, lawful processing, auditability, and consistent retention controls.
Any CVs or candidate details received outside the ATS (including via email, Slack, or direct message) will be treated as unsolicited, will not be considered as part of the recruitment process, and will not give rise to any fee or payment obligation.