Performance Engineer – Language Runtime (Contractor)

Huawei Technologies Research & Development (UK) LimitedCambridge, EnglandOn-siteContractListed 7 hours ago

Apply now

About this role

About Huawei Research and Development UK Limited

Founded in 1987, Huawei is a leading global provider of information and communications technology (ICT) infrastructure and smart devices. We have 207,000 employees and operate in over 170 countries and regions, serving more than three billion people around the world.

Our vision and mission is to bring digital to every person, home and organization for a fully connected, intelligent world. To this end, we will drive ubiquitous connectivity and promote equal access to networks; bring cloud and artificial intelligence to all four corners of the earth to provide superior computing power where you need it, when you need it; build digital platforms to help all industries and organizations become more agile, efficient, and dynamic; redefine user experience with AI, making it more personalized for people in all aspects of their life, whether they’re at home, in the office, or on the go.

This spirit of innovation has led Huawei to work in close partnership with leading academic institutions in the UK to develop and refine the latest technologies. With a shared commitment to innovation and progress, both parties have worked together to achieve common goals and establish a strong partnership. The partnership between UK and Huawei help to develop the technologies of the future that will transform the way we all communicate, work and live.

For the past 30 years we have maintained an unwavering focus, rejecting shortcuts and easy opportunities that don't align with our core business. With a practical approach to everything we do, we concentrate our efforts and invest patiently to drive technological breakthroughs.

This strategic focus is a reflection of our core values:

- Staying customer-centric,
- Inspiring dedication,
- Persevering,
- Growing by reflection.

Huawei Research and Development UK Limited Overview

Huawei’s vision is a fully connected, intelligent world. To achieve this, we work to inspire passion for basic research around the world. Our combined passion drives development across the global innovation value chain. Huawei has the largest Research and Development organization in the world with 96,000+ employees in research centers around the globe. In the UK, we already have design centers in Cambridge, London, Edinburgh and Ipswich. We continue to explore and define new research directions and new services. We have expanded our collaborations with academic researchers; researched new network architectures, integration of communications and key enabling technologies; and developed the fundamental theories of these technologies. We invite you to join us on this exciting journey and drive your career forward.

Job Summary

We are looking for a Performance Engineer to make managed language runtimes run measurably faster on ARM64 mobile-class CPUs. You will work across the full runtime stack – interpreter, JIT and AOT compilers, garbage collector and runtime libraries – and you will follow every optimisation down to the generated machine code and the hardware counters that explain it. This is a hands-on role for someone who finds the bottleneck, writes the patch and proves the speedup with rigorous measurement. You will sit inside the CPU team, so what you learn about how runtimes stress the pipeline, caches and branch predictors feeds directly into next-generation core design, and you will work with LLM-agent pipelines that automate parts of the analyse, patch, benchmark and validate loop.

Key Responsibilities:

– Runtime profiling : find and rank performance bottlenecks in interpreter, JIT/AOT, GC and runtime-library code on real application and benchmark workloads, using sampling profilers, PMU counters, top-down analysis and flame graphs.

– Execution engine optimisation : improve interpreter and compiler tiers end to end.

– Interpreter dispatch (threaded code, tail calls), inline caches, hidden classes / shapes, and tier-up and deoptimisation policy.

– JIT and AOT code quality: instruction selection, register allocation, inlining heuristics, escape analysis, check elimination and code-cache layout.

– Startup time: snapshots, AOT images, lazy compilation and warm-up behaviour.

– Memory management : tune allocation and collection for throughput, pause time and footprint.

– Allocation fast paths, write and read barriers, generational, concurrent and compacting collectors, and heap sizing policy.

– Object layout and pointer compression, and memory behaviour of pointer-chasing heaps.

– Microarchitectural analysis : explain runtime behaviour in hardware terms – indirect-branch misprediction in dispatch loops, I-cache and iTLB pressure from JIT code, load-to-use latency and prefetch behaviour on object graphs – and turn the findings into concrete software changes and hardware feedback.

– Measurement infrastructure : build benchmarking and regression-tracking pipelines with noise control, statistically sound A/B comparison and reproducible results, so that every claimed gain is trustworthy.

– Automated optimisation : build and use LLM-agent pipelines that analyse profiles, propose patches, benchmark them and validate correctness, and decide where automation can be trusted and where it cannot.

– Hardware–software co-design : work with CPU architects to identify where runtime workloads leave performance on the table and which hardware or ISA features would help.

– Upstream and share : land patches, write internal technical reports and, where appropriate, contribute to patents and publications.

This job description is only an outline of the tasks, responsibilities and outcomes required of the role. The jobholder will carry out any other duties as may be reasonably required by his/her line manager. The job description and personal specification may be reviewed on an ongoing basis in accordance with the changing needs of Huawei Research and Development UK Limited.

Required:

– Degree in Computer Science, Computer Engineering, Electrical Engineering or a related field, or equivalent practical experience.

– Strong C and C++ (or Rust), with the ability to work productively in a large, performance-critical codebase.

– Hands-on performance work in a language runtime, VM or compiler backend, with measured speedups you can walk us through in detail (for example V8, JavaScriptCore, SpiderMonkey, HotSpot, PyPy, LuaJIT, CPython, or an LLVM backend).

– Fluency with profiling tools (perf, PMU counters, flame graphs, or equivalents) and the ability to read disassembly and relate it to source.

– Solid grasp of a modern ISA (ARM64 or x86-64), calling conventions, memory model and the cost of branches and cache misses.

– Measurement rigour: you design experiments, control for noise and report results with appropriate statistics.

Please note that visa sponsorship is not available for this contractor role.

Desired:

– Working knowledge of CPU microarchitecture: out-of-order pipelines, branch prediction, caches, TLBs and prefetching.

– ARM64 experience, particularly on mobile or embedded SoCs.

– Contributions to an open-source runtime, compiler or GC, or a track record of upstream performance patches.

– Experience with LLM agents or ML-assisted tooling for code optimisation, or an end-to-end self-built system of that kind.

– Familiarity with cycle-accurate or architectural simulators such as gem5.

– Publications or patents in runtimes, compilers or computer architecture.