← Sottava, jobs the hour they open
2 mo agofound 4 h ago
AI Inference Engineer
Posted by Fuse Energy on 20 July 2026, 82 days ago. Still on their Workable board when we checked 20 min ago.
Read out of the posting
LevelNot stated
Experience askedNot stated
EmploymentNot stated
LocationLondon, England, United Kingdom
RemoteYes
Visa sponsorshipNot stated
SalaryNot published, and most postings do not
Posted2026-07-20
Found viaworkable, direct from their system
We saw it 3 months after it went up.
The posting, as the company wrote it
Employment: Full-time
Fuse Energy is an energy startup on a mission to make energy abundant and affordable, fast. We combine first-principles thinking with cutting-edge technology to build a radically better energy system.
We've raised over $200M from top-tier investors including Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, 20VC, Hummingbird and Collaborative Fund, alongside strategic angels including Nico Rosberg and GPs behind Meta, Revolut, Spotify and Uber.
We're building a fully integrated energy company: developing our own solar, batteries and other generation projects, building our own hardware, improving and developing grid infrastructure, trading power in real time, using AI across the business, and installing distributed energy in homes. By selling directly to consumers we cut out the middleman, lower costs and pass the savings on to our customers.
As data centres become one of the largest and fastest-growing sources of electricity demand, Fuse is expanding into high-performance compute infrastructure at the intersection of energy and AI. We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch, and we're looking for the founding engineer to own the latter. Reporting directly to the CTO, you'll own the layer above kernels and hardware: how models actually get served, scaled and delivered against committed performance targets. Few companies can pair real power delivery with real compute the way Fuse can, which puts inference serving at the heart of our offering.
Responsibilities
Define Fuse's inference serving strategy and architecture from first principles
Design and build the serving stack: request routing, batching, scheduling and autoscaling for high-throughput, latency-sensitive inference workloads
Own model-level optimisation strategy for serving, deciding where and how to apply quantisation, distillation, speculative decoding and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineers
Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server or equivalents)
Translate throughput, latency and uptime commitments into concrete technical specifications and serving capacity plans
Act as direct technical owner of inference performance and reliability
Work closely with the CUDA and GPU engineering teams to integrate custom kernels and hardware performance work cleanly into the serving layer
Set the standards, tooling and benchmarks this function will run on as it grows
Requirements
4+ years building or operating large-scale inference serving systems, or equivalent strong project/industry experience
Deep, hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding)
Strong systems thinking, able to reason about the full path from incoming request to served response across a large cluster
Comfortable working directly with GPU/CUDA engineers to integrate low-level performance work into a serving system
A track record of making high-stakes architecture calls and owning the outcome
Comfort operating without a playbook: this is a founding role shaping a new function, not joining an established one
Bonus: Triton or custom ML inference/training frameworks; autoscaling or capacity planning for large-scale inference; multi-tenant serving or SLA-driven infrastructure; background at a hyperscaler, frontier AI lab or large-scale distributed inference system; Kubernetes/Slurm; interest in energy markets, grid systems or sustainability-focused compute
Benefits
Competitive salary and eligibility for equity
Biannual bonus scheme
Fully expensed tech to match your needs
Private health insurance
Breakfast and dinner allowance for office-based employees
As we hire globally, benefits vary by location.
Copied from Fuse Energy’s own board, not rewritten. Original ↗
Also open at Fuse Energy
Why this page exists
We read companies’ own hiring systems every hour, 1,769 of them, and show a job the hour it opens instead of when a job board gets around to indexing it. We saw it 3 months after it went up.
The feed is free. No card, no trial to expire.
Apply at Fuse EnergyA free account first, no card