← Sottava, jobs the hour they open
2 mo agofound 1 h ago
Senior Site Reliability Engineer
Posted by Runware on 20 July 2026, 82 days ago. Still on their Workable board when we checked 1 h ago.
Read out of the posting
LevelNot stated
Experience askedNot stated
EmploymentNot stated
LocationUnited Kingdom
RemoteYes
Visa sponsorshipNot stated
SalaryNot published, and most postings do not
Posted2026-07-20
Found viaworkable, direct from their system
We saw it 3 months after it went up.
The posting, as the company wrote it
Employment: Full-time
Experience: Mid-Senior level
Education: Bachelor's Degree
Runware is building high-performance infrastructure and products to power the worlds intelligence. Our platform enables developers and businesses to run fast, scalable inference across image, video and emerging modalities, while our Serverless platform allows customers to deploy and scale their own AI models on production-grade GPU infrastructure.
As a Site Reliability Engineer at Runware, you will help ensure these systems remain reliable, performant and resilient as we scale. This is a highly technical, hands-on role working across software, infrastructure and production operations to improve observability, reduce incidents, eliminate operational toil and build lasting improvements across complex distributed systems.
What you’ll do
Own and improve the reliability, availability and performance of critical production services across the Runware platform
Define and evolve our reliability practices, including SLIs, SLOs, alerting, observability and production-readiness standards
Investigate complex production issues across distributed systems, APIs, networking, queues, databases and GPU-backed workloads, participating in our engineering on-call rotation
Lead and contribute to incident reviews and RCAs, turning recurring failure modes into lasting engineering improvements
Reduce operational toil through automation, automated remediation and improvements to deployment safety, recovery and system resilience
Work closely with Engineering and DevOps teams on capacity planning, performance, scaling and architectural improvements as the platform grows
Requirements
Have strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering or similar role
Have a strong understanding of distributed systems and are comfortable debugging across applications, databases, queues, containers, networking and infrastructure
Have experience designing and operating observability systems using metrics, logs and distributed tracing
Understand SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management and reducing operational toil
Have experience with Kubernetes, containers, IaC and automated deployment practices, alongside the ability to write software and automation using languages such as Python, Go or PHP
Take strong ownership of production problems and are comfortable participating in an engineering on-call rotation, taking issues from initial investigation through to long-term remediation
Bonus
Experience operating high-throughput or low-latency APIs and distributed systems
Experience with bare-metal infrastructure, GPU environments or AI and ML workloads
Experience with RabbitMQ or other distributed messaging and queueing systems
Experience operating MySQL, Redis, ClickHouse or similar production data systems
Experience with global traffic management, load balancing, CDN platforms and hybrid infrastructure environments
Experience building automated scaling, capacity management or self-healing systems
Benefits
We’re a remote-first collective, meeting in person twice a year to plan, brainstorm, celebrate wins, and enjoy some face-to-face time. We have core hours for cooperative working and calls, but outside of that your calendar is yours. Work the hours that let you perform at your peak while also building a healthy life.
We move quickly and hold a high bar, but we also believe sustainable performance matters. After major releases and milestones, we create space for the team to properly unplug, recharge and come back with the energy to tackle what's next.
Generous paid time off – vacation, sick days, public holidays
Meaningful stock options – share in the upside you create
Remote-first setup – work from home anywhere we can employ you
Flexible hours – own your schedule outside core collaboration blocks
Family leave – paid maternity, paternity, and caregiver time
Company retreats – twice-yearly gatherings in inspiring locations
Copied from Runware’s own board, not rewritten. Original ↗
Also open at Runware
Why this page exists
We read companies’ own hiring systems every hour, 1,769 of them, and show a job the hour it opens instead of when a job board gets around to indexing it. We saw it 3 months after it went up.
The feed is free. No card, no trial to expire.
Apply at RunwareA free account first, no card