Sottava
Scanning 1,123 companies · 22,542 open jobs · last pass 8 min ago

← Sottava, jobs the hour they open

2 mo agofound 7 d ago

Software Engineer, Cloud

Ollama·Palo Alto·via ashby
What the posting is about

Build and scale Ollama's cloud inference platform. Design routing and capacity layer for cost, latency, and availability. Own multi-tenant infrastructure and build reliability, observability, and cost controls.

Read out of the posting
LevelNot stated
Experience askedNot stated
EmploymentNot stated
LocationPalo Alto
RemoteNot stated
Visa sponsorshipNot stated
SalaryNot published, and most postings do not
Posted2026-07-09
Found viaashby, direct from their system

We saw it 3 months after it went up.

The posting, as the company wrote it
Ollama is the most popular way for developers to access open models. What started as an open-source, local-first runtime is now the largest developer network in the open-model ecosystem: 8.9 million monthly active developers and over 67,000+ community-built integrations. We're backed by Y Combinator, Benchmark, 8VC, and Theory Ventures. Our team is small and talent dense. We're flat, low-ego, and fast-moving. We like people who are truth-seeking, passionate, design-driven, and who enjoy shipping code. ABOUT THE ROLE You'll build Ollama’s cloud, a scalable inference platform that lets developers run large, capable open models in their workflow. You'll work on high-throughput, low-latency distributed systems — inference serving, GPU fleet management, routing, metering, and the platform that Pro, Max, Team, and Enterprise customers rely on to process trillions of tokens. WHAT YOU'LL DO - Build and scale the inference platform that serves every request from ollama.com http://ollama.com. - Design the routing and capacity layer that places workloads across GPUs and regions for cost, latency, and availability. - Own multi-tenant infrastructure: isolation, quotas, usage metering, billing, and Pro/Max/team/enterprise tiering. - Build the reliability, observability, and cost controls for our team and customers YOU MAY BE A FIT IF - You have deep experience with high-throughput, low-latency distributed systems — inference serving, traffic routing, real-time data pipelines, or large-scale APIs. - You're comfortable with cost/performance tradeoffs at scale and have owned a production service end-to-end. - You've worked with Kubernetes, GPU scheduling, or inference infrastructure. - You think in terms of reliability, SLOs, and honest capacity planning. - Bonus: experience building an inference platform, GPU fleet management, or billing/metering for an AI service.

Copied from Ollama’s own board, not rewritten. Original ↗

Also open at Ollama

Why this page exists

We read companies’ own hiring systems every hour, 1,123 of them, and show a job the hour it opens instead of when a job board gets around to indexing it. We saw it 3 months after it went up.

The feed is free. No card, no trial to expire.