About the Role
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimize telemetry at scale. Built on open source and open standards and designed for interoperability across any stack, Grafana Cloud brings AI to observability and observability to AI, giving teams (and their agents) unified visibility so they can see, understand, and act on all their disparate data, wherever it lives, and move at the speed of their ambitions. Customers, including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce, rely on Grafana Labs. We are a 100% remote company with team members across 40+ countries, backed by leading investors including Lightspeed Venture Partners, Sequoia Capital, GIC, Coatue, J.P. Morgan, CapitalG, and Lead Edge Capital. Learn more at grafana.com and follow us on LinkedIn and X.
We're scaling fast and staying true to what makes us different: an open-source legacy, a global collaborative culture, and a passion for meaningful work. Our team thrives in an innovation-driven environment where transparency, autonomy, and trust fuel everything we do.
You may not meet every requirement, and that's okay. If this role excites you, we'd love you to raise your hand for what could be a truly career-defining opportunity.
This is a remote opportunity and we would be interested in applicants living in UK time zones only at this time.
Software Engineer - Platform Productivity
The Opportunity:
Grafana Cloud moves millions of metrics, log lines, and traces per second from our customers' environments into a highly available, low-latency stack that processes and stores these data, and serves them to dashboards and alerting tools. We aim to grow this to hundreds of millions per second, and it's critical that as we grow, we improve our performance, increase our reliability, and, of course, do it efficiently and effectively.
The Internal Engineering Platform delivered by the Platform department provides application engineers with the tools, systems and Kubernetes clusters they need to build, deploy and run their workloads. Platform roles at Grafana Labs have an eye for engineers with a passion for performance and reliability, and who enjoy taking projects from conception to production. We organize ourselves into squads to allow focus on Cloud Infrastructure, Networking and Security; engineering Productivity; Capacity management, Client Administrative Tooling (CAT); and US Federal compliance.
Because we deploy production services, we have on-call rotations to ensure the health of the system. Everyone at Grafana Labs tries to incorporate and use our product line up into their day-to-day, so being on call is an important way to understand our system and how people use our products.
What You'll Be Doing:
We are hiring for the Platform Productivity squad. The squad is mostly responsible for helping our internal engineers release their software onto our infrastructure, in secure and measurable ways. They lead automation of the release processes (anywhere from CI/CD to bootstrapping) and help our internal engineering teams get on board with them using 'golden path' techniques, but also helping edge cases and making sure that teams can get the most out of our tools.
At the end of the day, we're the Platform Team for the teams that are building some of the most cherished observability tools– from Grafana, Mimir and Loki, to Tempo.
The squad is responsible for setting its own roadmap, and as a part of the team you'll have a part to play in that process. You'll help us maintain, improve and extend what we already have. You'll be involved in choosing what we focus on next and, just as importantly, when and how to gracefully sunset systems which are no longer needed. Your responsibilities will also include helping the team to design, compare, and choose appropriate solutions for (at least some) of the following things:
• Development and maintenance of our Internal Engineering Platform (IEP)
• CI/CD platform management and development
• Build, release and deployment automation
• Application configuration management tooling
• "Up to date" software automation
• Artefact management
• Working with diverse internal teams, from application development to security, to support implementation of their requirements
• Being part of an on-call rotation to support Platform tooling
We invest heavily in developer productivity. You can use modern AI coding assistants as part of your daily workflow (your choice of tools, within security guidelines), backed by a company-funded usage budget so you can iterate quickly without unnecessary friction. We encourage pragmatic AI-assisted development: faster prototyping, test generation, refactors, documentation, and incident follow-ups—always with strong code review and quality standards. You'll also have access to the latest frontier models from OpenAI, Anthropi