At the end of last year, our uptime was pretty shaky . You can see this trend on our status page , and that instability continued into the new year. Many of these outages were caused by a single bug, deep in SQLite . It took months of intense forensics to track it down. Now we’re in summer, we’re confident that we’ve found the bug, that we understand it—and more importantly, that we’ve fixed it. We know our customers expect Tailscale to be a reliable service, and for several months we didn’t live up to that promise. That’s disruptive, and we’re sorry. We’re publishing this blog post to explain what went wrong, how we responded, and how we ultimately helped to uncover a long-standing bug in the heart of the SQLite database. Tailscale’s database architecture While our clients interact with our control plane as a single public endpoint ( controlplane.tailscale.com ), internally, our control plane is split into a series of coordination servers (or “shards”).…