The Treacherous Tangle of Redundant Data: Resilience for Wallaroo

Introduction: we need data redundancy, but how, exactly? You now have your distributed system in production, congratulations! Your cluster is starting at six machines, but it is expected to grow quickly as it is assigned more work. The cluster’s main application is stateful, and that’s a problem. What if you lose a local disk drive? Or a sysadmin runs rm -rf on the wrong directory? Or else the entire machine cannot reboot, due to a power failure or administrator error that destroys an entire virtual machine?…

Keep reading

Checkpointing and Consistent Recovery Lines: How We Handle Failure in Wallaroo

In which we show you some of the key issues we considered when choosing how to handle failure in our system, and, in the process, introduce you to some concepts and resources that will help you in thinking about how to build resilient distributed systems of your own.

Keep reading

Utilizing Elixir as a lightweight tool to store real-time metrics data

How we use Elixir to store and aggregate Wallaroo’s metrics for end-user consumption.

Keep reading

Dynamic Keys

A look at how Wallaroo applications can now support new keys

Keep reading

Adventures with cgo: Part 2- Locks and other things that go bump in the night

We’ve learned quite a lot while working on the Go API for Wallaroo. Come along for the journey with us as we teach you about the fun and foibles that await when you go adventuring with cgo. In part 2, we follow up on some lessons learned in part 1.

Keep reading