The Treacherous Tangle of Redundant Data: Resilience for Wallaroo

Introduction: we need data redundancy, but how, exactly? You now have your distributed system in production, congratulations! Your cluster is starting at six machines, but it is expected to grow quickly as it is assigned more work. The cluster’s main application is stateful, and that’s a problem. What if you lose a local disk drive? Or a sysadmin runs rm -rf on the wrong directory? Or else the entire machine cannot reboot, due to a power failure or administrator error that destroys an entire virtual machine?…

Keep reading

Checkpointing and Consistent Recovery Lines: How We Handle Failure in Wallaroo

In which we show you some of the key issues we considered when choosing how to handle failure in our system, and, in the process, introduce you to some concepts and resources that will help you in thinking about how to build resilient distributed systems of your own.

Keep reading

Event Triggered Customer Segmentation

In which we quickly and easily create a full application to use Wallaroo to manage an ad campaign for a Marsupial Fan Club.

Keep reading

Stream processing, trending hashtags, and Wallaroo

See how you can chain Wallaroo state computations together for build a Twitter trending hashtags application.

Keep reading

Exploring The GitHub Archive

Learn to use Wallaroo with Python, Kafka, and the GitHub Archive

Keep reading