Skip to content
Archive

Author: Daniel Brooks

Daniel designs the collection and normalization pipelines behind Heroku's datasets. He cares about honest error bars, boring-reliable infrastructure, and documentation that people actually read.

Building Reliable Public-Web Data Pipelines

Public-web pipelines fail in boring ways: a layout changes, an encoding breaks, a rate limit appears. Engineering for those failures — detectably, recoverably — is what separates a dataset from a science project.

Daniel Brooks

Scaling Collection Responsibly

Collection at scale is a privilege the open web extends to well-behaved participants. The engineering and governance practices that let a collection program grow without becoming the problem it studies.

Daniel Brooks