Git at Any Scale
Git was designed by Linus Torvalds for his own use case: replacing BitKeeper for Linux kernel development, where a distributed workflow matched the project's decentralized maintainer structure. Twenty years later, Git is an industry standard, but most projects and companies actually rely on a centralized host, making its distributed nature more a hindrance than an advantage.
The fundamental challenge is that Git is inherently distributed: every repository instance is identical, so a server-side repository is not inherently special. While one could simply put an HTTP daemon in front of an on-disk copy, scalability and reliability issues arise from how Git stores data. Code and metadata are compressed into packfiles, which are binary serialization formats used both for storage and for network transfer of pushes and fetches.
This packfile-based design imposes a hard limitation on hosting at scale. Packfiles are large binary files that must exist on a filesystem for Git to access them, and the simple approach of serving an on-disk repo via HTTP has a very low ceiling. The article argues that servers are not tied to packfiles internally—only for network operations—but companies hosting Git at scale have found this design to be a major availability and scalability constraint.