Why Version Control, and Why Git Won

What a distributed model buys you that a centralized one cannot, why Git stores snapshots rather than deltas, and the properties that make branching cheap enough to change how teams work.

beginner 16 min lesson hands-on task included

Every team has a version control system. The ones without Git have a folder called final_v2_ACTUAL_final, and a person who knows which file is the real one.

Git is not the first version control system, and its interface is famously inconsistent. It won because of two design decisions that changed what was practical, and both of them are still visible in everything the tool does.


Topic 1: Centralized vs Distributed

CENTRALIZED (SVN, CVS, TFVC) the server dev 1 working copy only dev 2 working copy only dev 3 working copy only Every commit, log and diff is a network round trip. Server down = nobody commits. History lives in one place. DISTRIBUTED (Git) origin a convention, not an authority dev 1 FULL history dev 2 FULL history dev 3 FULL history Commit, branch, log, diff, blame, bisect — all local. Every clone is a complete backup of the whole history. THE CONSEQUENCE PEOPLE ACTUALLY FEEL Because branching is a 41-byte file rather than a server-side copy, Git made branches cheap — and cheap branches changed how teams work. Git also stores each commit as a full snapshot (deduplicated by content), not as a chain of deltas. That is why checkout is fast at any point in history.
The topology is the whole difference. In a centralized system the server holds the history and clients hold a working copy; in Git every clone holds the complete history, and 'origin' is a convention rather than an authority.

Centralized (SVN, CVS, Perforce, TFVC): one server holds the history. Your machine holds a working copy at one revision. Committing, viewing a log, or diffing against an old version is a network operation.

Distributed (Git, Mercurial): every clone contains the complete history. Committing, branching, logging, diffing, blaming and bisecting are local operations against a local database.

What actually falls out of that:

CentralizedGit
Commit while offlineNoYes
Speed of log, diff, blameNetwork round tripLocal, effectively instant
Server unavailableNobody commitsWork continues; you push later
BackupsThe server is the single copyEvery clone is a full backup
Branching costA server-side copy, sometimes expensiveA 41-byte file
Access controlPer directory, built inAll-or-nothing per repo; the forge adds the rest

That last row is a genuine Git weakness, not a talking point to dodge: Git has no per-directory permissions. If some files must be visible to some people only, the answer is a separate repository, not a clever Git configuration.

”Distributed” does not mean “no central server.” Almost every team designates one repository as canonical — GitHub, GitLab, Bitbucket, a bare repo on a box. The difference is that this is a social decision. Git itself does not know that origin is special, which is why a repository can be re-hosted in an afternoon and why a forge outage does not stop work.


Topic 2: Snapshots, Not Deltas

Most older systems store a file as an original plus a chain of diffs. Reconstructing an old version means replaying the chain.

Git stores a complete snapshot of the tree at every commit. A file that did not change between two commits is not copied — both commits reference the same object, because objects are addressed by the hash of their content. Identical content is stored exactly once, automatically, across the entire repository’s history.

Two consequences you will feel:

  • Checkout is fast at any point in history. There is no chain to replay. Checking out a commit from 2015 costs the same as checking out yesterday’s.
  • The .git directory is usually smaller than people expect — often smaller than the working tree — because deduplication plus zlib compression plus packfiles (which do use deltas, as a storage optimisation, on top of the object model) work well on text.

Where it goes badly: binary files. A 200 MB video that changes ten times is ten distinct objects, stored forever, in every clone anybody ever makes. This is the problem Git LFS exists to solve, and it is far cheaper to solve before the first commit than after.


Topic 3: Content Addressing and Integrity

Every object in Git is named by a cryptographic hash of its own contents. Historically SHA-1; the SHA-256 object format now exists, and the transition is gradual because a repository’s hash format is baked into every reference to it.

Because a commit’s hash covers its tree and its parent’s hash, a commit hash fingerprints the entire history behind it. Change anything — a byte in a file, an author’s email, a date, a commit message from 2019 — and every hash from that point forward changes.

That gives you three properties for free:

  1. Corruption is detectable. A damaged object no longer hashes to its own name. git fsck finds it.
  2. History is tamper-evident. You cannot quietly edit an old commit and leave the newer ones intact.
  3. A commit hash is a precise, permanent name. “Deploy 9f8e7d” is unambiguous in a way “deploy version 1.4” never is.

The SHA-1 collision attacks published from 2017 onward are worth knowing about and are not a practical break of Git: Git added collision detection to its SHA-1 implementation, and a forged object still has to pass review and signing. The SHA-256 work exists for the long term, not because your repository is at risk today.


Topic 4: Why Cheap Branching Changed Team Behaviour

In a system where a branch is a server-side copy, branching is a decision you make a few times a year. In Git a branch is a file containing one hash, so branching is a decision you make several times a day.

That single cost change is the reason modern workflows exist at all: feature branches, pull request review, short-lived topic branches, and “just try it on a branch” as an unremarkable habit. Every strategy in this module — trunk-based development, GitHub flow, GitFlow — is a different answer to a question that only became askable once branches became free.

The corollary, which people discover later: cheap branches also make it easy to accumulate forty stale branches nobody dares delete, and to keep a branch alive for three months until merging it becomes an event. Cheap to create is not cheap to keep.


Topic 5: What You Are Actually Learning

Git’s command surface is large and, in places, inconsistent — checkout used to do four unrelated jobs, which is exactly why switch and restore were added. Trying to memorise commands produces someone who is fine until something goes wrong.

The alternative, which this module follows, is to learn the model first:

Objects       blobs, trees, commits, tags — content-addressed, immutable
Refs          branches, tags, HEAD — mutable pointers into that graph
The index     the proposed next commit, sitting between your files and history
The reflog    a local record of where every ref has been

Four ideas. Every command is an operation on them, and every recovery is a matter of pointing a ref back at an object that still exists. When you can predict what a command does to those four things, the inconsistent naming stops mattering.

A worked example of that payoff: “I ran git reset --hard and lost two days of work” is a crisis if you memorised commands, and a two-minute fix if you know that reset moves a ref, that the commits are still objects in the database, and that the reflog recorded where the ref used to point.


Topic 6: Setting Up So the Rest of the Module Works

git config --global user.name "Your Name"
git config --global user.email "you@example.com"

# Which branch name new repositories get. 'main' is the current convention.
git config --global init.defaultBranch main

# Never guess how to reconcile divergent branches on pull — decide once.
git config --global pull.rebase true

# Refuse to push if it would overwrite work you have not seen.
git config --global push.default simple

# Remember conflict resolutions and reapply them automatically.
git config --global rerere.enabled true

# Better diffs: highlight the moved and changed parts, not just the lines.
git config --global diff.algorithm histogram
git config --global diff.colorMoved zebra

Two of those deserve a note now and a full treatment later. pull.rebase true avoids the “Merge branch ‘main’ of github.com:…” commits that clutter a busy repository’s history — covered properly in the rebase lesson. rerere.enabled records how you resolved a conflict and replays that resolution the next time the same conflict appears, which turns a long-running branch’s repeated conflicts from a daily tax into a one-time cost.

Try it yourself: run git count-objects -vH in a repository with real history, then git gc and run it again. Watching loose objects collapse into a packfile — often an order of magnitude smaller — is the storage model made visible.

Common mistake: treating Git as a backup tool and stopping at add, commit, push. Those three commands work until the first time something goes wrong, and then the model you skipped is the only thing that helps. The next lesson is the smallest possible version of that model: three places a change can be.