Objects, Refs and the Commit Graph

The four object types, why a commit hash fingerprints all of history, and how every branch, tag and HEAD is just a file containing one hash.

beginner 20 min lesson hands-on task included

This is the lesson that makes the rest of Git predictable. Everything else — branching, merging, rebasing, recovery — is an operation on a graph of immutable objects with mutable pointers into it. Twenty minutes here saves hours later.


Topic 1: Four Object Types

EVERY OBJECT IS ADDRESSED BY THE HASH OF ITS OWN CONTENT refs/heads/main a file containing one hash HEAD ref: refs/heads/main commit a1b2c3 tree → a tree hash parent → previous commit author + committer + date message: add parser commit 9f8e7d tree → a tree hash parent → previous commit author + committer + date message: fix quoting parent …older commits tree (a directory listing) 100644 blob 5f1a2b… README.md 100755 blob 7c3d9e… deploy.sh 040000 tree 2e4f6a… src/ mode + type + hash + name. That is the whole format. blob 5f1a2b file CONTENT only blob 7c3d9e no name, no path WHY IT MATTERS Two identical files are ONE blob. Changing a file changes every hash above it. So a commit hash fingerprints the whole history behind it. Rewrite = new hashes. git cat-file -p HEAD · git cat-file -p HEAD^{tree} · git rev-parse HEAD — the three commands that make this concrete.
Follow the arrows: a ref names a commit, a commit names a tree, a tree names blobs and other trees. Nothing points forward — which is why a commit can never know what came after it.

Everything Git stores is one of four types, and every one is addressed by the hash of its own content:

Blob — file contents. Nothing else: no name, no path, no permissions, no timestamp. Two identical files anywhere in the repository, in any commit, are one blob.

Tree — a directory listing. Each entry is a mode, a type, a hash and a name:

100644 blob 5f1a2b…    README.md
100755 blob 7c3d9e…    deploy.sh
040000 tree 2e4f6a…    src/

That is the whole format. Filenames live in trees, which is why renaming a file creates no new blob — Git detects renames by comparing content, rather than recording them.

Commit — a tree hash, zero or more parent hashes, an author, a committer, timestamps and a message:

tree 2e4f6a8b...
parent 9f8e7d6c...
author  Alice <alice@example.com> 1755500000 +0000
committer Alice <alice@example.com> 1755500000 +0000

fix(auth): reject tokens issued before a password change

Tag (annotated) — a pointer to an object plus a tagger, date, message, and optionally a signature.

Read them yourself; this is not an abstraction:

git cat-file -p HEAD              # the commit object, in full
git cat-file -p HEAD^{tree}       # its tree
git cat-file -p 5f1a2b            # a blob — just the file content
git cat-file -t 5f1a2b            # what type is this hash?
git rev-parse HEAD                # the full hash of HEAD

Topic 2: Author vs Committer, and Other Fields That Matter Later

A commit has two identities and two timestamps, and the difference explains behaviour people find mysterious:

  • Author — who wrote the change, and when. Preserved through rebase and cherry-pick.
  • Committer — who created this commit object, and when. Changes every time the commit is rewritten.

So after a rebase, git log still shows the original author and the original date — while git log --format='%h %ci %cn' shows that the commit was actually created five minutes ago by you. This is why a rebased branch can look, in a default log, as though nothing happened, and why --since filters sometimes behave unexpectedly.

It is also the mechanism behind a fact worth internalising early: the author field is text you supply. Nothing verifies it. That is what signing exists to fix, and it gets its own lesson.


Topic 3: The Graph Only Points Backwards

Every commit knows its parents. No commit knows its children.

A ←── B ←── C ←── D        each arrow is "parent"

Almost every confusing Git behaviour follows from this one asymmetry:

  • git log walks backwards from a ref. It can only show you commits reachable from where you start.
  • A commit with nothing pointing at it is unreachable — invisible to log, still present in the object database until garbage collection.
  • ”Deleting” a branch deletes a pointer. The commits remain until they are both unreachable and old enough for gc to remove.
  • Changing a commit changes its hash, therefore its children’s hashes are wrong, therefore they must be rebuilt with new hashes too. This is why rewriting history rewrites everything after the change.

Parent counts tell you what a commit is:

ParentsKind
0The root commit
1An ordinary commit
2A merge
3+An octopus merge — rare, and usually a sign of automation

The first parent of a merge is the branch you were on; the second is the branch you merged in. That ordering matters when you revert a merge (-m 1) and when you read history with --first-parent, which follows the mainline and skips the internals of merged branches — the single most useful log flag on a busy repository.


Topic 4: Refs Are Files Containing a Hash

There is no database of branches. A branch is a file:

cat .git/refs/heads/main
# 9f8e7d6c5b4a39281706f5e4d3c2b1a09f8e7d6c

cat .git/HEAD
# ref: refs/heads/main        ← HEAD names a branch: "attached"
# 9f8e7d6c…                   ← HEAD names a commit directly: "detached"

The ref namespace, which is worth recognising in error messages:

refs/heads/*      local branches
refs/remotes/*    remote-tracking branches (origin/main) — a cache, not your branch
refs/tags/*       tags
refs/stash        the stash

Committing does two things: it writes a new commit object, then it updates the file that HEAD points at. That is all “advancing a branch” means, and it is why branching is instantaneous and free.

Naming a commit — every one of these resolves to a hash, and knowing the syntax removes a lot of copy-pasting:

HEAD              # current commit
HEAD~3            # three commits back, following first parents
HEAD^             # first parent
HEAD^2            # SECOND parent — only meaningful on a merge
main@{2.days.ago} # where main pointed two days ago (reflog)
v1.2.0^{commit}   # the commit an annotated tag points to
:/fix parser      # the most recent commit whose message matches

~ walks generations; ^ selects among parents. HEAD~2 is “two commits back”; HEAD^2 is “the other parent of this merge”. Confusing them is common and, on a merge commit, produces the wrong answer silently.


Topic 5: What Is Inside .git

.git/
├── HEAD              which branch you are on
├── config            this repository's config
├── index             the staging area (a binary file)
├── objects/          every blob, tree, commit and tag
│   ├── 9f/8e7d…      loose objects, by hash prefix
│   └── pack/         packfiles — compressed, delta-encoded
├── refs/
│   ├── heads/        branches
│   ├── remotes/      remote-tracking branches
│   └── tags/
├── logs/             the reflog — every ref movement
└── hooks/            scripts, never cloned

Two housekeeping realities:

Packfiles. Loose objects are one file each. git gc packs them into a packfile with delta compression between objects — so the storage layer does use deltas, even though the object model does not. Git runs gc --auto for you; on very large repositories git maintenance start schedules the work properly in the background instead.

Garbage collection removes unreachable objects, but not immediately: the default grace period keeps unreachable objects for two weeks and reflog entries for 90 days. That window is why almost every “lost” commit is recoverable — and why git gc --prune=now is a command to run deliberately, never reflexively.

git count-objects -vH      # loose vs packed, and total size
git fsck --lost-found      # unreachable objects, written to .git/lost-found

Topic 6: Why This Makes Recovery Boring

Once the model is clear, recovery stops being about memorised incantations:

“Disaster”What actually happenedThe fix
reset --hard lost my commitsA ref moved. The objects are untouched.Point a ref back at them (reflog)
I deleted a branchOne file in refs/heads/ was removedRecreate it at the same hash
A rebase went wrongNew commits were created; the old ones are unreferencedreset --hard ORIG_HEAD
Amended over a good commitThe old commit is still an objectFind it in the reflog, branch from it
A force-push overwrote mainThe remote’s ref moved; objects survive on the forge and in every clonePush the correct hash back

Every row is the same sentence: objects are immutable, refs move, and the reflog remembers where refs have been. The only genuinely unrecoverable loss is work that was never committed — because nothing recorded it.

Try it yourself: run git cat-file -p HEAD, copy the tree hash, cat-file -p that, copy a blob hash, cat-file -p that. Three commands from a commit to a file’s bytes, with no filename anywhere in the blob. That is the model, and after this you can predict what any command does to it.

Common mistake: believing git gc or a deleted branch destroys work. In normal use Git is close to append-only — it is very hard to lose a committed object within the grace period. Which cuts the other way too: a secret you committed is not removed by deleting the file, or the branch, or by a fresh commit. It is an object, it is reachable from history, and removing it takes a deliberate rewrite.