This is the lesson that makes the rest of Git predictable. Everything else — branching, merging, rebasing, recovery — is an operation on a graph of immutable objects with mutable pointers into it. Twenty minutes here saves hours later.
Topic 1: Four Object Types
Everything Git stores is one of four types, and every one is addressed by the hash of its own content:
Blob — file contents. Nothing else: no name, no path, no permissions, no timestamp. Two identical files anywhere in the repository, in any commit, are one blob.
Tree — a directory listing. Each entry is a mode, a type, a hash and a name:
100644 blob 5f1a2b… README.md
100755 blob 7c3d9e… deploy.sh
040000 tree 2e4f6a… src/
That is the whole format. Filenames live in trees, which is why renaming a file creates no new blob — Git detects renames by comparing content, rather than recording them.
Commit — a tree hash, zero or more parent hashes, an author, a committer, timestamps and a message:
tree 2e4f6a8b...
parent 9f8e7d6c...
author Alice <alice@example.com> 1755500000 +0000
committer Alice <alice@example.com> 1755500000 +0000
fix(auth): reject tokens issued before a password change
Tag (annotated) — a pointer to an object plus a tagger, date, message, and optionally a signature.
Read them yourself; this is not an abstraction:
git cat-file -p HEAD # the commit object, in full
git cat-file -p HEAD^{tree} # its tree
git cat-file -p 5f1a2b # a blob — just the file content
git cat-file -t 5f1a2b # what type is this hash?
git rev-parse HEAD # the full hash of HEAD
Topic 2: Author vs Committer, and Other Fields That Matter Later
A commit has two identities and two timestamps, and the difference explains behaviour people find mysterious:
- Author — who wrote the change, and when. Preserved through rebase and cherry-pick.
- Committer — who created this commit object, and when. Changes every time the commit is rewritten.
So after a rebase, git log still shows the original author and the original date — while git log --format='%h %ci %cn' shows that the commit was actually created five minutes ago by you. This is why a rebased branch can look, in a default log, as though nothing happened, and why --since filters sometimes behave unexpectedly.
It is also the mechanism behind a fact worth internalising early: the author field is text you supply. Nothing verifies it. That is what signing exists to fix, and it gets its own lesson.
Topic 3: The Graph Only Points Backwards
Every commit knows its parents. No commit knows its children.
A ←── B ←── C ←── D each arrow is "parent"
Almost every confusing Git behaviour follows from this one asymmetry:
git logwalks backwards from a ref. It can only show you commits reachable from where you start.- A commit with nothing pointing at it is unreachable — invisible to
log, still present in the object database until garbage collection. - ”Deleting” a branch deletes a pointer. The commits remain until they are both unreachable and old enough for
gcto remove. - Changing a commit changes its hash, therefore its children’s hashes are wrong, therefore they must be rebuilt with new hashes too. This is why rewriting history rewrites everything after the change.
Parent counts tell you what a commit is:
| Parents | Kind |
|---|---|
| 0 | The root commit |
| 1 | An ordinary commit |
| 2 | A merge |
| 3+ | An octopus merge — rare, and usually a sign of automation |
The first parent of a merge is the branch you were on; the second is the branch you merged in. That ordering matters when you revert a merge (-m 1) and when you read history with --first-parent, which follows the mainline and skips the internals of merged branches — the single most useful log flag on a busy repository.
Topic 4: Refs Are Files Containing a Hash
There is no database of branches. A branch is a file:
cat .git/refs/heads/main
# 9f8e7d6c5b4a39281706f5e4d3c2b1a09f8e7d6c
cat .git/HEAD
# ref: refs/heads/main ← HEAD names a branch: "attached"
# 9f8e7d6c… ← HEAD names a commit directly: "detached"
The ref namespace, which is worth recognising in error messages:
refs/heads/* local branches
refs/remotes/* remote-tracking branches (origin/main) — a cache, not your branch
refs/tags/* tags
refs/stash the stash
Committing does two things: it writes a new commit object, then it updates the file that HEAD points at. That is all “advancing a branch” means, and it is why branching is instantaneous and free.
Naming a commit — every one of these resolves to a hash, and knowing the syntax removes a lot of copy-pasting:
HEAD # current commit
HEAD~3 # three commits back, following first parents
HEAD^ # first parent
HEAD^2 # SECOND parent — only meaningful on a merge
main@{2.days.ago} # where main pointed two days ago (reflog)
v1.2.0^{commit} # the commit an annotated tag points to
:/fix parser # the most recent commit whose message matches
~ walks generations; ^ selects among parents. HEAD~2 is “two commits back”; HEAD^2 is “the other parent of this merge”. Confusing them is common and, on a merge commit, produces the wrong answer silently.
Topic 5: What Is Inside .git
.git/
├── HEAD which branch you are on
├── config this repository's config
├── index the staging area (a binary file)
├── objects/ every blob, tree, commit and tag
│ ├── 9f/8e7d… loose objects, by hash prefix
│ └── pack/ packfiles — compressed, delta-encoded
├── refs/
│ ├── heads/ branches
│ ├── remotes/ remote-tracking branches
│ └── tags/
├── logs/ the reflog — every ref movement
└── hooks/ scripts, never cloned
Two housekeeping realities:
Packfiles. Loose objects are one file each. git gc packs them into a packfile with delta compression between objects — so the storage layer does use deltas, even though the object model does not. Git runs gc --auto for you; on very large repositories git maintenance start schedules the work properly in the background instead.
Garbage collection removes unreachable objects, but not immediately: the default grace period keeps unreachable objects for two weeks and reflog entries for 90 days. That window is why almost every “lost” commit is recoverable — and why git gc --prune=now is a command to run deliberately, never reflexively.
git count-objects -vH # loose vs packed, and total size
git fsck --lost-found # unreachable objects, written to .git/lost-found
Topic 6: Why This Makes Recovery Boring
Once the model is clear, recovery stops being about memorised incantations:
| “Disaster” | What actually happened | The fix |
|---|---|---|
reset --hard lost my commits | A ref moved. The objects are untouched. | Point a ref back at them (reflog) |
| I deleted a branch | One file in refs/heads/ was removed | Recreate it at the same hash |
| A rebase went wrong | New commits were created; the old ones are unreferenced | reset --hard ORIG_HEAD |
| Amended over a good commit | The old commit is still an object | Find it in the reflog, branch from it |
| A force-push overwrote main | The remote’s ref moved; objects survive on the forge and in every clone | Push the correct hash back |
Every row is the same sentence: objects are immutable, refs move, and the reflog remembers where refs have been. The only genuinely unrecoverable loss is work that was never committed — because nothing recorded it.
Try it yourself: run git cat-file -p HEAD, copy the tree hash, cat-file -p that, copy a blob hash, cat-file -p that. Three commands from a commit to a file’s bytes, with no filename anywhere in the blob. That is the model, and after this you can predict what any command does to it.
Common mistake: believing git gc or a deleted branch destroys work. In normal use Git is close to append-only — it is very hard to lose a committed object within the grace period. Which cuts the other way too: a secret you committed is not removed by deleting the file, or the branch, or by a fresh commit. It is an object, it is reachable from history, and removing it takes a deliberate rewrite.