Large Repositories: LFS, Partial Clone and Worktrees

Why a clone gets slow, the four ways to take less of it, and choosing between submodules, subtrees and a monorepo without regretting it in a year.

advanced 20 min lesson hands-on task included

Git was designed for the Linux kernel: a large repository of text, with many contributors and a long history. It handles that beautifully. It handles a repository full of 200 MB binaries badly, and the difference between the two cases is worth understanding before your clone takes eleven minutes.


Topic 1: What Actually Makes a Clone Slow

WHEN A CLONE STOPS BEING FREE — FOUR WAYS TO TAKE LESS OF IT full clone every object, every commit the default shallow --depth=1 CI only — cannot push reliably partial --filter=blob:none history now, file cont ents on demand sparse sparse-checkout set DIR full history, part of the tree AND THREE WAYS TO SPLIT OR JOIN REPOSITORIES worktree two branches checked out at once, one object store no second clone, no stashing to swit ch submodule a pinned pointer to another repo exact version control, constant fric tion subtree the other repo's files merged in no extra commands for consumers, mes sier history LARGE BINARIES ARE A DIFFERENT PROBLEM — GIT LFS Every version of a 200 MB asset lives in every clone forever. LFS stores a pointer in Git and the bytes elsewhere. Decide before the first commit.
The top row is how much of the repository you take; the bottom row is how repositories are split or joined. Shallow clones are the popular choice and the one with the most caveats — full history is what makes blame and bisect possible.

Three independent causes, and they need different fixes:

  1. Many objects — a long history with many commits and files. Fixed with partial clone and a current commit-graph.
  2. Large objects — binaries stored in Git. Fixed with LFS, and only really fixed before they are committed.
  3. A large working tree — millions of files checked out. Fixed with sparse checkout.

Measure before choosing:

git count-objects -vH                # loose vs packed, total size

# The twenty largest objects in the entire history
git rev-list --objects --all |
  git cat-file --batch-check='%(objecttype) %(objectname) %(objectsize) %(rest)' |
  awk '$1=="blob" {print $3, $4}' | sort -rn | head -20 |
  numfmt --to=iec --field=1

That second pipeline is worth keeping. It usually reveals that 80% of a repository’s size comes from a handful of files somebody committed once — a database dump, a design asset, a vendored binary.


Topic 2: The Four Ways to Take Less

# Shallow — newest commit only
git clone --depth=1 <url>

# Partial — all commits, file contents fetched on demand
git clone --filter=blob:none <url>
git clone --filter=blob:limit=1m <url>     # only blobs under 1 MB up front

# Single branch
git clone --single-branch --branch main <url>

# Sparse — full history, part of the working tree
git clone --filter=blob:none --sparse <url>
cd repo
git sparse-checkout set services/api libs/common
Keeps historyCan blame/bisectCan push safelyGood for
--depth=1NoNoAwkwardCI, throwaway builds
--filter=blob:noneYesYes (fetches on demand)YesHumans on big repos
--single-branchYes, one branchYesYesFocused work
sparse-checkoutYesYesYesMonorepos

Shallow clones are the common choice and the most misused. --depth=1 is right for a CI job that builds once and discards the checkout. It is wrong for a working clone: no history means no blame, no bisect, no log, and pushing from a shallow clone has enough edge cases to be worth avoiding. For a human on a large repository, --filter=blob:none gives you everything a shallow clone gives, and history as well — blobs arrive when you actually touch a file.

Sparse checkout with cone mode is what makes a large monorepo usable:

git sparse-checkout init --cone
git sparse-checkout set services/api libs/common
git sparse-checkout list
git sparse-checkout disable      # back to everything

Cone mode restricts patterns to directory prefixes, which is much faster than arbitrary pattern matching and is what you want in practice.


Topic 3: Git LFS

Every version of every binary lives in every clone forever. LFS replaces the file in Git with a small pointer and stores the bytes on a separate server:

git lfs install
git lfs track "*.psd" "*.mp4" "*.zip"
git add .gitattributes            # tracking rules must be committed
git add design.psd
git commit -m "add design source"
*.psd filter=lfs diff=lfs merge=lfs -text

What to know before adopting it:

  • Decide before the first commit. Converting an existing repository means a history rewrite (git lfs migrate import) with all the coordination that implies.
  • Everyone needs the LFS client, or they get pointer files containing a hash instead of their assets.
  • It is a separate server and a separate quota, often separately priced, and it is another thing that can be down.
  • Not every forge or mirror supports it, which matters for open-source contributors and for disaster recovery.

The alternative worth considering first: do not put the binary in version control at all. An artifact registry, an object store with a version in the filename, or a package manager reference is frequently a better fit — and it removes the problem instead of relocating it.


Topic 4: Worktrees

A worktree is a second (or third) checkout backed by the same object store:

git worktree add ../repo-hotfix hotfix/token-expiry
git worktree add -b review/pr-482 ../repo-review origin/pr-482
git worktree list
git worktree remove ../repo-hotfix
git worktree prune            # clean up after a directory was deleted manually

Why this beats a second clone: one object store (no duplicated history on disk), one set of remotes and config, and instant creation. Why it beats stashing: you keep your working state completely untouched while you look at something else.

The uses that come up constantly:

  • An urgent hotfix while a long build or a messy refactor sits in your main checkout.
  • Reviewing a colleague’s branch while running it, without disturbing your own tree.
  • Building two versions simultaneously to compare behaviour.

Two constraints: the same branch cannot be checked out in two worktrees at once (Git refuses, which is a feature), and each worktree has its own index and HEAD — so git status in one says nothing about the other.


Topic 5: Submodules, Subtrees and Monorepos

Three answers to “this repository needs code from another one”.

Submodule — a pinned pointer to a specific commit of another repository:

git submodule add https://github.com/org/lib vendor/lib
git clone --recurse-submodules <url>
git submodule update --init --recursive
git submodule update --remote          # move pointers to the tracked branch

Precise version control, and constant friction: every clone needs an extra flag, a forgotten submodule update produces a confusing “modified: vendor/lib (new commits)” state, and detached HEAD inside a submodule is the default rather than an accident. Choose it when you genuinely need to pin an external dependency by commit and cannot use a package manager.

git config --global submodule.recurse true    # removes most of the daily pain

Subtree — the other repository’s files merged into a subdirectory:

git subtree add --prefix=vendor/lib https://github.com/org/lib main --squash
git subtree pull --prefix=vendor/lib https://github.com/org/lib main --squash
git subtree push --prefix=vendor/lib https://github.com/org/lib main

Consumers need no special commands — a plain clone gets everything. The cost is a messier history and awkward contribution back upstream.

Monorepo — one repository, many projects, no cross-repo mechanics at all. Atomic changes across projects, one version of everything, and a genuine tooling requirement: sparse checkout, partial clone, change-detecting CI, and a build system that understands the dependency graph. A monorepo without that investment is a repository where every change runs every test.

The honest guidance: prefer a package manager to a submodule. Prefer a monorepo to a web of submodules when the projects change together. Reach for subtrees when consumers must not need extra commands and you rarely push upstream.


Topic 6: Keeping a Repository Fast

git maintenance start        # background gc, commit-graph, prefetch, loose-object cleanup

That one command replaces most manual maintenance folklore, and the commit-graph in particular makes git log, git merge-base and history traversal dramatically faster on a large repository.

git gc                            # pack loose objects
git commit-graph write --reachable
git repack -Ad --write-bitmap-index    # bitmaps make clones and fetches much faster
git config feature.manyFiles true      # bundles the settings for huge working trees
git config core.fsmonitor true         # OS-level file watching — big status speedup

core.fsmonitor is the one to try first on a monorepo: git status on a two-million-file tree goes from seconds to milliseconds because Git stops walking the whole tree.

Try it yourself: run the largest-blobs pipeline on a repository that has grown mysteriously. The result is almost always a small number of files somebody committed once and deleted later — still in every clone, forever, because deletion adds a commit and removes nothing.

Common mistake: using --depth=1 for a developer’s working clone to make it faster. It works until the first git blame, the first git bisect, or the first time someone needs a branch that was not fetched — and unshallowing later (git fetch --unshallow) downloads everything anyway. Use --filter=blob:none, which is faster than a shallow clone in most real cases and keeps every capability.