Raw
1 gitformat-pack(5)
2 =================
3
4 NAME
5 ----
6 gitformat-pack - Git pack format
7
8
9 SYNOPSIS
10 --------
11 [verse]
12 $GIT_DIR/objects/pack/pack-*.{pack,idx}
13 $GIT_DIR/objects/pack/pack-*.rev
14 $GIT_DIR/objects/pack/pack-*.mtimes
15 $GIT_DIR/objects/pack/multi-pack-index
16
17 DESCRIPTION
18 -----------
19
20 The Git pack format is how Git stores most of its primary repository
21 data. Over the lifetime of a repository, loose objects (if any) and
22 smaller packs are consolidated into larger pack(s). See
23 linkgit:git-gc[1] and linkgit:git-pack-objects[1].
24
25 The pack format is also used over-the-wire, see
26 e.g. linkgit:gitprotocol-v2[5], as well as being a part of
27 other container formats in the case of linkgit:gitformat-bundle[5].
28
29 == Checksums and object IDs
30
31 In a repository using the traditional SHA-1, pack checksums, index checksums,
32 and object IDs (object names) mentioned below are all computed using SHA-1.
33 Similarly, in SHA-256 repositories, these values are computed using SHA-256.
34
35 CRC32 checksums are always computed over the entire packed object, including
36 the header (n-byte type and length); the base object name or offset, if any;
37 and the entire compressed object. The CRC32 algorithm used is that of zlib.
38
39 == pack-*.pack files have the following format:
40
41 - A header appears at the beginning and consists of the following:
42
43 4-byte signature:
44 The signature is: {'P', 'A', 'C', 'K'}
45
46 4-byte version number (network byte order):
47 Git currently accepts version number 2 or 3 but
48 generates version 2 only.
49
50 4-byte number of objects contained in the pack (network byte order)
51
52 Observation: we cannot have more than 4G versions ;-) and
53 more than 4G objects in a pack.
54
55 - The header is followed by a number of object entries, each of
56 which looks like this:
57
58 (undeltified representation)
59 n-byte type and length (3-bit type, (n-1)*7+4-bit length)
60 compressed data
61
62 (deltified representation)
63 n-byte type and length (3-bit type, (n-1)*7+4-bit length)
64 base object name if OBJ_REF_DELTA or a negative relative
65 offset from the delta object's position in the pack if this
66 is an OBJ_OFS_DELTA object
67 compressed delta data
68
69 Observation: the length of each object is encoded in a variable
70 length format and is not constrained to 32-bit or anything.
71
72 - The trailer records a pack checksum of all of the above.
73
74 === Object types
75
76 Valid object types are:
77
78 - OBJ_COMMIT (1)
79 - OBJ_TREE (2)
80 - OBJ_BLOB (3)
81 - OBJ_TAG (4)
82 - OBJ_OFS_DELTA (6)
83 - OBJ_REF_DELTA (7)
84
85 Type 5 is reserved for future expansion. Type 0 is invalid.
86
87 === Object encoding
88
89 Unlike loose objects, packed objects do not have a prefix containing the type,
90 size, and a NUL byte. These are not necessary because they can be determined by
91 the n-byte type and length that prefixes the data and so they are omitted from
92 the compressed and deltified data.
93
94 The computation of the object ID still uses this prefix by reconstructing it
95 from the type and length as needed.
96
97 === Size encoding
98
99 This document uses the following "size encoding" of non-negative
100 integers: From each byte, the seven least significant bits are
101 used to form the resulting integer. As long as the most significant
102 bit is 1, this process continues; the byte with MSB 0 provides the
103 last seven bits. The seven-bit chunks are concatenated. Later
104 values are more significant.
105
106 This size encoding should not be confused with the "offset encoding",
107 which is also used in this document.
108
109 When encoding the size of an undeltified object in a pack, the size is that of
110 the uncompressed raw object. For deltified objects, it is the size of the
111 uncompressed delta. The base object name or offset is not included in the size
112 computation.
113
114 === Deltified representation
115
116 Conceptually there are only four object types: commit, tree, tag and
117 blob. However to save space, an object could be stored as a "delta" of
118 another "base" object. These representations are assigned new types
119 ofs-delta and ref-delta, which is only valid in a pack file.
120
121 Both ofs-delta and ref-delta store the "delta" to be applied to
122 another object (called 'base object') to reconstruct the object. The
123 difference between them is, ref-delta directly encodes base object
124 name. If the base object is in the same pack, ofs-delta encodes
125 the offset of the base object in the pack instead.
126
127 The base object could also be deltified if it's in the same pack.
128 Ref-delta can also refer to an object outside the pack (i.e. the
129 so-called "thin pack"). When stored on disk however, the pack should
130 be self contained to avoid cyclic dependency.
131
132 The delta data starts with the size of the base object and the
133 size of the object to be reconstructed. These sizes are
134 encoded using the size encoding from above. The remainder of
135 the delta data is a sequence of instructions to reconstruct the object
136 from the base object. If the base object is deltified, it must be
137 converted to canonical form first. Each instruction appends more and
138 more data to the target object until it's complete. There are two
139 supported instructions so far: one for copying a byte range from the
140 source object and one for inserting new data embedded in the
141 instruction itself.
142
143 Each instruction has variable length. Instruction type is determined
144 by the seventh bit of the first octet. The following diagrams follow
145 the convention in RFC 1951 (Deflate compressed data format).
146
147 ==== Instruction to copy from base object
148
149 +----------+---------+---------+---------+---------+-------+-------+-------+
150 | 1xxxxxxx | offset1 | offset2 | offset3 | offset4 | size1 | size2 | size3 |
151 +----------+---------+---------+---------+---------+-------+-------+-------+
152
153 This is the instruction format to copy a byte range from the source
154 object. It encodes the offset to copy from and the number of bytes to
155 copy. Offset and size are in little-endian order.
156
157 All offset and size bytes are optional. This is to reduce the
158 instruction size when encoding small offsets or sizes. The first seven
159 bits in the first octet determine which of the next seven octets is
160 present. If bit zero is set, offset1 is present. If bit one is set
161 offset2 is present and so on.
162
163 Note that a more compact instruction does not change offset and size
164 encoding. For example, if only offset2 is omitted like below, offset3
165 still contains bits 16-23. It does not become offset2 and contains
166 bits 8-15 even if it's right next to offset1.
167
168 +----------+---------+---------+
169 | 10000101 | offset1 | offset3 |
170 +----------+---------+---------+
171
172 In its most compact form, this instruction only takes up one byte
173 (0x80) with both offset and size omitted, which will have default
174 values zero. There is another exception: size zero is automatically
175 converted to 0x10000.
176
177 ==== Instruction to add new data
178
179 +----------+============+
180 | 0xxxxxxx | data |
181 +----------+============+
182
183 This is the instruction to construct the target object without the base
184 object. The following data is appended to the target object. The first
185 seven bits of the first octet determine the size of data in
186 bytes. The size must be non-zero.
187
188 ==== Reserved instruction
189
190 +----------+============
191 | 00000000 |
192 +----------+============
193
194 This is the instruction reserved for future expansion.
195
196 == Original (version 1) pack-*.idx files have the following format:
197
198 - The header consists of 256 4-byte network byte order
199 integers. N-th entry of this table records the number of
200 objects in the corresponding pack, the first byte of whose
201 object name is less than or equal to N. This is called the
202 'first-level fan-out' table.
203
204 - The header is followed by sorted 24-byte entries, one entry
205 per object in the pack. Each entry is:
206
207 4-byte network byte order integer, recording where the
208 object is stored in the packfile as the offset from the
209 beginning.
210
211 one object name of the appropriate size.
212
213 - The file is concluded with a trailer:
214
215 A copy of the pack checksum at the end of the corresponding
216 packfile.
217
218 Index checksum of all of the above.
219
220 Pack Idx file:
221
222 -- +--------------------------------+
223 fanout | fanout[0] = 2 (for example) |-.
224 table +--------------------------------+ |
225 | fanout[1] | |
226 +--------------------------------+ |
227 | fanout[2] | |
228 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ |
229 | fanout[255] = total objects |---.
230 -- +--------------------------------+ | |
231 main | offset | | |
232 index | object name 00XXXXXXXXXXXXXXXX | | |
233 table +--------------------------------+ | |
234 | offset | | |
235 | object name 00XXXXXXXXXXXXXXXX | | |
236 +--------------------------------+<+ |
237 .-| offset | |
238 | | object name 01XXXXXXXXXXXXXXXX | |
239 | +--------------------------------+ |
240 | | offset | |
241 | | object name 01XXXXXXXXXXXXXXXX | |
242 | ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ |
243 | | offset | |
244 | | object name FFXXXXXXXXXXXXXXXX | |
245 --| +--------------------------------+<--+
246 trailer | | packfile checksum |
247 | +--------------------------------+
248 | | idxfile checksum |
249 | +--------------------------------+
250 .-------.
251 |
252 Pack file entry: <+
253
254 packed object header:
255 1-byte size extension bit (MSB)
256 type (next 3 bit)
257 size0 (lower 4-bit)
258 n-byte sizeN (as long as MSB is set, each 7-bit)
259 size0..sizeN form 4+7+7+..+7 bit integer, size0
260 is the least significant part, and sizeN is the
261 most significant part.
262 packed object data:
263 If it is not DELTA, then deflated bytes (the size above
264 is the size before compression).
265 If it is REF_DELTA, then
266 base object name (the size above is the
267 size of the delta data that follows).
268 delta data, deflated.
269 If it is OFS_DELTA, then
270 n-byte offset (see below) interpreted as a negative
271 offset from the type-byte of the header of the
272 ofs-delta entry (the size above is the size of
273 the delta data that follows).
274 delta data, deflated.
275
276 offset encoding:
277 n bytes with MSB set in all but the last one.
278 The offset is then the number constructed by
279 concatenating the lower 7 bit of each byte, and
280 for n >= 2 adding 2^7 + 2^14 + ... + 2^(7*(n-1))
281 to the result.
282
283
284
285 == Version 2 pack-*.idx files support packs larger than 4 GiB, and
286 have some other reorganizations. They have the format:
287
288 - A 4-byte magic number '\377tOc' which is an unreasonable
289 fanout[0] value.
290
291 - A 4-byte version number (= 2)
292
293 - A 256-entry fan-out table just like v1.
294
295 - A table of sorted object names. These are packed together
296 without offset values to reduce the cache footprint of the
297 binary search for a specific object name.
298
299 - A table of 4-byte CRC32 values of the packed object data.
300 This is new in v2 so compressed data can be copied directly
301 from pack to pack during repacking without undetected
302 data corruption.
303
304 - A table of 4-byte offset values (in network byte order).
305 These are usually 31-bit pack file offsets, but large
306 offsets are encoded as an index into the next table with
307 the msbit set.
308
309 - A table of 8-byte offset entries (empty for pack files less
310 than 2 GiB). Pack files are organized with heavily used
311 objects toward the front, so most object references should
312 not need to refer to this table.
313
314 - The same trailer as a v1 pack file:
315
316 A copy of the pack checksum at the end of the
317 corresponding packfile.
318
319 Index checksum of all of the above.
320
321 == pack-*.rev files have the format:
322
323 - A 4-byte magic number '0x52494458' ('RIDX').
324
325 - A 4-byte version identifier (= 1).
326
327 - A 4-byte hash function identifier (= 1 for SHA-1, 2 for SHA-256).
328
329 - A table of index positions (one per packed object, num_objects in
330 total, each a 4-byte unsigned integer in network order), sorted by
331 their corresponding offsets in the packfile.
332
333 - A trailer, containing a:
334
335 checksum of the corresponding packfile, and
336
337 a checksum of all of the above.
338
339 All 4-byte numbers are in network order.
340
341 == pack-*.mtimes files have the format:
342
343 All 4-byte numbers are in network byte order.
344
345 - A 4-byte magic number '0x4d544d45' ('MTME').
346
347 - A 4-byte version identifier (= 1).
348
349 - A 4-byte hash function identifier (= 1 for SHA-1, 2 for SHA-256).
350
351 - A table of 4-byte unsigned integers. The ith value is the
352 modification time (mtime) of the ith object in the corresponding
353 pack by lexicographic (index) order. The mtimes count standard
354 epoch seconds.
355
356 - A trailer, containing a checksum of the corresponding packfile,
357 and a checksum of all of the above (each having length according
358 to the specified hash function).
359
360 == multi-pack-index (MIDX) files have the following format:
361
362 The multi-pack-index files refer to multiple pack-files and loose objects.
363
364 In order to allow extensions that add extra data to the MIDX, we organize
365 the body into "chunks" and provide a lookup table at the beginning of the
366 body. The header includes certain length values, such as the number of packs,
367 the number of base MIDX files, hash lengths and types.
368
369 All 4-byte numbers are in network order.
370
371 HEADER:
372
373 4-byte signature:
374 The signature is: {'M', 'I', 'D', 'X'}
375
376 1-byte version number:
377 Git writes the version specified by the "midx.version"
378 configuration option, which defaults to 2. It recognizes
379 both versions 1 and 2.
380
381 1-byte Object Id Version
382 We infer the length of object IDs (OIDs) from this value:
383 1 => SHA-1
384 2 => SHA-256
385 If the hash type does not match the repository's hash algorithm,
386 the multi-pack-index file should be ignored with a warning
387 presented to the user.
388
389 1-byte number of "chunks"
390
391 1-byte number of base multi-pack-index files:
392 This value is currently always zero.
393
394 4-byte number of pack files
395
396 CHUNK LOOKUP:
397
398 (C + 1) * 12 bytes providing the chunk offsets:
399 First 4 bytes describe chunk id. Value 0 is a terminating label.
400 Other 8 bytes provide offset in current file for chunk to start.
401 (Chunks are provided in file-order, so you can infer the length
402 using the next chunk position if necessary.)
403
404 The CHUNK LOOKUP matches the table of contents from
405 the chunk-based file format, see linkgit:gitformat-chunk[5].
406
407 The remaining data in the body is described one chunk at a time, and
408 these chunks may be given in any order. Chunks are required unless
409 otherwise specified.
410
411 CHUNK DATA:
412
413 Packfile Names (ID: {'P', 'N', 'A', 'M'})
414 Store the names of packfiles as a sequence of NUL-terminated
415 strings. There is no extra padding between the filenames,
416 and they are listed in lexicographic order. The chunk itself
417 is padded at the end with between 0 and 3 NUL bytes to make the
418 chunk size a multiple of 4 bytes. Version 1 MIDXs are required to
419 list their packs in lexicographic order, but version 2 MIDXs may
420 list their packs in any arbitrary order.
421
422 Bitmapped Packfiles (ID: {'B', 'T', 'M', 'P'})
423 Stores a table of two 4-byte unsigned integers in network order.
424 Each table entry corresponds to a single pack (in the order that
425 they appear above in the `PNAM` chunk). The values for each table
426 entry are as follows:
427 - The first bit position (in pseudo-pack order, see below) to
428 contain an object from that pack.
429 - The number of bits whose objects are selected from that pack.
430
431 OID Fanout (ID: {'O', 'I', 'D', 'F'})
432 The ith entry, F[i], stores the number of OIDs with first
433 byte at most i. Thus F[255] stores the total
434 number of objects.
435
436 OID Lookup (ID: {'O', 'I', 'D', 'L'})
437 The OIDs for all objects in the MIDX are stored in lexicographic
438 order in this chunk.
439
440 Object Offsets (ID: {'O', 'O', 'F', 'F'})
441 Stores two 4-byte values for every object.
442 1: The pack-int-id for the pack storing this object.
443 2: The offset within the pack.
444 If all offsets are less than 2^32, then the large offset chunk
445 will not exist and offsets are stored as in IDX v1.
446 If there is at least one offset value larger than 2^32-1, then
447 the large offset chunk must exist, and offsets larger than
448 2^31-1 must be stored in it instead. If the large offset chunk
449 exists and the 31st bit is on, then removing that bit reveals
450 the row in the large offsets containing the 8-byte offset of
451 this object.
452
453 [Optional] Object Large Offsets (ID: {'L', 'O', 'F', 'F'})
454 8-byte offsets into large packfiles.
455
456 [Optional] Bitmap pack order (ID: {'R', 'I', 'D', 'X'})
457 A list of MIDX positions (one per object in the MIDX, num_objects in
458 total, each a 4-byte unsigned integer in network byte order), sorted
459 according to their relative bitmap/pseudo-pack positions.
460
461 TRAILER:
462
463 Index checksum of the above contents.
464
465 == multi-pack-index reverse indexes
466
467 Similar to the pack-based reverse index, the multi-pack index can also
468 be used to generate a reverse index.
469
470 Instead of mapping between offset, pack-, and index position, this
471 reverse index maps between an object's position within the MIDX, and
472 that object's position within a pseudo-pack that the MIDX describes
473 (i.e., the ith entry of the multi-pack reverse index holds the MIDX
474 position of ith object in pseudo-pack order).
475
476 To clarify the difference between these orderings, consider a multi-pack
477 reachability bitmap (which does not yet exist, but is what we are
478 building towards here). Each bit needs to correspond to an object in the
479 MIDX, and so we need an efficient mapping from bit position to MIDX
480 position.
481
482 One solution is to let bits occupy the same position in the oid-sorted
483 index stored by the MIDX. But because oids are effectively random, their
484 resulting reachability bitmaps would have no locality, and thus compress
485 poorly. (This is the reason that single-pack bitmaps use the pack
486 ordering, and not the .idx ordering, for the same purpose.)
487
488 So we'd like to define an ordering for the whole MIDX based around
489 pack ordering, which has far better locality (and thus compresses more
490 efficiently). We can think of a pseudo-pack created by the concatenation
491 of all of the packs in the MIDX. E.g., if we had a MIDX with three packs
492 (a, b, c), with 10, 15, and 20 objects respectively, we can imagine an
493 ordering of the objects like:
494
495 |a,0|a,1|...|a,9|b,0|b,1|...|b,14|c,0|c,1|...|c,19|
496
497 where the ordering of the packs is defined by the MIDX's pack list,
498 and then the ordering of objects within each pack is the same as the
499 order in the actual packfile.
500
501 Given the list of packs and their counts of objects, you can
502 naïvely reconstruct that pseudo-pack ordering (e.g., the object at
503 position 27 must be (c,1) because packs "a" and "b" consumed 25 of the
504 slots). But there's a catch. Objects may be duplicated between packs, in
505 which case the MIDX only stores one pointer to the object (and thus we'd
506 want only one slot in the bitmap).
507
508 Callers could handle duplicates themselves by reading objects in order
509 of their bit-position, but that's linear in the number of objects, and
510 much too expensive for ordinary bitmap lookups. Building a reverse index
511 solves this, since it is the logical inverse of the index, and that
512 index has already removed duplicates. But, building a reverse index on
513 the fly can be expensive. Since we already have an on-disk format for
514 pack-based reverse indexes, let's reuse it for the MIDX's pseudo-pack,
515 too.
516
517 Objects from the MIDX are ordered as follows to string together the
518 pseudo-pack. Let `pack(o)` return the pack from which `o` was selected
519 by the MIDX, and define an ordering of packs based on their numeric ID
520 (as stored by the MIDX). Let `offset(o)` return the object offset of `o`
521 within `pack(o)`. Then, compare `o1` and `o2` as follows:
522
523 - If one of `pack(o1)` and `pack(o2)` is preferred and the other
524 is not, then the preferred one sorts first.
525 +
526 (This is a detail that allows the MIDX bitmap to determine which
527 pack should be used by the pack-reuse mechanism, since it can ask
528 the MIDX for the pack containing the object at bit position 0).
529
530 - If `pack(o1) ≠ pack(o2)`, then sort the two objects in descending
531 order based on the pack ID.
532
533 - Otherwise, `pack(o1) = pack(o2)`, and the objects are sorted in
534 pack-order (i.e., `o1` sorts ahead of `o2` exactly when `offset(o1)
535 < offset(o2)`).
536
537 In short, a MIDX's pseudo-pack is the de-duplicated concatenation of
538 objects in packs stored by the MIDX, laid out in pack order, and the
539 packs arranged in MIDX order (with the preferred pack coming first).
540
541 The MIDX's reverse index is stored in the optional 'RIDX' chunk within
542 the MIDX itself.
543
544 === `BTMP` chunk
545
546 The Bitmapped Packfiles (`BTMP`) chunk encodes additional information
547 about the objects in the multi-pack index's reachability bitmap. Recall
548 that objects from the MIDX are arranged in "pseudo-pack" order (see
549 above) for reachability bitmaps.
550
551 From the example above, suppose we have packs "a", "b", and "c", with
552 10, 15, and 20 objects, respectively. In pseudo-pack order, those would
553 be arranged as follows:
554
555 |a,0|a,1|...|a,9|b,0|b,1|...|b,14|c,0|c,1|...|c,19|
556
557 When working with single-pack bitmaps (or, equivalently, multi-pack
558 reachability bitmaps with a preferred pack), linkgit:git-pack-objects[1]
559 performs ``verbatim'' reuse, attempting to reuse chunks of the bitmapped
560 or preferred packfile instead of adding objects to the packing list.
561
562 When a chunk of bytes is reused from an existing pack, any objects
563 contained therein do not need to be added to the packing list, saving
564 memory and CPU time. But a chunk from an existing packfile can only be
565 reused when the following conditions are met:
566
567 - The chunk contains only objects which were requested by the caller
568 (i.e. does not contain any objects which the caller didn't ask for
569 explicitly or implicitly).
570
571 - All objects stored in non-thin packs as offset- or reference-deltas
572 also include their base object in the resulting pack.
573
574 The `BTMP` chunk encodes the necessary information in order to implement
575 multi-pack reuse over a set of packfiles as described above.
576 Specifically, the `BTMP` chunk encodes three pieces of information (all
577 32-bit unsigned integers in network byte-order) for each packfile `p`
578 that is stored in the MIDX, as follows:
579
580 `bitmap_pos`:: The first bit position (in pseudo-pack order) in the
581 multi-pack index's reachability bitmap occupied by an object from `p`.
582
583 `bitmap_nr`:: The number of bit positions (including the one at
584 `bitmap_pos`) that encode objects from that pack `p`.
585
586 For example, the `BTMP` chunk corresponding to the above example (with
587 packs ``a'', ``b'', and ``c'') would look like:
588
589 [cols="1,2,2"]
590 |===
591 | |`bitmap_pos` |`bitmap_nr`
592
593 |packfile ``a''
594 |`0`
595 |`10`
596
597 |packfile ``b''
598 |`10`
599 |`15`
600
601 |packfile ``c''
602 |`25`
603 |`20`
604 |===
605
606 With this information in place, we can treat each packfile as
607 individually reusable in the same fashion as verbatim pack reuse is
608 performed on individual packs prior to the implementation of the `BTMP`
609 chunk.
610
611 == cruft packs
612
613 The cruft packs feature offer an alternative to Git's traditional mechanism of
614 removing unreachable objects. This document provides an overview of Git's
615 pruning mechanism, and how a cruft pack can be used instead to accomplish the
616 same.
617
618 === Background
619
620 To remove unreachable objects from your repository, Git offers `git repack -Ad`
621 (see linkgit:git-repack[1]). Quoting from the documentation:
622
623 ----
624 [...] unreachable objects in a previous pack become loose, unpacked objects,
625 instead of being left in the old pack. [...] loose unreachable objects will be
626 pruned according to normal expiry rules with the next 'git gc' invocation.
627 ----
628
629 Unreachable objects aren't removed immediately, since doing so could race with
630 an incoming push which may reference an object which is about to be deleted.
631 Instead, those unreachable objects are stored as loose objects and stay that way
632 until they are older than the expiration window, at which point they are removed
633 by linkgit:git-prune[1].
634
635 Git must store these unreachable objects loose in order to keep track of their
636 per-object mtimes. If these unreachable objects were written into one big pack,
637 then either freshening that pack (because an object contained within it was
638 re-written) or creating a new pack of unreachable objects would cause the pack's
639 mtime to get updated, and the objects within it would never leave the expiration
640 window. Instead, objects are stored loose in order to keep track of the
641 individual object mtimes and avoid a situation where all cruft objects are
642 freshened at once.
643
644 This can lead to undesirable situations when a repository contains many
645 unreachable objects which have not yet left the grace period. Having large
646 directories in the shards of `.git/objects` can lead to decreased performance in
647 the repository. But given enough unreachable objects, this can lead to inode
648 starvation and degrade the performance of the whole system. Since we
649 can never pack those objects, these repositories often take up a large amount of
650 disk space, since we can only zlib compress them, but not store them in delta
651 chains.
652
653 === Cruft packs
654
655 A cruft pack eliminates the need for storing unreachable objects in a loose
656 state by including the per-object mtimes in a separate file alongside a single
657 pack containing all loose objects.
658
659 A cruft pack is written by `git repack --cruft` when generating a new pack.
660 linkgit:git-pack-objects[1]'s `--cruft` option. Note that `git repack --cruft`
661 is a classic all-into-one repack, meaning that everything in the resulting pack is
662 reachable, and everything else is unreachable. Once written, the `--cruft`
663 option instructs `git repack` to generate another pack containing only objects
664 not packed in the previous step (which equates to packing all unreachable
665 objects together). This progresses as follows:
666
667 1. Enumerate every object, marking any object which is (a) not contained in a
668 kept-pack, and (b) whose mtime is within the grace period as a traversal
669 tip.
670
671 2. Perform a reachability traversal based on the tips gathered in the previous
672 step, adding every object along the way to the pack.
673
674 3. Write the pack out, along with a `.mtimes` file that records the per-object
675 timestamps.
676
677 This mode is invoked internally by linkgit:git-repack[1] when instructed to
678 write a cruft pack. Crucially, the set of in-core kept packs is exactly the set
679 of packs which will not be deleted by the repack; in other words, they contain
680 all of the repository's reachable objects.
681
682 When a repository already has a cruft pack, `git repack --cruft` typically only
683 adds objects to it. An exception to this is when `git repack` is given the
684 `--cruft-expiration` option, which allows the generated cruft pack to omit
685 expired objects instead of waiting for linkgit:git-gc[1] to expire those objects
686 later on.
687
688 It is linkgit:git-gc[1] that is typically responsible for removing expired
689 unreachable objects.
690
691 === Alternatives
692
693 Notable alternatives to this design include:
694
695 - The location of the per-object mtime data.
696
697 On the location of mtime data, a new auxiliary file tied to the pack was chosen
698 to avoid complicating the `.idx` format. If the `.idx` format were ever to gain
699 support for optional chunks of data, it may make sense to consolidate the
700 `.mtimes` format into the `.idx` itself.
701
702 GIT
703 ---
704 Part of the linkgit:git[1] suite