docs: Update readme
This commit is contained in:
@@ -1,32 +1,258 @@
|
||||
# LVM Salvage Tools
|
||||
# LVM/ext4 Salvage Tools
|
||||
|
||||
Tools to scan and recover files from damaged LVM volumes.
|
||||
Tools for conservative, file-level recovery from a damaged ext4 filesystem on an LVM logical volume when one physical volume (PV) is missing, failed, or has been replaced with a blank disk.
|
||||
|
||||
## Scan:
|
||||
These scripts were written after a failed PV in a **linear** LVM volume. The replacement PV could be recreated with the original UUID and geometry, which made the LV structurally activatable again, but the filesystem blocks that had lived on the failed disk were now effectively zero-filled. That means ordinary reads can appear to succeed even when the original data is gone.
|
||||
|
||||
```bash
|
||||
./scan.sh '/lvm-vol/path1'
|
||||
./scan.sh '/lvm-vol/path2'
|
||||
./scan.sh '/lvm-vol/path3'
|
||||
The important rule is therefore:
|
||||
|
||||
> **Do not treat a successful `debugfs dump` as proof that a file is intact. First prove that every physical extent for that file is outside the known-dead filesystem block range.**
|
||||
|
||||
The scripts do exactly that and only recover files classified `SAFE`.
|
||||
|
||||
## Files
|
||||
|
||||
```plain
|
||||
scan.sh Scan an arbitrary source directory and classify files.
|
||||
recover.sh Recover only files classified SAFE by the matching scan.
|
||||
salvage.py Shared recovery engine.
|
||||
```
|
||||
|
||||
## Recover:
|
||||
## Basic usage
|
||||
|
||||
Scan any directory inside the damaged filesystem:
|
||||
|
||||
```bash
|
||||
./recover.sh '/lvm-vol/path1'
|
||||
./scan.sh '/Warez/Books'
|
||||
./scan.sh '/Warez/3D Printing'
|
||||
./scan.sh '/Media/Some Folder'
|
||||
```
|
||||
|
||||
Then recover the `SAFE` files from the same path:
|
||||
|
||||
```bash
|
||||
./recover.sh '/Warez/Books'
|
||||
```
|
||||
|
||||
Always quote paths containing spaces.
|
||||
|
||||
The source tree is mirrored beneath `/mnt/recovery`:
|
||||
|
||||
```bash
|
||||
/lvm-vol/path1 -> /mnt/recovery/lvm-vol/path1
|
||||
/lvm-vol/path2 -> /mnt/recovery/lvm-vol/path2
|
||||
/lvm-vol/path3 -> /mnt/recovery/lvm-vol/path3
|
||||
```plain
|
||||
/Warez/Books -> /mnt/recovery/Warez/Books
|
||||
/Warez/3D Printing -> /mnt/recovery/Warez/3D Printing
|
||||
/Media/Some Folder -> /mnt/recovery/Media/Some Folder
|
||||
```
|
||||
|
||||
The scan creates uniquely named manifest/summary files under `/mnt/recovery`.
|
||||
Each scan creates uniquely named manifest and summary files under `/mnt/recovery`. The recovery run creates a TSV log and summary there as well.
|
||||
|
||||
Classifications:
|
||||
## Before using the scripts on another recovery
|
||||
|
||||
`salvage.py` currently contains recovery-specific constants near the top:
|
||||
|
||||
```python
|
||||
LV = "/dev/mapper/plex--new-plex--new"
|
||||
DEST_MOUNT = Path("/mnt/recovery")
|
||||
DEAD_START = 3418636288
|
||||
DEAD_END = 4395389951
|
||||
```
|
||||
|
||||
**Do not blindly reuse the dead block range on another filesystem.** Determine it again from the damaged LV layout using the process below.
|
||||
|
||||
The current scripts assume **one contiguous dead filesystem-block range**. If a failed PV supplied multiple non-contiguous LV segments, or more than one PV failed, the overlap logic must be changed to support multiple dead ranges.
|
||||
|
||||
---
|
||||
|
||||
# Determining the known-dead block range
|
||||
|
||||
There are two coordinate systems involved:
|
||||
|
||||
- `dmsetup` reports logical-volume offsets and lengths in **512-byte sectors**.
|
||||
- ext4 extents reported by `debugfs` use **filesystem blocks** (4096 bytes on the filesystem this project was written for).
|
||||
|
||||
The dead range is the LV sector range backed by the failed/blank PV, converted into ext4 filesystem-block numbers.
|
||||
|
||||
## 1. Keep the damaged LV read-only
|
||||
|
||||
Activate the VG/LV as needed, but do not mount it read-write and do not run a repairing `e2fsck` before salvage is complete.
|
||||
|
||||
When working with a missing PV, partial activation may be required:
|
||||
|
||||
```bash
|
||||
vgchange -ay --activationmode partial <vg-name>
|
||||
```
|
||||
|
||||
Then force the LV read-only at the block-device layer:
|
||||
|
||||
```bash
|
||||
blockdev --setro /dev/mapper/<vg--name>-<lv--name>
|
||||
blockdev --getro /dev/mapper/<vg--name>-<lv--name>
|
||||
```
|
||||
|
||||
`blockdev --getro` must return:
|
||||
|
||||
```plain
|
||||
1
|
||||
```
|
||||
|
||||
The scripts refuse to run unless the configured LV is read-only.
|
||||
|
||||
## 2. Determine the ext4 filesystem block size
|
||||
|
||||
Normal `debugfs` may refuse to open a badly damaged filesystem because of bitmap/checksum failures. Catastrophic mode (`-c`) is intentionally used here because it opens the filesystem read-only and avoids depending on allocation bitmaps.
|
||||
|
||||
```bash
|
||||
DEBUGFS_PAGER=cat debugfs -c \
|
||||
-R 'stats' \
|
||||
/dev/mapper/<vg--name>-<lv--name> | grep 'Block size'
|
||||
```
|
||||
|
||||
For this recovery the result was:
|
||||
|
||||
```plain
|
||||
Block size: 4096
|
||||
```
|
||||
|
||||
With a 4096-byte filesystem block, each filesystem block contains exactly 8 `dmsetup` sectors:
|
||||
|
||||
```plain
|
||||
4096 / 512 = 8
|
||||
```
|
||||
|
||||
## 3. Identify the LV segment that belongs to the failed/blank PV
|
||||
|
||||
### Best case: the PV is still missing
|
||||
|
||||
With the LV activated in partial mode, inspect the device-mapper table:
|
||||
|
||||
```bash
|
||||
dmsetup table /dev/mapper/<vg--name>-<lv--name>
|
||||
```
|
||||
|
||||
During this incident the relevant table was:
|
||||
|
||||
```plain
|
||||
0 15628050432 linear 8:16 2048
|
||||
15628050432 11721039872 linear 8:32 2048
|
||||
27349090304 7814029312 linear 252:0 0
|
||||
35163119616 11721039872 linear 8:48 2048
|
||||
46884159488 3907010560 linear 8:32 11721041920
|
||||
```
|
||||
|
||||
The suspicious segment was:
|
||||
|
||||
```plain
|
||||
27349090304 7814029312 linear 252:0 0
|
||||
```
|
||||
|
||||
Here:
|
||||
|
||||
```plain
|
||||
27349090304 = LV start sector
|
||||
7814029312 = segment length in sectors
|
||||
252:0 = device-mapper device backing that segment
|
||||
```
|
||||
|
||||
Inspect the referenced mapper device:
|
||||
|
||||
```bash
|
||||
cat /sys/dev/block/252:0/dm/name
|
||||
```
|
||||
|
||||
Then inspect its table:
|
||||
|
||||
```bash
|
||||
dmsetup table "$(cat /sys/dev/block/252:0/dm/name)"
|
||||
```
|
||||
|
||||
For this failure it returned:
|
||||
|
||||
```plain
|
||||
0 7814029312 error
|
||||
```
|
||||
|
||||
That `error` target proved that the entire LV segment was the missing PV region.
|
||||
|
||||
### If the failed PV has already been recreated on a blank replacement disk
|
||||
|
||||
Once the replacement PV is recreated and the historical VG metadata is restored, the LV may no longer contain an `error` target. The blank replacement disk will appear as an ordinary `linear` target.
|
||||
|
||||
In that case:
|
||||
|
||||
1. identify the replacement PV (`pvs`, `lsblk`, and the restored VG metadata are useful);
|
||||
2. identify its major:minor number with `lsblk -o NAME,MAJ:MIN`;
|
||||
3. inspect `dmsetup table` for the LV;
|
||||
4. find every LV segment backed by that replacement PV;
|
||||
5. treat those LV sector ranges as dead, because the replacement contains no original filesystem data.
|
||||
|
||||
Historical metadata is extremely useful here:
|
||||
|
||||
```bash
|
||||
vgcfgrestore -l <vg-name>
|
||||
lvs --segments -o lv_name,vg_name,seg_start,seg_size,devices
|
||||
pvs -o pv_name,pv_uuid,pv_size,pv_free
|
||||
```
|
||||
|
||||
If exact sector arithmetic is needed, prefer the exact `dmsetup table` values over rounded human-readable LVM sizes.
|
||||
|
||||
## 4. Convert the dead LV sector range to ext4 block numbers
|
||||
|
||||
`dmsetup` units are always 512-byte sectors.
|
||||
|
||||
Use:
|
||||
|
||||
```plain
|
||||
START_BLOCK = START_SECTOR * 512 / FILESYSTEM_BLOCK_SIZE
|
||||
BLOCK_COUNT = LENGTH_SECTORS * 512 / FILESYSTEM_BLOCK_SIZE
|
||||
END_BLOCK = START_BLOCK + BLOCK_COUNT - 1
|
||||
```
|
||||
|
||||
For this recovery:
|
||||
|
||||
```plain
|
||||
START_SECTOR = 27349090304
|
||||
LENGTH_SECTORS = 7814029312
|
||||
FILESYSTEM_BLOCK_SIZE = 4096
|
||||
```
|
||||
|
||||
Therefore:
|
||||
|
||||
```plain
|
||||
START_BLOCK = 27349090304 * 512 / 4096
|
||||
= 3418636288
|
||||
|
||||
BLOCK_COUNT = 7814029312 * 512 / 4096
|
||||
= 976753664
|
||||
|
||||
END_BLOCK = 3418636288 + 976753664 - 1
|
||||
= 4395389951
|
||||
```
|
||||
|
||||
The known-dead ext4 block range for this incident is therefore:
|
||||
|
||||
```plain
|
||||
3418636288 - 4395389951
|
||||
```
|
||||
|
||||
Equivalent Bash arithmetic:
|
||||
|
||||
```bash
|
||||
block_size=4096
|
||||
start_sector=27349090304
|
||||
length_sectors=7814029312
|
||||
|
||||
start_block=$(( start_sector * 512 / block_size ))
|
||||
block_count=$(( length_sectors * 512 / block_size ))
|
||||
end_block=$(( start_block + block_count - 1 ))
|
||||
|
||||
printf 'dead range: %s-%s\n' "$start_block" "$end_block"
|
||||
```
|
||||
|
||||
Before using the generic scripts on a different failure, update `DEAD_START` and `DEAD_END` in `salvage.py` with the calculated values.
|
||||
|
||||
---
|
||||
|
||||
# What the classifications mean
|
||||
|
||||
```plain
|
||||
SAFE
|
||||
@@ -37,17 +263,226 @@ Classifications:
|
||||
OTHER
|
||||
```
|
||||
|
||||
Only `SAFE` files are recovered.
|
||||
## `SAFE`
|
||||
|
||||
Safety:
|
||||
- damaged LV must remain kernel read-only;
|
||||
All physical data extents that `debugfs` can resolve lie outside the known-dead filesystem-block range.
|
||||
|
||||
Only `SAFE` files are automatically recovered.
|
||||
|
||||
A `SAFE` classification means the file does not depend on the known failed-PV region. It is still sensible to validate particularly important files afterward (hash against another known copy, open/test archives, use application-specific validation, etc.).
|
||||
|
||||
## `LOST`
|
||||
|
||||
At least one resolved data extent overlaps the dead filesystem-block range.
|
||||
|
||||
The file may be partially present, but it cannot be trusted as intact and is not automatically recovered.
|
||||
|
||||
## `ETB_LOST`
|
||||
|
||||
The inode survives, but the external ext4 **extent-tree block** (`ETB`) needed to locate the file's data lives inside the dead block range.
|
||||
|
||||
Example seen during this incident:
|
||||
|
||||
```plain
|
||||
EXTENTS:
|
||||
(ETB0):4338778112
|
||||
```
|
||||
|
||||
`4338778112` is inside `3418636288-4395389951`, so the mapping from file offsets to physical data blocks is gone.
|
||||
|
||||
This does **not** prove that all of the file's actual data was on the failed disk. It means ext4 has lost the map needed to find that data. Further recovery would require carving/forensic work rather than ordinary filesystem traversal.
|
||||
|
||||
## `ETB_SURVIVES`
|
||||
|
||||
An external extent-tree block exists outside the known-dead range, but usable data extents were not resolved by the scanner. Investigate manually before deciding the file is lost.
|
||||
|
||||
## `UNKNOWN`
|
||||
|
||||
The scanner could not obtain enough trustworthy extent information to classify the file.
|
||||
|
||||
Do not recover it automatically.
|
||||
|
||||
## `OTHER`
|
||||
|
||||
The directory entry is not a regular file (for example, a symlink or special file). The generic recovery script currently handles regular files only.
|
||||
|
||||
---
|
||||
|
||||
# Directory metadata can be lost too
|
||||
|
||||
A top-level directory entry can survive while the inode it points to is gone.
|
||||
|
||||
During this incident `/Backups` still appeared in the root directory, but:
|
||||
|
||||
```plain
|
||||
Inode: 496369665 Type: bad type Mode: 0000
|
||||
Size: 0
|
||||
Links: 0
|
||||
Blockcount: 0
|
||||
```
|
||||
|
||||
That means the directory's inode metadata was in the dead region. Without the directory inode, `debugfs` has no tree to traverse, even though some child inodes or file data may still physically survive elsewhere.
|
||||
|
||||
The generic scanner will stop immediately in this situation. Recovering such a tree requires forensic carving or targeted searches for known filenames/signatures and is outside the normal workflow.
|
||||
|
||||
---
|
||||
|
||||
# Recovery destination
|
||||
|
||||
`/mnt/recovery` must be a separate mounted filesystem. Confirm it before starting:
|
||||
|
||||
```bash
|
||||
findmnt /mnt/recovery
|
||||
df -hT /mnt/recovery
|
||||
```
|
||||
|
||||
Do not accidentally recover back onto the damaged LV.
|
||||
|
||||
The recovery drive used during this incident was exFAT. That is fine for file contents, but exFAT does not preserve normal Unix ownership/mode metadata. Do **not** rely on `debugfs dump -p` to preserve metadata on such a destination; it can report an ownership error after successfully writing the file.
|
||||
|
||||
The generic recovery script intentionally dumps file contents without `-p`.
|
||||
|
||||
---
|
||||
|
||||
# Why the scan matters even if `debugfs dump` appears to work
|
||||
|
||||
The failed PV in this incident was replaced with a same-sized blank disk so the original LVM geometry could be reconstructed.
|
||||
|
||||
That creates an important trap: reads from the formerly missing region no longer necessarily return an I/O error. They can return zero-filled blocks from the replacement disk.
|
||||
|
||||
Therefore this is **not sufficient**:
|
||||
|
||||
```bash
|
||||
debugfs -c -R 'dump <inode> /some/output/file' <LV>
|
||||
```
|
||||
|
||||
A dump can have the expected byte count while containing zeroes/corruption where the original file used blocks on the dead PV.
|
||||
|
||||
The scanner prevents this by checking physical extents first and only permitting automatic recovery when every resolved extent avoids the dead block range.
|
||||
|
||||
---
|
||||
|
||||
# Manual inspection commands
|
||||
|
||||
Inspect a directory:
|
||||
|
||||
```bash
|
||||
DEBUGFS_PAGER=cat debugfs -c \
|
||||
-R 'ls -l "/path/to/directory"' \
|
||||
/dev/mapper/<vg--name>-<lv--name>
|
||||
```
|
||||
|
||||
Inspect a file by path:
|
||||
|
||||
```bash
|
||||
DEBUGFS_PAGER=cat debugfs -c \
|
||||
-R 'stat "/path/to/file"' \
|
||||
/dev/mapper/<vg--name>-<lv--name>
|
||||
```
|
||||
|
||||
Inspect by inode (safer when filenames contain awkward characters):
|
||||
|
||||
```bash
|
||||
DEBUGFS_PAGER=cat debugfs -c \
|
||||
-R 'stat <123456789>' \
|
||||
/dev/mapper/<vg--name>-<lv--name>
|
||||
```
|
||||
|
||||
Ask `debugfs` for the blocks it can resolve:
|
||||
|
||||
```bash
|
||||
DEBUGFS_PAGER=cat debugfs -c \
|
||||
-R 'blocks <123456789>' \
|
||||
/dev/mapper/<vg--name>-<lv--name>
|
||||
```
|
||||
|
||||
Dump a manually verified-safe file by inode:
|
||||
|
||||
```bash
|
||||
DEBUGFS_PAGER=cat debugfs -c \
|
||||
-R 'dump <123456789> /mnt/recovery/recovered-file' \
|
||||
/dev/mapper/<vg--name>-<lv--name>
|
||||
```
|
||||
|
||||
Verify the output byte count against the inode's `Size:` value:
|
||||
|
||||
```bash
|
||||
stat -c '%s' /mnt/recovery/recovered-file
|
||||
```
|
||||
|
||||
For important files, also use a format-specific validator or compare a cryptographic hash against another known-good copy when available.
|
||||
|
||||
---
|
||||
|
||||
# Script safety behavior
|
||||
|
||||
The generic scripts intentionally make several conservative choices:
|
||||
|
||||
- the damaged LV must be kernel read-only;
|
||||
- `/mnt/recovery` must be a separate mounted filesystem;
|
||||
- filesystem root `"/"` is refused;
|
||||
- `".."` is refused;
|
||||
- each file is dumped by inode;
|
||||
- dumped byte size must match inode size;
|
||||
- existing same-size recovered files are skipped;
|
||||
- filesystem root `/` is refused as a scan target;
|
||||
- `..` path components are refused;
|
||||
- only regular files classified `SAFE` are recovered;
|
||||
- each file is dumped by inode to a temporary filename first;
|
||||
- the dumped byte size must exactly match the inode's `Size:` value;
|
||||
- only after the size check is the temporary file moved into the reconstructed tree;
|
||||
- existing destination files with the expected size are skipped, making recovery restartable;
|
||||
- at least 2 GiB free is preserved on `/mnt/recovery`.
|
||||
|
||||
Known dead ext4 block range:
|
||||
Do not remove these checks just to make a troublesome file copy. Investigate the troublesome file separately instead.
|
||||
|
||||
---
|
||||
|
||||
# Incident-specific reference values
|
||||
|
||||
These values apply to the original `plex-new` recovery only and are retained here as a worked example for future troubleshooting.
|
||||
|
||||
```plain
|
||||
VG: plex-new
|
||||
LV: plex-new
|
||||
Filesystem: ext4
|
||||
Filesystem block size: 4096 bytes
|
||||
|
||||
Failed/recreated PV UUID:
|
||||
QKoHgS-dCwc-tqVD-7o6R-ZoQr-6ab7-kdjYSK
|
||||
|
||||
Failed PV geometry:
|
||||
dev_size = 7814033408 sectors
|
||||
pe_start = 2048 sectors
|
||||
pe_count = 953861
|
||||
extent size = 8192 sectors (4 MiB)
|
||||
|
||||
Dead LV segment:
|
||||
start = 27349090304 sectors
|
||||
length = 7814029312 sectors
|
||||
|
||||
Dead ext4 filesystem blocks:
|
||||
3418636288 - 4395389951
|
||||
```
|
||||
|
||||
Historical five-segment LV layout after metadata restoration:
|
||||
|
||||
```plain
|
||||
start 0 -> /dev/sdb(0)
|
||||
start ~7.28T -> /dev/sdc(0)
|
||||
start ~12.74T -> failed/replacement /dev/sda1(0)
|
||||
start ~16.37T -> /dev/sdd(0)
|
||||
start ~21.83T -> remainder of /dev/sdc
|
||||
```
|
||||
|
||||
The exact `dmsetup` values, not these rounded TiB values, should be used for future arithmetic.
|
||||
|
||||
---
|
||||
|
||||
# Practical recovery order
|
||||
|
||||
When space or time is limited, recover in value order rather than disk order:
|
||||
|
||||
1. irreplaceable personal/project data;
|
||||
2. rare firmware, installers, keys, configs, archives, and documentation;
|
||||
3. collections that are difficult to reconstruct;
|
||||
4. replaceable media last.
|
||||
|
||||
Scan first, review the summary, then recover.
|
||||
|
||||
Once everything worth saving has been copied and independently verified, rebuild the filesystem/array cleanly rather than attempting to return this damaged ext4 filesystem to normal service.
|
||||
|
||||
Reference in New Issue
Block a user