Skip to content
All modules
CoreModule 16 of 2790 minutes

File forensics and carving

Identify files by their bytes, find data appended past a format's end marker, and pull evidence out of images, archives and documents.

Assumes1. Recognising encodings

By the end you can

  • Identify a file's real type regardless of extension
  • Find and extract data hidden after a format's terminator
  • Read EXIF, PNG chunks, and JPEG segments for metadata and appended payloads
  • Take a PDF or Office document apart by object and stream, and find the active content in it
  • Read an archive's central directory against its local headers, and spot the mismatch that hides a file
  • Repair a deliberately corrupted header

1. Read

2. Use the tools

In the order they come up while solving. Read what each one does, or go straight to the workspace tab that runs it.

3. Try one now

Generated in your browser and checked in your browser. No account, nothing to download, and a fresh one whenever you want another.

4. Practise on the real thing

Real picoCTF challenges that use these techniques, easiest first. 45 match in total - see the full index.

Common mistakes

The wrong turns this topic reliably produces. Written as the mistake rather than the rule, because the rule is easy to agree with and easy to walk straight past.

  • Reaching for a steganography solver before running strings and checking the magic bytes. Most 'stego' challenges at this level are an archive glued onto an image.
  • Trusting the file size. Every container format has a defined end, and the bytes after it are invisible to anything that respects the format.
  • Extracting an archive to look inside it. Extraction runs the format's own logic, including path traversal in stored names - read the directory first.
  • Repairing a broken header by pasting one in from another file. Dimensions, colour type and chunk offsets live in that header too, so the image opens, renders as noise, and looks repaired.
  • Reading a document's text instead of its objects. The interesting part of a PDF or an Office file is the stream nobody renders - an embedded object, a macro, an external reference - and a text extractor is built to skip exactly that.

Checkpoint

Given a polyglot file, extract every embedded file it contains and state the offset and signature of each.

Teaching note

Students reach for the steganography solver too early here. Make them run strings and check the magic bytes first, every time - most 'stego' challenges at this level are a ZIP glued onto a PNG.

Go deeper

The lessons above are written to get you through a challenge. These go after the subject instead. Each one opens our notes on that chapter - what it argues, what to take from it and where it stops - so this is somewhere to read now rather than a book to buy first. Nothing here is affiliate-linked or sold by us.

Every book the curriculum cites has a page in the library.