File forensics and carving
Identify files by their bytes, find data appended past a format's end marker, and pull evidence out of images, archives and documents.
Assumes1. Recognising encodings
By the end you can
- Identify a file's real type regardless of extension
- Find and extract data hidden after a format's terminator
- Read EXIF, PNG chunks, and JPEG segments for metadata and appended payloads
- Take a PDF or Office document apart by object and stream, and find the active content in it
- Read an archive's central directory against its local headers, and spot the mismatch that hides a file
- Repair a deliberately corrupted header
1. Read
The first ten minutes: a triage playbook for any CTF challenge
Most challenges are lost to flailing, not to difficulty. Here is a repeatable order of operations for an unknown blob, an unknown file, and an unknown service - and the point at which you should stop guessing and start reading.
Archive attacks: ZIP crypto, known plaintext, and Zip Slip
Cracking a password-protected archive without cracking the password, why legacy ZipCrypto falls to twelve known bytes, path traversal through an entry name, and the structural tricks that make one archive hold two different sets of files.
Document forensics: taking apart a PDF and an Office file
A PDF is a graph of objects and an Office document is a zip of XML. Both hide things in places a viewer never renders. How to enumerate the structure, pull out streams and macros, and follow what the document tries to fetch.
2. Use the tools
In the order they come up while solving. Read what each one does, or go straight to the workspace tab that runs it.
- File type identifier by magic bytes
Drop a file and identify what it really is from its signature, regardless of extension. Also finds file headers embedded inside other files.
Open it in the workspace - Magic byte and file signature table
Look up any file format by its magic bytes, or any byte sequence by format - including the trailers that mark where a file ends.
Open it in the workspace - Strings extractor for binaries and blobs
Pull printable ASCII and UTF-16 strings out of any file, filtered by length, with flag-pattern highlighting.
Open it in the workspace - Hex viewer and hexdump
Inspect any file byte by byte with a side-by-side hex and ASCII view, offsets, and structure highlighting.
Open it in the workspace - PNG chunk analyzer
Walk a PNG chunk by chunk, validate CRCs, read tEXt and zTXt metadata, and find data hidden after IEND or in non-standard chunks.
Open it in the workspace - JPEG marker and segment analyzer
Parse JPEG markers, read APP segments and comments, and extract data appended after the end-of-image marker.
Open it in the workspace - EXIF metadata viewer
Read EXIF, GPS coordinates, camera details, timestamps, and embedded comments from images - the metadata that answers OSINT challenges.
Open it in the workspace - ZIP archive inspector
Read a ZIP’s central directory and local headers, spot mismatches used to hide files, and check encryption and compression per entry.
Open it in the workspace - PDF object and stream analyzer
Walk PDF objects, decompress streams, extract embedded files and JavaScript, and find text hidden under redaction boxes.
Open it in the workspace - Office macro and OLE extractor
Extract VBA macros from DOCM, XLSM, and legacy OLE documents, and deobfuscate the string-concatenation tricks they hide behind.
Open it in the workspace
3. Try one now
Generated in your browser and checked in your browser. No account, nothing to download, and a fresh one whenever you want another.
4. Practise on the real thing
Real picoCTF challenges that use these techniques, easiest first. 45 match in total - see the full index.
Common mistakes
The wrong turns this topic reliably produces. Written as the mistake rather than the rule, because the rule is easy to agree with and easy to walk straight past.
- Reaching for a steganography solver before running strings and checking the magic bytes. Most 'stego' challenges at this level are an archive glued onto an image.
- Trusting the file size. Every container format has a defined end, and the bytes after it are invisible to anything that respects the format.
- Extracting an archive to look inside it. Extraction runs the format's own logic, including path traversal in stored names - read the directory first.
- Repairing a broken header by pasting one in from another file. Dimensions, colour type and chunk offsets live in that header too, so the image opens, renders as noise, and looks repaired.
- Reading a document's text instead of its objects. The interesting part of a PDF or an Office file is the stream nobody renders - an embedded object, a macro, an external reference - and a text extractor is built to skip exactly that.
Checkpoint
Given a polyglot file, extract every embedded file it contains and state the offset and signature of each.
Teaching note
Students reach for the steganography solver too early here. Make them run strings and check the magic bytes first, every time - most 'stego' challenges at this level are a ZIP glued onto a PNG.
Go deeper
The lessons above are written to get you through a challenge. These go after the subject instead. Each one opens our notes on that chapter - what it argues, what to take from it and where it stops - so this is somewhere to read now rather than a book to buy first. Nothing here is affiliate-linked or sold by us.
Chapter 3, Forensic Image Formats
Practical Forensic Imaging - Bruce Nikkel
Container formats treated as evidence, with the metadata and integrity structures spelled out.
Chapter 9, Extracting Subsets of Forensic Images
Practical Forensic Imaging - Bruce Nikkel
Carving at offsets, done rigorously enough to defend afterwards rather than just quickly.
Chapter 1, Basic Static Malware Analysis
Malware Data Science - Joshua Saxe with Hillary Sanders
What static file structure alone can tell you before anything is executed or decoded.
Chapter 2, Malware Triage and Behavioral Analysis
Evasive Malware - Kyle Cucci
A triage order for unknown files that scales past the one-file-at-a-time habit CTFs teach.
Every book the curriculum cites has a page in the library.