Skip to content

October 6, 2026 · Hash Verification & Integrity

NSRL Hash Sets: Filtering Known Files in an Examination

The National Software Reference Library maintains a database of cryptographic hashes of known software files, allowing forensic examiners to automatically filter out operating systems, applications, and other known content during an investigation. A hash match identifies what a file is; it does not establish who placed it there, when, or whether its presence is lawful.

What the NSRL Is and Why It Exists

The National Software Reference Library (NSRL) is a database of cryptographic hash values—digital signatures that uniquely identify files by their content—collected from legitimate software packages and maintained by the National Institute of Standards and Technology (NIST) in partnership with the U.S. Department of Justice's National Institute of Justice and federal, state, and local law enforcement.[1][8]

The NSRL has three distinct components: a physical collection of software packages, an internal database containing detailed information (metadata) about the files that compose those packages, and a public dataset called the Reference Data Set (RDS), which contains a subset of that metadata available for free download and updated quarterly.[1][3][8] The purpose is straightforward: to promote efficient computer forensics by allowing investigators to eliminate files known to be legitimate so that examination effort focuses on files that matter as evidence.[2][6]

How Hash Values Identify Files

A hash (or cryptographic hash) is a one-way mathematical function that accepts a file of any size and returns a fixed-length alphanumeric string—typically 32 characters for MD5, 40 for SHA-1, and 64 for SHA-256—that represents the file's exact content.[7][10] The function is one-way: knowing a hash value cannot reveal the original file, and no two different files will produce the same hash if the hashing algorithm is working correctly.[12]

The power of hashing for file identification lies in its determinism: the same file, on the same system or a different one, will always generate the same hash value, even if the file has been renamed, moved, or located in a different directory.[7][12] This means a hash value is a unique identifier of content, independent of the file's name, location, or metadata (timestamp, owner, permissions).

Modern NSRL Reference Data Sets (version 3.X and later) include SHA-256 hashes alongside the older MD5 and SHA-1 values, together with detailed product versioning, manufacturer information, original string data, and the file's location within the software package.[1][4] This additional context allows examiners not only to identify a file as known but to understand precisely which version of which software product it belongs to.

Known-File Filtering: Reducing Manual Review Burden

In a typical criminal investigation involving a seized computer, the file system may contain hundreds of thousands or millions of files: the operating system, bundled applications, libraries, drivers, and user data all intermixed. Manual review of every file is impractical. Known-file filtering reduces that burden by automatically eliminating files already catalogued in the NSRL.[5]

Forensic tools automate this process by computing the hash value of each file on the examined system and comparing it against the NSRL RDS.[5][10] Files that match a hash in the RDS are flagged as known; files that do not match are flagged as unknown or of interest. This immediately separates the known (operating system components, standard applications, legitimate libraries) from the potentially significant (user-created documents, unauthorized software, suspicious binaries).[5]

The NSRL is not limited to law enforcement. System administrators use it to verify that critical system files have not been altered or replaced by malware. Digital archivists use it to distinguish applications from user-created data. Corporate security teams use it to detect unauthorized software installations. Defense teams use it to discover exculpatory evidence by identifying legitimacy that the prosecution may not have considered.[1][5]

What a Hash Match Establishes: Content and Origin

A hash match against the NSRL establishes precisely three facts: the file's exact content (byte-for-byte), the software product it originates from, and (in modern RDS releases) the product version, manufacturer, and location within the installation package.[1][7][10]

That specificity is the hash's strength. If a Windows operating system file on the examined system produces a hash that matches an entry in the NSRL, the examiner can be certain the file is that operating system component, in that version. No guesswork about file names or locations is required. Hash matching is also far faster and more reliable than traditional file signature or metadata matching: the hash function detects even a single-bit change in a file's content, whereas a file name or modification date might not.[10]

For this reason, known-file filtering via hash matching has become the standard first step in forensic triage: it immediately removes the known and focuses analysis on the unknown or anomalous.

What a Hash Match Does Not Establish: Circumstance and Legality

This point is critical and often misunderstood. A hash match does not establish any fact about how the file came to be present on the system.[1][7]

Specifically, a hash match does not establish:

  • Who placed the file there. A match identifies Windows as the source of a system file, not whether the user, the administrator, malware, or an attacker installed or retained it.
  • When it was placed there. The hash is a function of content alone; it says nothing about the file's creation date, modification history, or when the file was first written to the system.
  • Whether the user or owner was aware of the file. An application file might have been installed legitimately and used, or bundled with legitimate software but never run, or present without the user's knowledge.
  • Whether possession is lawful or unlawful. A hash match against a legitimate software database establishes that the file is what the NSRL says it is. It does not establish whether installation, retention, or use of that software is legal, licensed, or appropriate in that user's hands or context.

The NSRL contains hashes of known, traceable software—operating systems, commercial applications, libraries, and legitimate utilities.[1][8] It does not contain hash values of illicit content, such as known child sexual abuse material.[1] Therefore, a negative match (a file that does not match any hash in the NSRL) does not prove the file is illegal or evidence of a crime. It merely indicates the file is not in the known-software database.[1]

This boundary is operationally vital. Hash-based filtering is a tool for triage—for separating known from unknown—not for judgment about legality, culpability, or the presence of crime.[1][5] The examiner who finds an unknown file must still investigate its origin, purpose, and relevance. The investigator who identifies an unauthorized software installation must still determine whether it was intentional, negligent, or imposed by malware. The attorney presenting evidence must still prove the circumstances that make possession or use of a known file incriminating in the specific case.

Technical Foundation: Why Hash Matching is Reliable

The reliability of NSRL-based identification rests on the collision resistance of the hashing algorithms used—the extreme improbability that two different files will ever produce the same hash value.[12]

NIST has confirmed that there are no detectable collisions between files in the NSRL database for MD5 or SHA-1, and no systematic bias introduced by the hashing process itself.[12] The probability of future collisions is negligible for all practical purposes in file identification work.[12]

While the Scientific Working Group on Digital Evidence (SWGDE) encourages the adoption of SHA-256 and SHA-3 by tool vendors and practitioners for new work, MD5 and SHA-1 remain acceptable for file identification and integrity verification purposes in digital forensics.[13] Modern NSRL RDS releases include SHA-256 hashes, providing examiners with the option to use more current algorithms while maintaining access to established historical data.

Accessing and Updating NSRL Hash Sets

The NSRL RDS is published and updated quarterly, with releases typically issued on the first Friday of March, June, September, and December.[3][9] The full database is available as a free download from NIST.[9] An annual subscription option (NIST Special Database 28) is also available with the same quarterly release schedule and additional metadata features.[3]

Examiners should verify when their local NSRL database was last updated. A hash match against an out-of-date RDS may miss newer software versions or recently released applications. Conversely, quarterly updates ensure that as new legitimate software is catalogued, examiners benefit from the most current filtering capability.

Using Known-File Sets in an Examination: Practical Application

The typical workflow is straightforward. When a forensic tool processes a seized device or disk image, the examiner selects the NSRL RDS (or a local database derived from it) for comparison. The tool computes the hash of every accessible file and compares each against the RDS. Results are typically presented in at least two categories: files matching the RDS (identified as known) and files not matching (unknown or unmatched).

The examiner then focuses detailed examination on the unknown set. This might include manual review of file content, malware scanning, timeline analysis, metadata examination, and correlation with investigative leads. Files in the known set are typically noted and catalogued but not subjected to detailed analysis—unless a specific investigative question targets them (such as verifying that a system was not modified by replacing operating system components with malicious versions).

It is important to recognize that hash matching is a filter, not a conclusion. A file that matches the NSRL is identified; a file that does not match is flagged for attention. Neither result, standing alone, resolves any evidentiary question about what occurred or who was responsible. Hash matching narrows the search space and reduces noise; it does not make the investigation automatic or fact-free.

Common questions

What is the National Software Reference Library?
The NSRL is a database of cryptographic hash values of files from known, legitimate software packages, maintained by NIST in partnership with law enforcement and the Department of Justice.[1] It contains hashes of operating systems, commercial applications, drivers, libraries, and other standard software, together with metadata identifying the product, version, manufacturer, and file location. The public NSRL Reference Data Set (RDS) is updated quarterly and available for free download.[3][9]
How are known files filtered out of an examination?
A forensic tool computes a cryptographic hash (MD5, SHA-1, or SHA-256) of each file on the examined system and compares it against hashes in the NSRL RDS database.[5][10] Files whose hashes match entries in the RDS are identified as known and separated from files that do not match. This automatic filtering allows examiners to focus detailed analysis on files not in the known-software database, dramatically reducing manual review time.
What does a hash match against a reference set establish?
A hash match establishes that the file's exact content matches a known software product and identifies the product name, version, manufacturer, and original file location within that software package.[1][7][10] Hash matching proves the file's identity and origin with certainty: a single differing byte in the file would produce a different hash, so a match is content-specific, not approximate or based on file name or metadata.
What does a hash match not establish?
A hash match does not establish who placed the file on the system, when it arrived, whether the user was aware of it, or whether its presence is lawful.[1][7] The NSRL contains hashes of legitimate software only, not illicit content, so a match means the file is identified as known software—it does not establish legality, intent, or culpability.[1] Circumstance and meaning require separate investigation beyond the hash match itself.

Sources

  1. [1] National Software Reference Library (NSRL) — National Institute of Standards and Technology
  2. [2] National Software Reference Library — National Institute of Standards and Technology
  3. [3] National Software Reference Library (NSRL) Reference Data Set (RDS) - NIST Special Database 28 — National Institute of Standards and Technology
  4. [4] Step Inside the National Software Reference Library — National Institute of Standards and Technology
  5. [5] Library Contents — National Institute of Standards and Technology
  6. [6] Digital Caseload Processing with the NIST National Software Reference Library — National Institute of Justice, U.S. Department of Justice
  7. [7] NSRL Introduction — National Institute of Standards and Technology
  8. [8] About the NSRL — National Institute of Standards and Technology
  9. [9] Current RDS Hash Sets — National Institute of Standards and Technology
  10. [10] Identification of Known Files on Computer Systems — National Institute of Standards and Technology
  11. [11] Digital Forensics at the National Institute of Standards and Technology — National Institute of Standards and Technology
  12. [12] Unique File Identification in the National Software Reference Library — National Institute of Standards and Technology
  13. [13] SWGDE Position on the Use of MD5 and SHA1 Hash — Scientific Working Group on Digital Evidence
  14. [14] Federal Rule of Evidence 901 — Authenticating or Identifying Evidence — Legal Information Institute, Cornell Law School
  15. [15] Federal Rule of Evidence 902 — Evidence That Is Self-Authenticating (including 902(13) and 902(14) and the Advisory Committee Notes) — Legal Information Institute, Cornell Law School
  16. [16] NIST SP 800-86 — Guide to Integrating Forensic Techniques into Incident Response — National Institute of Standards and Technology
  17. [17] H.R. Doc. 115-34 — Amendments to the Federal Rules of Evidence adopted by the Supreme Court on April 27, 2017 and effective December 1, 2017, adding Rules 902(13) and 902(14), with Advisory Committee Notes — U.S. Government Publishing Office
  18. [18] FIPS 180-4 — Secure Hash Standard (SHS) — National Institute of Standards and Technology
  19. [19] FIPS 202 — SHA-3 Standard: Permutation-Based Hash and Extendable-Output Functions — National Institute of Standards and Technology
  20. [20] NIST IR 8202 — Blockchain Technology Overview — National Institute of Standards and Technology
  21. [21] Computer Forensics Tool Testing Program (CFTT) — National Institute of Standards and Technology

CustodyTrack creates tamper-evident chain-of-custody records that any third party can verify. See how it works →

For this audience: Chain of Custody for Corporate Legal, IT & eDiscovery