A Content Identifier (CID) is a self-describing, cryptographically derived label used in the InterPlanetary File System and IPLD ecosystem to uniquely identify and verify content through its hash, encoding the hash function used, the hash digest, and a multicodec descriptor for the serialised data format into a compact, version-aware multihash structure.

Content

  • The Content Identifier was designed by Juan Benet and the Protocol Labs team as part of the IPFS project, first specified around 2015–2016. IPFS’s foundational insight was that location-addressed URLs (http://server/path) are fragile and censor-able, whereas content-addressed identifiers that encode the hash of the data itself are permanent and verifiable regardless of where the data is stored. CID formalised this by creating a single compact binary structure that is self-describing: it encodes which version of the CID format is being used, which codec serialises the data, and which hash function produced the digest.
  • A CID is constructed as a concatenation of four components: a version byte (CIDv0 or CIDv1), a multicodec varint identifying the data format (dag-pb for legacy protobuf trees, dag-cbor for CBOR-encoded IPLD, raw for unstructured bytes, etc.), a multihash prefix identifying the hash algorithm (SHA2-256, BLAKE2b, etc.), and the hash digest itself. CIDv0 is a legacy format compatible with original IPFS addresses and is always a base58-encoded SHA2-256 multihash of a dag-pb block. CIDv1 introduces multibase encoding (base32, base64url) and codec flexibility. Human-readable CIDs begin with “bafy…” (base32, dag-cbor) or “QmY…” (base58, dag-pb) prefixes recognisable in practice.
  • CIDs enable a powerful set of properties in distributed systems. Content deduplication is automatic: identical data always produces the same CID regardless of origin. Tamper detection is unconditional: any modification to content changes the hash and thus the CID, making inconsistency immediately detectable without trusting the server. Permanent links are achievable when content is pinned by storage providers: a CID published today remains valid and retrievable years later even if the original publisher goes offline. These properties have made CIDs the addressing mechanism of choice for NFT metadata storage, decentralised web applications, and verifiable dataset archiving.
  • By 2024–2025, CIDs are embedded in a wide range of Web3 infrastructure. The Ethereum ecosystem uses CIDs in ERC-721 NFT token URIs pointing to IPFS-stored metadata, and ERC-4973 and related standards are deepening on-chain CID verification. The W3C Verifiable Credentials Data Model 2.0 uses IPLD and CIDs for content-addressing credential schemas. Filecoin’s storage deals are indexed by CID, and the IPNI (InterPlanetary Network Indexer) provides a scalable lookup service mapping CIDs to providers. Research into CID caching, provider reputation, and content routing at scale is active, as IPFS networks have grown to hundreds of millions of stored CIDs.