A file system is the data structure and associated management software that an operating system uses to organise, store, retrieve, and manage data on storage media. It defines how data is logically partitioned into named files and directories, governs access permissions, maintains metadata such as timestamps and ownership, and translates logical file operations into physical block I/O against underlying storage hardware. File systems range from local single-disk formats (NTFS, ext4, APFS) to distributed and network-attached systems (NFS, HDFS, GFS) that span many physical storage nodes.

Content

  • File systems are among the oldest abstractions in computing, predating modern operating systems in primitive form. Early systems such as IBM’s OS/360 PDS (Partitioned Data Set) and the Cambridge File Server formalised hierarchical naming and sequential access conventions in the 1960s. Unix’s virtual file system (VFS) abstraction, introduced in the early 1980s, decoupled application I/O calls from specific storage formats, establishing the inode-based model — where a file is a named reference to an inode containing ownership, permissions, timestamps, and block pointers — that underpins Linux ext2/3/4 and many other POSIX file systems to this day.
  • Modern file systems implement several layers of sophistication beyond simple file-to-block mapping. Journaling (ext3/4, NTFS, HFS+) maintains a write-ahead log to recover consistently from unexpected power loss. Copy-on-write semantics (Btrfs, ZFS, APFS) enable atomic snapshots and self-healing data integrity through per-block checksums. Log-structured file systems (F2FS, designed for flash storage) optimise write patterns for NAND characteristics. Network file systems (NFS, SMB/CIFS) extend the file system namespace across IP networks, presenting remote storage to local processes with native path semantics.
  • The file system is a critical infrastructure component whose design choices cascade through entire software stacks. AI training pipelines, for instance, depend heavily on file system throughput to feed GPU clusters — Lustre and GPFS (IBM Spectrum Scale) are specifically optimised for the large sequential reads that characterise training data ingestion. In container environments, overlay file systems (OverlayFS) enable copy-on-write layering of container images, reducing per-container storage footprint while maintaining isolation. File system performance characteristics (random vs. sequential IOPS, metadata operation latency) directly govern the performance envelope of databases, object stores, and analytics workloads running above them.
  • In 2024-2025, file system development is shaped by three vectors. First, persistent memory (CXL-attached DRAM and Optane successors) demands file systems with sub-microsecond path overhead, driving interest in NOVA and related PMEM-optimised designs. Second, the proliferation of AI workloads has intensified demand for parallel file systems with POSIX compatibility and multi-hundred-gigabyte-per-second aggregate bandwidth, leading hyperscalers to develop bespoke distributed file systems. Third, content-addressed and verifiable storage — exemplified by IPFS and blockchain-anchored provenance systems — challenges the traditional path-based naming model with hash-based addressing that enables global deduplication and tamper-evidence.