A web crawler is an automated program that systematically traverses the web by following hyperlinks, fetching pages, and queuing newly discovered URLs for further retrieval. It underpins search engine indexing, SEO analysis, and large-scale dataset construction for training language models, typically respecting robots.txt directives and rate limits to avoid overloading target servers. Crawlers must handle duplicate content detection, URL normalisation, and politeness policies at scale, distinguishing them from targeted scraping of a known set of pages. Modern crawlers increasingly render JavaScript to capture content generated by single-page applications.