Schema.org is a collaborative, community-maintained vocabulary project founded in 2011 by Google, Microsoft, Yahoo, and Yandex to define a shared set of structured data markup schemas for web pages, enabling search engines and other consumers to understand the semantic content of web resources. The vocabulary defines types and properties for entities such as persons, organisations, events, products, reviews, and creative works, expressed using JSON-LD, Microdata, or RDFa. Schema.org markup embedded in web pages allows search engines to generate rich snippets, knowledge panels, and structured results. The vocabulary is extensible and hosted at schema.org, with governance managed by a W3C community group.

Content

  • Schema.org was created in response to a fragmented landscape of proprietary structured data formats—rich snippets, open graph, microformats—that each search engine interpreted differently. By providing a single, shared vocabulary endorsed by all major search engines, Schema.org created an incentive for web publishers to invest in structured markup, knowing it would be interpreted consistently across platforms. Within a few years of launch it became one of the most widely deployed web standards, appearing in tens of millions of websites.
  • The vocabulary is organised as a type hierarchy rooted in schema:Thing, with major branches for creative works, events, organisations, persons, places, products, and actions. Each type has associated properties specifying the data that can be attached, with expected value types (Text, URL, another schema.org type, etc.). Extensions to the core vocabulary are managed through hosted extensions and external vocabularies that follow the schema.org linking conventions. The governance model, operated through a W3C community group, allows community contributions whilst maintaining vocabulary coherence through editorial review.
  • In the era of large language models, Schema.org markup has taken on additional significance. LLM training corpora that include structured web content benefit from schema.org annotations that provide machine-readable labels for entity types, relationships, and factual claims. Conversely, knowledge graph construction pipelines for search engines use schema.org as the target vocabulary when extracting entities and relationships from unstructured text, closing a feedback loop between web markup and AI training data.
  • Schema.org’s relationship with Ontology formalisms is complementary rather than competitive. Whilst Schema.org prioritises broad adoption and simplicity over logical rigour, the vocabulary is mapped to more expressive formalisms including OWL and the Wikidata property model. This allows schema.org-annotated data to be lifted into richer knowledge representation frameworks when more precise semantic reasoning is required, making it a pragmatic entry point into the linked-data ecosystem rather than a terminal destination.