Text mining is the automated discovery of useful patterns, structure and knowledge from large collections of unstructured natural-language text. It combines natural-language processing, information retrieval and data-mining techniques to transform documents into structured representations amenable to analysis. Applications span extracting entities and relations, classifying and clustering documents, and surfacing trends across corpora too large to read manually.

Overview

  • Most enterprise and web knowledge resides in unstructured text that resists direct quantitative analysis.
  • Text mining converts such text into structured representations, then applies statistical and machine-learning methods to find meaning across documents.
  • The pipeline typically moves from preprocessing and tokenisation through linguistic analysis to extraction, classification, clustering and summarisation.
  • It differs from information retrieval, which finds relevant documents, by aiming to derive new, aggregate knowledge from the documents themselves.

Key aspects

  • Transformation of unstructured text into analysable representations.
  • Extraction of entities, relations and events.
  • Classification, clustering and topic discovery over corpora.
  • Trend and pattern detection across large document sets.
  • Evaluation against linguistic and task-specific ground truth.

Mechanisms

  • Tokenisation, normalisation and linguistic preprocessing.
  • Feature representation from bag-of-words to learned embeddings.
  • Supervised and unsupervised models for classification and clustering.
  • Rule-based and neural extraction of entities and relations.
  • Aggregation and visualisation of mined results.

Applications

  • Customer-feedback and social-media sentiment analysis.
  • Biomedical and scientific literature mining.
  • Legal and regulatory document analysis.
  • Fraud, risk and compliance monitoring.
  • Market and competitive intelligence.

Provenance