Permutation invariant matrix statistics and computational language tasks

preprint OA: closed
View at publisher

Abstract

Abstract The Linguistic Matrix Theory (LMT) programme introduced by Kartsaklis, Ramgoolam and Sadrzadeh is an approach to the statistics of matrices generated and used in the fields of Computational Linguistics and Natural Language Processing. These matrices provide semantic representations for words and phrases of natural language. LMT is based on permutation invariant polynomial functions of matrices, which are regarded as the key observables encoding the significant statistics. In this paper, we generalize the previous results on the approximate Gaussianity of matrix distributions arising from a compositional distributional approach to natural language semantics. We also introduce two geometries of observable vectors for words, defined by exploiting the graph-theoretic basis for the permutation invariants and the statistical characteristics of the ensemble of matrices associated with the words. We describe successful applications of this unified framework to four tasks in semantic relation identification: synonym versus antonym distinction, a three-way classification of synonym pairs, antonym pairs and no-relation pairs, distinction between hypernym/hyponym pairs and co-hyponym pairs, and the distinction between a hypernym and co-hyponym within a pair.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00