Abstract
High dimensional, sparsely populated data spaces have been characterized in terms of ultrametric topology. This implies that there are natural, no tnecessarily unique, tree or hierarchy structures defined by the ultrametrictopology. In this note we study the extent of local ultrametric topology in texts, with the aim of finding unique ``fingerprints'' for a text or corpus, discriminating between texts from different domains, and opening up the possibility of exploiting hierarchical structures in the data. We use coherent and meaningful collections of over 1000 texts, comprising over 1.3 millionwords.
Original language | English |
---|---|
Publication status | Published - 27 Jan 2007 |
Keywords
- cs.CL
- I.5.3; I.7.2; H.3