Extracting Shared Topics of Multiple Documents (original) (raw)

Abstract

In this paper, we present a weighted graph based method to simultaneously compare the textual content of two or more documents and extract the shared (sub)topics of them, if available. A set of documents are modelled with a set of pairwise weighted bipartite graphs. A generalized mutual reinforcement principle is applied to the pairwise bipartite graphs to calculate the saliency scores of sentences in each documents based on pairwise weighted bipartite graphs. Sentences with advantaged saliency are selected, and they together convey the dominant shared topic. If there are more than one shared subtopics among the documents, a spectral min-max cut algorithm can be used to partition a derived sentence similarity graph into several subgraphs. For a subgraph, if all documents contribute some sentences(nodes) to it, then these sentences(nodes) in the subgraph may convey a shared subtopic. The generalized mutual reinforcement principle is applied to them to verify and extract the shared subtopic.

Preview

Unable to display preview. Download preview PDF.

References

J. Dean and M. Henzinger. Finding Related Pages in the World Wide Web. Proceedings of WWW8, pp. 1467–1479. 1999.
Google Scholar
C. Ding, X. He, H. Zha, M. Gu, and H. Simon. Spectral Min-max Cut for Graph Partitioning and Data Clustering. Proceedings of the First IEEE International Conference on Data Mining, pp. 107–114, 2001.
Google Scholar
L. Ertoz, and M. Steinbach and V. Kumar. Finding Topics in Collections of Documents: A Shared Nearest Neighbor Approach. Proceedings of 1st SIAM International Conference on Data Mining, 2001.
Google Scholar
T. H. Haveliwala, A. Gionis, D. Klein, and P. Indyk. Evaluating Strategies for Similarity Search on the Web. Proceedings of WWW11, 2002.
Google Scholar
J. Mani and E. Bloedorn. Multi-document Summarization by Graph Search and Matching. Proceedings of the Fourteenth National Conference on Artificial Intelligence, pp. 622–628, 1997.
Google Scholar
I. Mani and M. Maybury. Advances in automatic text summarization. MIT Press, 1999.
Google Scholar
M. Porter. The Porter Stemming Algorithm. http://www.tartarus.org/martin/PorterStemmer
D. R. Radev and K. R. McKeown. Generating Natural Language Summaries from Multiple On-Line Sources. Computational Linguistics, Vol. 24, pp. 469–500, 1999.
Google Scholar
Topic Detection and Tracking. http://www.nist.gov/speech/tests/tdt/
C. Wayne. Multilingual Topic Detection and Tracking: Successful Research Enabled by Corpora and Evaluation. Language Resources and Evaluation Conference, pp. 1487–1494, 2000.
Google Scholar
H. Zha, X. He, C. Ding, M. Gu, and H. Simon. Bipartite Graph Partitioning and Data Clustering. Proceedings of ACM CIKM, pp. 25–32, 2001.
Google Scholar
H. Zha. Generic Summarization Keyphrase Extraction Using Mutual Reinforcement Principle and Sentence Clustering. Proceedings of the 25th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 113–120, 2002.
Google Scholar
H. Zha and X. Ji. Correlating Multilingual Documents via Bipartite Graph Modeling. Proceedings of the 25th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 443–444, 2002.
Google Scholar

Download references

Author information

Authors and Affiliations

Department of Computer Science and Engineering, The Pennsylvania State University, University Park, PA, 16802, USA
Xiang Ji & Hongyuan Zha

Authors

Xiang Ji
Hongyuan Zha

Editor information

Editors and Affiliations

Computer Science Department, Korea Advanced Institute of Science and Technology, 373-1 Koo-Sung Dong, Yoo-Sung Ku, Daejeon, 305-701, Korea
Kyu-Young Whang
Department of Statistics, Seoul National University, Sillimdong Kwanakgu, Seoul, 151-742, Korea
Jongwoo Jeon
School of Electrical Engineering and Computer Science, Seoul National University, Kwanak P.O. Box 34, Seoul, 151-742, Korea
Kyuseok Shim
Department of Computer Science and Engineering, University of Minnesota, 200 Union St SE, Minneapolis, MN, 55455, USA
Jaideep Srivastava

Rights and permissions

Copyright information

About this paper

Cite this paper

Ji, X., Zha, H. (2003). Extracting Shared Topics of Multiple Documents. In: Whang, KY., Jeon, J., Shim, K., Srivastava, J. (eds) Advances in Knowledge Discovery and Data Mining. PAKDD 2003. Lecture Notes in Computer Science(), vol 2637. Springer, Berlin, Heidelberg. https://doi.org/10.1007/3-540-36175-8\_10

Download citation

.RIS
.ENW
.BIB
DOI: https://doi.org/10.1007/3-540-36175-8\_10
Published: 30 April 2003
Publisher Name: Springer, Berlin, Heidelberg
Print ISBN: 978-3-540-04760-5
Online ISBN: 978-3-540-36175-6
eBook Packages: Springer Book Archive