Science Explorer Interactive view Map

Text and Document Classification Technologies

Text and Document Classification Technologies is a research topic within Artificial Intelligence. Science Explorer counts 23k research works in it since 1954. 22.9% of them reached the world's top 10% most cited for their field and year.

This cluster of papers focuses on the application of machine learning algorithms for multi-label text classification, with an emphasis on techniques such as feature selection, Naive Bayes classifier, K-nearest Neighbor (KNN), hierarchical classification, and support vector machines (SVM). The research covers various aspects of document categorization and information retrieval in the context of text mining and natural language processing.

  • Multi-label Learning
  • Text Classification
  • Feature Selection
  • Naive Bayes Classifier
  • K-nearest Neighbor (KNN)
  • Hierarchical Classification
  • Machine Learning Algorithms
  • Document Categorization
  • Support Vector Machines (SVM)
  • Information Retrieval
Research works
23k
fractional, since 1954
In the world top 10%
5.2k
per year above
Top-10% rate
22.9%
share of its works in the world top 10%
Growth, 2013–17 → 2018–22
+48%
the tick is no change

Which countries lead Text and Document Classification Technologies research?

By volume, China and India publish the most (2.6k and 980 works in 2022–2025).

By volume, 2022–2025

  1. 1 China 2.6k works
  2. 2 India 980 works
  3. 3 United States 400 works
  4. 4 Indonesia 195 works
  5. 5 United Kingdom 105 works
  6. 6 South Korea 103 works
  7. 7 Japan 102 works
  8. 8 Türkiye 101 works
  9. 9 Malaysia 92 works
  10. 10 Saudi Arabia 90 works

How concentrated that is

The same countries as shares of everything the list above accounts for. A node where two countries do two thirds of the work and one spread evenly across twelve read alike as a ranking and not at all alike here.

China: 54.4%India: 20.6%United States: 8.4%Indonesia: 4.1%6 others listed: 12.5%54%largest
China2,581 · 54.4%India980 · 20.6%United States400 · 8.4%Indonesia195 · 4.1%6 others listed592 · 12.5%

Shares of the rows listed above, not of the whole node.

Which institutions lead Text and Document Classification Technologies research?

By volume in 2022–2025, National University of Defense Technology publishes the most Text and Document Classification Technologies research, followed by Harbin Institute of Technology and Xidian University.

Who are the leading researchers in Text and Document Classification Technologies?

The most-cited researchers publishing on Text and Document Classification Technologies include Wei Liu, Chih‐Jen Lin and Thomas S. Huang.

  1. 1 Wei Liu China 9.7k citations
  2. 2 Chih‐Jen Lin Taiwan 9.6k citations
  3. 3 Thomas S. Huang United States 7.4k citations
  4. 4 Philip S. Yu United States 6.6k citations
  5. 5 Francisco Herrera Spain 6.5k citations
  6. 6 Dacheng Tao Australia 6.1k citations
  7. 7 Witold Pedrycz Canada 5.4k citations
  8. 8 Lei Zhang Hong Kong 5.3k citations
  9. 9 Jiawei Han United States 5.2k citations

Ranked by citations received across their whole record, among researchers with at least three works on this topic.

Where is Text and Document Classification Technologies research done?

The largest centres of Text and Document Classification Technologies research in 2022–2025 are Beijing (China), Shanghai (China), Nanjing (China) and Guangzhou (China). Among places with at least 20 works in it, it is an unusually large share of all research in Chandigarh.

Largest cities, 2022–2025

  1. 1 Beijing China 463 works
  2. 2 Shanghai China 141 works
  3. 3 Nanjing China 141 works
  4. 4 Guangzhou China 136 works
  5. 5 Xi'an China 126 works
  6. 6 Chengdu China 113 works
  7. 7 Wuhan China 100 works
  8. 8 Chennai India 88 works
  9. 9 Changsha China 85 works
  10. 10 Hefei China 75 works

Where it is the local speciality

  1. ChandigarhIN · 32.8 works5.1×
← less than its size predictsmore →

Location quotient: how much more of its research is in Text and Document Classification Technologies than the world average.

See Text and Document Classification Technologies on the map

Where is the best place to study Text and Document Classification Technologies?

Among universities, judged by research, Guangdong University of Technology, South China Normal University and Anhui University score highest, combining excellence, specialisation, size, growth and international reach. Research strength is one signal when choosing where to study; it does not measure teaching.

0%20%40%mean 26.78%fractional works in this node (log) →share in the world top 10% →Guangdong University of Technology: 31, 17.7%South China Normal University: 19, 21.6%Anhui University: 23, 33.2%Chongqing University of Posts and Telecommunications: 27, 16.2%Amrita Vishwa Vidyapeetham: 22, 22.0%National University of Defense Technology: 37, 28.2%National Institute of Technology Raipur: 9, 30.9%Xihua University: 8, 40.5%Minnan Normal University: 12, 26.8%Xidian University: 31, 30.7%Anhui UniversitySouth China Normal U…Guangdong University…Chongqing University…
above the meannear itbelow it

One dot per university in the table below. The upper left is the interesting corner: small places doing unusually strong work.

#UniversityScoreTop 10%SpecialisationWorksGrowth
1 Guangdong University of TechnologyChina 65.217.7%9.2×31 +276.2%
2 South China Normal UniversityChina 58.021.6%8.7×19 +181.2%
3 Anhui UniversityChina 57.133.2%8.3×23 +55.6%
4 Chongqing University of Posts and TelecommunicationsChina 56.816.2%15.6×27 +77.3%
5 Amrita Vishwa VidyapeethamIndia 56.022.0%6.9×22 +231.6%
6 National University of Defense TechnologyChina 55.528.2%7.9×37 +4.7%
7 National Institute of Technology RaipurIndia 55.330.9%9.9×9 +566.7%
8 Xihua UniversityChina 55.140.5%8.1×8 +233.3%
9 Minnan Normal UniversityChina 53.126.8%40.3×12 +122.1%
10 Xidian UniversityChina 53.030.7%7.0×31 -0.7%

Universities only. Score blends excellence (30%), specialisation (25%), size (20%), growth (15%) and international reach (10%), 2015–2022; growth compares 2010–14 with 2015–19.

Is Text and Document Classification Technologies research growing?

Output in 2018–2022 was 48% higher than in 2013–2017, peaking in 2025. The fastest-growing topics are Text and Document Classification Technologies.

19801990200020102020
grewheldshrank

The same series as a ribbon — one cell per year, darker for more. The line above answers how much; this answers when.

Which topics inside it are moving

Growth and decline on one axis around a shared zero. Two lists side by side hide the thing that matters: whether the growth dwarfs the decline, or the other way round.