Science Explorer Interactive view Map

Multimodal Machine Learning Applications

Multimodal Machine Learning Applications is a research topic within Computer Vision and Pattern Recognition. Science Explorer counts 22k research works in it since 1975. 24.9% of them reached the world's top 10% most cited for their field and year.

This cluster of papers focuses on the development and improvement of visual question answering systems, image captioning techniques, and neural networks for understanding and generating descriptions of images and videos. The research involves semantic reasoning, multimodal fusion, scene graph generation, attention mechanisms, and deep learning approaches to bridge the gap between vision and language.

  • Visual Question Answering
  • Image Captioning
  • Neural Networks
  • Semantic Reasoning
  • Multimodal Fusion
  • Scene Graph Generation
  • Video Description
  • Attention Mechanism
  • Language Understanding
  • Deep Learning
Research works
22k
fractional, since 1975
In the world top 10%
5.5k
per year above
Top-10% rate
24.9%
share of its works in the world top 10%
Growth, 2013–17 → 2018–22
+338%
the tick is no change

Which countries lead Multimodal Machine Learning Applications research?

By volume, China and the United States publish the most (4.9k and 1.4k works in 2022–2025).

By volume, 2022–2025

  1. 1 China 4.9k works
  2. 2 United States 1.4k works
  3. 3 India 683 works
  4. 4 South Korea 333 works
  5. 5 United Kingdom 319 works
  6. 6 Japan 308 works
  7. 7 Germany 260 works
  8. 8 Australia 218 works
  9. 9 Singapore 189 works
  10. 10 Italy 178 works

How concentrated that is

The same countries as shares of everything the list above accounts for. A node where two countries do two thirds of the work and one spread evenly across twelve read alike as a ranking and not at all alike here.

China: 55.7%United States: 15.9%India: 7.8%South Korea: 3.8%6 others listed: 16.8%56%largest
China4,871 · 55.7%United States1,390 · 15.9%India683 · 7.8%South Korea333 · 3.8%6 others listed1,472 · 16.8%

Shares of the rows listed above, not of the whole node.

Which institutions lead Multimodal Machine Learning Applications research?

By volume in 2022–2025, Tsinghua University publishes the most Multimodal Machine Learning Applications research, followed by University of Science and Technology of China and Peking University.

By volume, 2022–2025

  1. 1 Tsinghua University China 134 works
  2. 2 University of Science and Technology of China China 104 works
  3. 3 Peking University China 102 works
  4. 4 University of Electronic Science and Technology of China China 102 works
  5. 5 Chinese Academy of Sciences China 97 works
  6. 6 Zhejiang University China 94 works
  7. 7 Beijing University of Posts and Telecommunications China 93 works
  8. 8 Shanghai Jiao Tong University China 91 works
  9. 9 University of Chinese Academy of Sciences China 90 works
  10. 10 Harbin Institute of Technology China 85 works

Who are the leading researchers in Multimodal Machine Learning Applications?

The most-cited researchers publishing on Multimodal Machine Learning Applications include Andrew Zisserman, Karen Simonyan and Ilya Sutskever.

  1. 1 Andrew Zisserman 25k citations
  2. 2 Karen Simonyan 23k citations
  3. 3 Ilya Sutskever 21k citations
  4. 4 Piotr Dollár 20k citations
  5. 5 Kaiming He 17k citations
  6. 6 Li Fei-Fei 17k citations
  7. 7 Yoshua Bengio 17k citations
  8. 8 Vincent Vanhoucke 15k citations
  9. 9 Serge Belongie 14k citations
  10. 10 Xiaogang Wang 13k citations

Ranked by citations received across their whole record, among researchers with at least three works on this topic.

Where is Multimodal Machine Learning Applications research done?

The largest centres of Multimodal Machine Learning Applications research in 2022–2025 are Beijing (China), Shanghai (China), Shenzhen (China) and Hangzhou (China). Among places with at least 20 works in it, it is an unusually large share of all research in Shanghai and Mountain View.

Largest cities, 2022–2025

  1. 1 Beijing China 1.2k works
  2. 2 Shanghai China 404 works
  3. 3 Shenzhen China 236 works
  4. 4 Hangzhou China 229 works
  5. 5 Nanjing China 228 works
  6. 6 Xi'an China 228 works
  7. 7 Guangzhou China 219 works
  8. 8 Singapore Singapore 189 works
  9. 9 Wuhan China 183 works
  10. 10 Seoul South Korea 181 works

Where it is the local speciality

  1. Shanghai26.7 works18×
  2. Mountain ViewUS · 58.0 works10×
← less than its size predictsmore →

Location quotient: how much more of its research is in Multimodal Machine Learning Applications than the world average.

See Multimodal Machine Learning Applications on the map

Where is the best place to study Multimodal Machine Learning Applications?

Among universities, judged by research, Hong Kong University of Science and Technology, Mohamed bin Zayed University of Artificial Intelligence and Nanyang Technological University score highest, combining excellence, specialisation, size, growth and international reach. Research strength is one signal when choosing where to study; it does not measure teaching.

0%20%40%60%mean 33.77%fractional works in this node (log) →share in the world top 10% →Hong Kong University of Science and Technology: 39, 37.7%Mohamed bin Zayed University of Artificial Intelligence: 18, 49.8%Nanyang Technological University: 64, 39.3%Carnegie Mellon University: 50, 27.6%University of Science and Technology of China: 104, 24.5%Tsinghua University: 134, 28.6%Singapore University of Technology and Design: 10, 51.9%Beijing University of Posts and Telecommunications: 93, 15.9%National University of Singapore: 62, 35.5%Renmin University of China: 33, 26.9%Mohamed bin Zayed Un…Nanyang Technologica…Hong Kong University…Carnegie Mellon Univ…
above the meannear itbelow it

One dot per university in the table below. The upper left is the interesting corner: small places doing unusually strong work.

#UniversityScoreTop 10%SpecialisationWorksGrowth
1Hong Kong University of Science and Technology Hong Kong 81.037.7%9.9×39 +258.9%
2Mohamed bin Zayed University of Artificial Intelligence United Arab Emirates 75.349.8%43.0×18
3Nanyang Technological University Singapore 73.339.3%7.9×64 +106.1%
4Carnegie Mellon University United States 71.827.6%12.1×50 +249.9%
5University of Science and Technology of China China 70.524.5%8.6×104 +282.9%
6Tsinghua University China 69.828.6%6.7×134 +557.8%
7Singapore University of Technology and Design Singapore 69.751.9%10.3×10
8Beijing University of Posts and Telecommunications China 67.615.9%15.0×93 +440.6%
9National University of Singapore Singapore 67.635.5%6.4×62 +121.4%
10Renmin University of China China 67.626.9%9.9×33 +795.8%

Universities only. Score blends excellence (30%), specialisation (25%), size (20%), growth (15%) and international reach (10%), 2015–2022; growth compares 2010–14 with 2015–19.

Is Multimodal Machine Learning Applications research growing?

Output in 2018–2022 was 338% higher than in 2013–2017, peaking in 2025. The fastest-growing topics are Multimodal Machine Learning Applications.

19801990200020102020
grewheldshrank

The same series as a ribbon — one cell per year, darker for more. The line above answers how much; this answers when.

Which topics inside it are moving

Growth and decline on one axis around a shared zero. Two lists side by side hide the thing that matters: whether the growth dwarfs the decline, or the other way round.