Science Explorer Interactive view Map

Speech Recognition and Synthesis

Speech Recognition and Synthesis is a research topic within Artificial Intelligence. Science Explorer counts 35k research works in it since 1950. 16.9% of them reached the world's top 10% most cited for their field and year.

This cluster of papers focuses on the advances in speech recognition technology, covering topics such as acoustic modeling using deep neural networks, speaker verification, convolutional neural networks for speech recognition, end-to-end speech recognition systems, hidden Markov models, sequence-to-sequence models, automatic speech recognition, speaker diarization, and statistical language modeling.

  • Deep Neural Networks
  • Acoustic Modeling
  • Speaker Verification
  • Convolutional Neural Networks
  • End-to-End Speech Recognition
  • Hidden Markov Models
  • Sequence-to-Sequence Models
  • Automatic Speech Recognition
  • Speaker Diarization
  • Statistical Language Modeling
Research works
35k
fractional, since 1950
In the world top 10%
6k
per year above
Top-10% rate
16.9%
share of its works in the world top 10%
Growth, 2013–17 → 2018–22
+27%
the tick is no change

Which countries lead Speech Recognition and Synthesis research?

By volume, China and India publish the most (2k and 1.1k works in 2022–2025).

By volume, 2022–2025

  1. 1 China 2k works
  2. 2 India 1.1k works
  3. 3 United States 1k works
  4. 4 Japan 428 works
  5. 5 United Kingdom 296 works
  6. 6 South Korea 277 works
  7. 7 Germany 220 works
  8. 8 France 183 works
  9. 9 Taiwan 171 works
  10. 10 Canada 131 works

How concentrated that is

The same countries as shares of everything the list above accounts for. A node where two countries do two thirds of the work and one spread evenly across twelve read alike as a ranking and not at all alike here.

China: 33.8%India: 19.3%United States: 17.9%Japan: 7.3%6 others listed: 21.8%34%largest
China1,983 · 33.8%India1,130 · 19.3%United States1,047 · 17.9%Japan428 · 7.3%6 others listed1,279 · 21.8%

Shares of the rows listed above, not of the whole node.

Which institutions lead Speech Recognition and Synthesis research?

By volume in 2022–2025, University of Science and Technology of China publishes the most Speech Recognition and Synthesis research, followed by Google (United States) and Shanghai Jiao Tong University.

By volume, 2022–2025

  1. 1 University of Science and Technology of China China 64 works
  2. 2 Google (United States) United States 58 works
  3. 3 Shanghai Jiao Tong University China 56 works
  4. 4 Tsinghua University China 55 works
  5. 5 Xinjiang University China 50 works
  6. 6 Carnegie Mellon University United States 48 works
  7. 7 Zhejiang University China 40 works
  8. 8 Northwestern Polytechnical University China 40 works
  9. 9 Amrita Vishwa Vidyapeetham India 37 works
  10. 10 Amazon (United States) United States 37 works

Who are the leading researchers in Speech Recognition and Synthesis?

The most-cited researchers publishing on Speech Recognition and Synthesis include Andrew Zisserman, Geoffrey E. Hinton and Yoshua Bengio.

  1. 1 Andrew Zisserman 25k citations
  2. 2 Geoffrey E. Hinton 22k citations
  3. 3 Yoshua Bengio 17k citations
  4. 4 Wei Liu 9.7k citations
  5. 5 Quoc V. Le 8.6k citations
  6. 6 Thomas S. Huang 7.4k citations
  7. 7 Xiangyu Zhang 6.9k citations
  8. 8 Trevor Darrell 6.7k citations
  9. 9 Dacheng Tao 6.1k citations

Ranked by citations received across their whole record, among researchers with at least three works on this topic.

Where is Speech Recognition and Synthesis research done?

The largest centres of Speech Recognition and Synthesis research in 2022–2025 are Beijing (China), Tokyo (Japan), Seoul (South Korea) and Shanghai (China). Among places with at least 20 works in it, it is an unusually large share of all research in Redmond.

Largest cities, 2022–2025

  1. 1 Beijing China 464 works
  2. 2 Tokyo Japan 192 works
  3. 3 Seoul South Korea 148 works
  4. 4 Shanghai China 144 works
  5. 5 Shenzhen China 116 works
  6. 6 Chennai India 102 works
  7. 7 Xi'an China 96 works
  8. 8 Hangzhou China 88 works
  9. 9 Bengaluru India 88 works
  10. 10 Taipei Taiwan 87 works

Where it is the local speciality

  1. RedmondUS · 24.9 works34×
← less than its size predictsmore →

Location quotient: how much more of its research is in Speech Recognition and Synthesis than the world average.

See Speech Recognition and Synthesis on the map

Where is the best place to study Speech Recognition and Synthesis?

Among universities, judged by research, Carnegie Mellon University, Chinese University of Hong Kong, Shenzhen and Amrita Vishwa Vidyapeetham score highest, combining excellence, specialisation, size, growth and international reach. Research strength is one signal when choosing where to study; it does not measure teaching.

0%20%40%mean 24.03%fractional works in this node (log) →share in the world top 10% →Carnegie Mellon University: 48, 24.9%Chinese University of Hong Kong, Shenzhen: 16, 25.0%Amrita Vishwa Vidyapeetham: 37, 12.1%Shenzhen Research Institute of Big Data: 8, 30.3%Brno University of Technology: 19, 27.5%Northwestern Polytechnical University: 40, 25.3%University of Science and Technology of China: 64, 22.4%Dhirubhai Ambani University: 20, 21.0%Aalto University: 20, 21.9%Korea University: 21, 29.9%Shenzhen Research In…Chinese University o…Carnegie Mellon Univ…Amrita Vishwa Vidyap…
above the meannear itbelow it

One dot per university in the table below. The upper left is the interesting corner: small places doing unusually strong work.

#UniversityScoreTop 10%SpecialisationWorksGrowth
1Carnegie Mellon University United States 65.924.9%15.4×48 -7.3%
2Chinese University of Hong Kong, Shenzhen China 61.825.0%11.2×16
3Amrita Vishwa Vidyapeetham India 61.712.1%8.9×37 +340.1%
4Shenzhen Research Institute of Big Data China 60.330.3%57.3×8
5Brno University of Technology Czechia 58.927.5%15.2×19 -19.0%
6Northwestern Polytechnical University China 57.925.3%4.6×40 +154.3%
7University of Science and Technology of China China 57.722.4%7.0×64 -4.5%
8Dhirubhai Ambani University India 57.421.0%127.0×20 +72.3%
9Aalto University Finland 57.321.9%8.7×20 +24.1%
10Korea University South Korea 56.829.9%5.1×21 +117.3%

Universities only. Score blends excellence (30%), specialisation (25%), size (20%), growth (15%) and international reach (10%), 2015–2022; growth compares 2010–14 with 2015–19.

Is Speech Recognition and Synthesis research growing?

Output in 2018–2022 was 27% higher than in 2013–2017, peaking in 2025. The fastest-growing topics are Speech Recognition and Synthesis.

19801990200020102020
grewheldshrank

The same series as a ribbon — one cell per year, darker for more. The line above answers how much; this answers when.

Which topics inside it are moving

Growth and decline on one axis around a shared zero. Two lists side by side hide the thing that matters: whether the growth dwarfs the decline, or the other way round.