Science Explorer Interactive view Map

Speech Recognition and Synthesis

Speech Recognition and Synthesis is a research topic within Artificial Intelligence. Science Explorer counts 35k research works in it since 1950. 16.9% of them reached the world's top 10% most cited for their field and year.

This cluster of papers focuses on the advances in speech recognition technology, covering topics such as acoustic modeling using deep neural networks, speaker verification, convolutional neural networks for speech recognition, end-to-end speech recognition systems, hidden Markov models, sequence-to-sequence models, automatic speech recognition, speaker diarization, and statistical language modeling.

  • Deep Neural Networks
  • Acoustic Modeling
  • Speaker Verification
  • Convolutional Neural Networks
  • End-to-End Speech Recognition
  • Hidden Markov Models
  • Sequence-to-Sequence Models
  • Automatic Speech Recognition
  • Speaker Diarization
  • Statistical Language Modeling
Research works
35k
fractional, since 1950
In the world top 10%
6k
per year above
Top-10% rate
16.9%
share of its works in the world top 10%
Growth, 2013–17 → 2018–22
+27%
the tick is no change

Which countries lead Speech Recognition and Synthesis research?

By volume, China and India publish the most (2k and 1.1k works in 2022–2025).

By volume, 2022–2025

  1. 1 China 2k works
  2. 2 India 1.1k works
  3. 3 United States 1k works
  4. 4 Japan 428 works
  5. 5 United Kingdom 296 works
  6. 6 South Korea 277 works
  7. 7 Germany 220 works
  8. 8 France 183 works
  9. 9 Taiwan 171 works
  10. 10 Canada 131 works

How concentrated that is

The same countries as shares of everything the list above accounts for. A node where two countries do two thirds of the work and one spread evenly across twelve read alike as a ranking and not at all alike here.

China: 33.8%India: 19.3%United States: 17.9%Japan: 7.3%6 others listed: 21.8%34%largest
China1,983 · 33.8%India1,130 · 19.3%United States1,047 · 17.9%Japan428 · 7.3%6 others listed1,279 · 21.8%

Shares of the rows listed above, not of the whole node.

Which institutions lead Speech Recognition and Synthesis research?

By volume in 2022–2025, University of Science and Technology of China publishes the most Speech Recognition and Synthesis research, followed by Google (United States) and Shanghai Jiao Tong University.

Who are the leading researchers in Speech Recognition and Synthesis?

The most-cited researchers publishing on Speech Recognition and Synthesis include Andrew Zisserman, Geoffrey E. Hinton and Yoshua Bengio.

  1. 1 Andrew Zisserman United Kingdom 25k citations
  2. 2 Geoffrey E. Hinton Canada 22k citations
  3. 3 Yoshua Bengio Canada 17k citations
  4. 4 Wei Liu China 9.7k citations
  5. 5 Quoc V. Le United States 8.6k citations
  6. 6 Thomas S. Huang United States 7.4k citations
  7. 7 Xiangyu Zhang 6.9k citations
  8. 8 Trevor Darrell United States 6.7k citations
  9. 9 Dacheng Tao Australia 6.1k citations

Ranked by citations received across their whole record, among researchers with at least three works on this topic.

Where is Speech Recognition and Synthesis research done?

The largest centres of Speech Recognition and Synthesis research in 2022–2025 are Beijing (China), Tokyo (Japan), Seoul (South Korea) and Shanghai (China). Among places with at least 20 works in it, it is an unusually large share of all research in Redmond.

Largest cities, 2022–2025

  1. 1 Beijing China 464 works
  2. 2 Tokyo Japan 192 works
  3. 3 Seoul South Korea 148 works
  4. 4 Shanghai China 144 works
  5. 5 Shenzhen China 116 works
  6. 6 Chennai India 102 works
  7. 7 Xi'an China 96 works
  8. 8 Hangzhou China 88 works
  9. 9 Bengaluru India 88 works
  10. 10 Taipei Taiwan 87 works

Where it is the local speciality

  1. RedmondUS · 24.9 works34×
← less than its size predictsmore →

Location quotient: how much more of its research is in Speech Recognition and Synthesis than the world average.

See Speech Recognition and Synthesis on the map

Where is the best place to study Speech Recognition and Synthesis?

Among universities, judged by research, Carnegie Mellon University, Chinese University of Hong Kong, Shenzhen and Amrita Vishwa Vidyapeetham score highest, combining excellence, specialisation, size, growth and international reach. Research strength is one signal when choosing where to study; it does not measure teaching.

0%20%40%mean 24.03%fractional works in this node (log) →share in the world top 10% →Carnegie Mellon University: 48, 24.9%Chinese University of Hong Kong, Shenzhen: 16, 25.0%Amrita Vishwa Vidyapeetham: 37, 12.1%Shenzhen Research Institute of Big Data: 8, 30.3%Brno University of Technology: 19, 27.5%Northwestern Polytechnical University: 40, 25.3%University of Science and Technology of China: 64, 22.4%Dhirubhai Ambani University: 20, 21.0%Aalto University: 20, 21.9%Korea University: 21, 29.9%Shenzhen Research In…Chinese University o…Carnegie Mellon Univ…Amrita Vishwa Vidyap…
above the meannear itbelow it

One dot per university in the table below. The upper left is the interesting corner: small places doing unusually strong work.

#UniversityScoreTop 10%SpecialisationWorksGrowth
1 Carnegie Mellon UniversityUnited States 65.924.9%15.4×48 -7.3%
2 Chinese University of Hong Kong, ShenzhenChina 61.825.0%11.2×16
3 Amrita Vishwa VidyapeethamIndia 61.712.1%8.9×37 +340.1%
4 Shenzhen Research Institute of Big DataChina 60.330.3%57.3×8
5 Brno University of TechnologyCzechia 58.927.5%15.2×19 -19.0%
6 Northwestern Polytechnical UniversityChina 57.925.3%4.6×40 +154.3%
7 University of Science and Technology of ChinaChina 57.722.4%7.0×64 -4.5%
8 Dhirubhai Ambani UniversityIndia 57.421.0%127.0×20 +72.3%
9 Aalto UniversityFinland 57.321.9%8.7×20 +24.1%
10 Korea UniversitySouth Korea 56.829.9%5.1×21 +117.3%

Universities only. Score blends excellence (30%), specialisation (25%), size (20%), growth (15%) and international reach (10%), 2015–2022; growth compares 2010–14 with 2015–19.

Is Speech Recognition and Synthesis research growing?

Output in 2018–2022 was 27% higher than in 2013–2017, peaking in 2025. The fastest-growing topics are Speech Recognition and Synthesis.

19801990200020102020
grewheldshrank

The same series as a ribbon — one cell per year, darker for more. The line above answers how much; this answers when.

Which topics inside it are moving

Growth and decline on one axis around a shared zero. Two lists side by side hide the thing that matters: whether the growth dwarfs the decline, or the other way round.