Project

Shaping Efficiency-Aware Speech and Language Technology (SEASALT)

The training and deployment of modern artificial intelligence systems for speech and language require considerable computing resources. This is especially true for automatic speech recognition, where large neural networks are trained on hundreds to many thousands of hours of speech and many experiments are needed to compare model architectures, training strategies, search methods and hardware constraints. The goal of the SEASALT project was to improve the resource efficiency of speech and language technology without sacrificing recognition quality. We therefore investigated both algorithmic improvements and the interaction between models and modern hardware, including energy-efficient hardware concepts that perform parts of neural-network computation directly in memory. High-performance computing was essential for this work. Individual training runs often require several days on modern graphics processors, while meaningful scientific conclusions require many such runs. The granted resources allowed us to scale models and datasets to more realistic sizes and to perform systematic parameter studies that would not have been feasible on our local infrastructure alone.

Project Details

Project term

April 1, 2025–June 30, 2026

Affiliations

RWTH Aachen University

Institute

Chair of Machine Learning and Reasoning

Principal Investigator

Dr. Ralf Schlüter

Methods

We used neural-network models for automatic speech recognition and language modeling and trained them on large speech and text datasets. The project compared several recognition architectures and investigated more efficient model components, training procedures and decoding strategies. We also studied pre-trained language models in speech recognition, both by integrating them directly into speech models and by combining them with existing recognition systems during decoding. For data-efficient training, we generated synthetic speech from large text collections and used it to create additional training material. We further investigated simplified learnable feature extraction, which replaces parts of conventional hand-designed signal processing with trainable components. For efficient decoding, we developed and evaluated new search methods and performed extensive parameter sweeps. A further part of the project simulated emerging in-memory computing hardware on graphics processors. This allowed us to study how realistic hardware effects, limited numerical precision and device noise influence speech recognition before suitable physical hardware is available. Large training and simulation experiments were executed on the high-performance computing system, while preprocessing and smaller tasks were partly handled locally.

Results

The granted computing resources enabled progress in all planned research areas. We investigated three new speech-recognition architectures, simplified learnable feature extraction, new search methods, large-scale synthetic data generation, training on a substantially larger speech corpus, and the simulation of in-memory hardware for speech recognition. For pre-trained language models, we found that direct integration into a speech model could achieve reasonable performance after limited fine-tuning. Although this approach did not yet match the baseline system on data similar to the training domain, it performed better on data from different domains. Combining a pre-trained language model with an existing recognition system during decoding produced clear improvements while requiring only limited additional training. We also obtained promising results with alternative language-modeling approaches that allow more parallel computation during decoding. New search strategies improved recognition compared with the underlying speech-recognition baseline and reduced computational cost compared with conventional model combinations at similar recognition quality. In feature extraction, we showed that the neural front-end can be simplified substantially while maintaining competitive recognition performance when the data-augmentation procedure is adapted. This reduces model complexity and removes several hand-designed processing steps. Synthetic-data experiments were scaled to hundreds of millions of words. Parallel processing was important because the total generation and recognition workload would otherwise have taken too long for practical iteration. We also trained several speech-synthesis and recognition systems on a more realistic dataset containing about 25,000 hours of transcribed speech. For the hardware simulations, we increased model size to about 100 million parameters and extended the simulation so that all major fixed-weight operations of the recognition model could be mapped to the simulated devices. We identified limitations caused by restricted numerical ranges and developed architectural changes that restored performance. We also improved the workflow by moving expensive initialization steps away from the graphics processors, substantially reducing idle accelerator time.

Discussion

The first project phase showed that high-performance computing is required not only to train larger models, but also to study efficiency itself. Many relevant questions can only be answered through systematic comparisons across architectures, data sizes, search parameters and simulated hardware settings. Modern graphics processors with large memory enabled configurations that
previously exceeded local hardware limits, while parallel jobs enabled detailed parameter studies. The work also revealed several challenges. Larger models increase memory requirements, synthetic-data pipelines can create excessive storage load, and hardware simulations may introduce substantial overhead if the workflow is not carefully organized. Some planned studies on model size, depth and numerical precision had to be postponed because additional experiments were required to resolve unexpected issues with positional information in the simulated hardware. The results provide a strong basis for the next project phase. We will extend the work towards training with little or no labeled speech data, larger multilingual settings, further large-scale modeling, and a deeper analysis of model design for energy-efficient hardware. A central goal remains to improve the balance between recognition quality, training cost, decoding speed and hardware efficiency under realistic conditions.

Additional Project Information

DFG classification: 409-05 Interactive and Intelligent Systems, Image and Language Processing, Computer Graphics and Visualisation
Software: PyTorch, RETURNN, Huggingface Transformers
Cluster: CLAIX

Publications

Unified Learnable 2D Convolutional Feature Extraction for AS,
Peter Vieting, Benedikt Hilmes, Ralf Schlüter, Hermann Ney,
https://dx.doi.org/10.48550/arXiv.2509.10031, September 2025

Reproducing and Dissecting Denoising Language Models for Speech Recognition,
Dorian Koch, Albert Zeyer, Nick Rossenbach, Ralf Schlüter, Hermann Ney,
https://dx.doi.org/10.48550/arXiv.2512.13576, December 2025

LLMs and Speech: Integration vs. Combination,
Robin Schmitt, Albert Zeyer, Mohammad Zeineldeen, Ralf Schlüter, Hermann Ney,
https://dx.doi.org/10.48550/arXiv.2603.15045, March 2026

Supplementary Resources and Analysis for Automatic Speech Recognition Systems Trained on the Loquacious Dataset,
Nick Rossenbach, Robin Schmitt, Tina Raissi, Simon Berger, Larissa Kleppel, Ralf Schlüter,
https://dx.doi.org/10.63317/4zsvhm25r7zf, May 2026

 

Links

RETURNN https://github.com/rwth-i6/returnn
Huggingface Transformers https://github.com/huggingface/transformers
PyTorch https://github.com/pytorch/pytorch