Submitted:
06 November 2025
Posted:
07 November 2025
You are already at the latest version
Abstract

Keywords:
1. Introduction
1.1. Background and Motivation
1.2. Challenges
- Heterogeneous Audio Data: Varying recording quality, background noise, poorly recorded audio, and multilingual content require robust preprocessing and adaptable models.
- Large-Scale Processing: With thousands of hours of content, the system must support high-throughput, GPU-intensive inference without disrupting ongoing user services.
- Complex Orchestration: Coordinating multiple AI models, each with distinct runtime requirements, demands reliable task scheduling, monitoring, and failure recovery.
- Computational Workload Management: Efficiently distributing computational workloads between front-facing services and heavy inference tasks is essential for performance and scalability.
- Personalization & Retrieval Accuracy: Generating accurate recommendations and search results depends on effective integration of metadata filtering, semantic embeddings, and user preference modeling.
1.3. Scope of This Work
1.4. Contributions
- End-to-End AI Podcast Pipeline: An integrated framework combining audio preprocessing, ASR, audio classification, topic modeling, semantic retrieval, and recommendation into a single, unified workflow.
- Multi-Server Deployment: A separation of web services and inference tasks, enabling efficient GPU utilization and uninterrupted user access.
- Orchestration with Apache Airflow: Automated, fault-tolerant scheduling and coordination of processing tasks, ensuring scalability and maintainability.
- Dense Retrieval and Sparse Retrieval: Integration of dense semantic retrieval with metadata filtering to improve both relevance and personalization.
- Recommendation System: Generating accurate recommendations and search results depends on effective integration of metadata filtering, semantic embeddings, and user preference modeling.
- Real-World Implementation: Deployment and evaluation on a large corpus of multilingual podcasts, demonstrating the practicality of the proposed system.
2. Related Works
3. System Overview
- Web Servers host the user interface, search and browse features, and REST APIs for authentication, CRUD operations, and metadata filtering. They manage podcast metadata (MySQL and MongoDB) and caching (Redis) to ensure fast user interaction.
- AI Server runs the podcast processing pipeline as Apache Airflow DAGs. This includes speech-to-text transcription (FasterWhisperXXL) [14], post-ASR correction (Gemini 2.5 Flash) [15], audio classification (BEATs) [16], topic modeling (BERTopic) [17], topic classification via text similarity (MPNet) [18], embedding generation (LaBSE) [19], and recommendation system. Vector data is stored in Qdrant, while job orchestration relies on a MySQL queue.
- Network Attached Storage (NAS) provides centralized object storage for podcasts, user, and school data.

4. AI Pipeline
4.1. Transcription
4.2. Context-Dependent Grammar Correction
4.3. Audio Classification
4.4. Topic Modeling
4.5. Topic Classification via Text Similarity
4.6. Embedding Generation & Retrieval
4.7. Recommendation System
- Discrete Embedding Aggregation
- Capture Multi-faceted Interests: Users often have diverse, seemingly disparate interests that would be lost in averaging. For instance, a user might enjoy both true crime podcasts and meditation content - maintaining separate embeddings preserves both preference clusters.
- Preserve Semantic Granularity: Individual embeddings retain the specific semantic fingerprints of each podcast, including genre-specific terminology, thematic elements, and stylistic characteristics that contribute to the user's preference profile.
- Enable Dynamic Weighting: The system can apply sophisticated weighting schemes based on recency, explicit user ratings, listening duration, or engagement metrics without losing the underlying preference diversity.
- 2.
- Weighted Average Embedding
- Weighted Vector Summation: Individual podcast embeddings are multiplied by their respective weight coefficients, which may be derived from listening duration, explicit ratings, recency factors, or engagement frequency metrics.
- Normalization Operations: The weighted sum undergoes L2 normalization to maintain unit vector properties, ensuring consistent similarity computation across different user profiles regardless of their listening volume.
- Dimensional Consistency: The aggregation process preserves the original embedding dimensionality of the 768 dimensions, maintaining compatibility with the pre-trained podcast embedding model.
5. Airflow Orchestration & Deployment
5.1. Overview
5.2. Podcast and User Pipelines
- Coordinate Data Pipelines: It seamlessly integrates with both relational and vector databases, allowing for efficient ETL (extract, transform, load) processes for podcast and user data.
- Handle Data Updates: It supports incremental updates, which keeps both content and user interaction data fresh without needing to re-index everything.
- Provide a Modular Structure: Its DAG (Directed Acyclic Graph) framework allows for pipelines to be broken down into a series of modular, interdependent tasks, making it easy to add new data sources or processing stages.
- Podcast Pipelines address ingestion, updates, and deletions of content. Insert DAGs begin by retrieving pending podcast entries from the job queue, then apply audio classification (BEATs), generate transcripts (FasterWhisperXXL and VAD), refine them with Gemini-2.5 flash-lite, label content with MPNet, and generate embeddings with LaBSE. These embeddings are inserted into Qdrant, synchronized with external APIs, and jobs are marked as complete. Update DAGs follow a similar path but emphasize refreshing labels, embeddings, and user profiles linked to the modified podcast. Delete DAGs ensure that obsolete embeddings are removed from Qdrant and that affected user embeddings are updated accordingly. To balance timeliness with efficiency, podcast DAGs execute in 30-minute cycles, which ensures new uploads and updates are integrated quickly without overloading resources.
- User Pipelines govern how user data influences the recommendation system. Insert DAGs process new user accounts by generating embeddings from initial listening histories or engagement data. Update DAGs refresh metadata (e.g., user country flag) and adjust embeddings to align with recent activity. Delete DAGs handle account removals or the loss of specific attributes, recomputing embeddings so user profiles remain accurate and aligned with the current dataset. Because user actions (listening, liking, following/unfollowing) directly affect recommendation quality, these DAGs run in 5-minute cycles, allowing the system to adapt almost immediately to behavioral changes.
5.3. End-to-End Workflow
6. Implementation Considerations
7. Conclusions
Author Contributions
Acknowledgments
Funding
Data Availability Statement
Conflicts of Interest
Abbreviations
| Acronym | Explanation |
| AI | Artificial Intelligence |
| ASR | Automatic Speech Recognition |
| API | Application Programming Interface |
| ANN | Approximate Nearest Neighbor (search) |
| BEATs | Bootstrap your Own Audio Transformer with Self-supervised Learning |
| BERTopic | Bidirectional Encoder Representations Topic Modeling (framework based on transformers) |
| BART | Bidirectional and Auto-Regressive Transformers |
| CER | Character Error Rate |
| CRUD | Create, Read, Update, Delete (basic database operations) |
| CSV | Comma-Separated Values |
| CTF-IDF | Class-based Term Frequency–Inverse Document Frequency |
| DAG | Directed Acyclic Graph (used in Airflow for workflow orchestration) |
| DCMI | Dublin Core Metadata Initiative |
| DSA ESR |
Digital Services Act European School Radio |
| ETL | Extract, Transform, Load (data pipeline process) |
| EU AI Act | European Union Artificial Intelligence Act |
| GDPR | General Data Protection Regulation |
| GPU | Graphics Processing Unit |
| HDBSCAN | Hierarchical Density-Based Spatial Clustering of Applications with Noise |
| HNSW | Hierarchical Navigable Small World (graph-based ANN index) |
| IR | Information Retrieval |
| LaBSE | Language-agnostic BERT Sentence Embedding |
| MDPI | Multidisciplinary Digital Publishing Institute |
| ML | Machine Learning |
| NAS | Network Attached Storage |
| NDCG | Normalized Discounted Cumulative Gain (ranking metric) |
| PUA | Private Use Area (Unicode codepoints) |
| Qdrant | Open-source Vector Database for semantic search |
| RAID | Redundant Array of Independent Disks (RAID-5 specifically used here) |
| REST | Representational State Transfer (API style) |
| TF-IDF | Term Frequency–Inverse Document Frequency |
| TREC | Text REtrieval Conference (Podcast Track) |
| UI | User Interface |
| VAD | Voice Activity Detection |
| WER | Word Error Rate |
Appendix A: Model Performance Metrics

Appendix B: List of Predefined Labels Used for Classification
- Humor
- School Dialogue, Expressions
- Cinema and Directing
- Bullying, School Safety
- Theatre and Performing Arts
- European Elections and Democracy
- Charity, Volunteering
- Christmas and Holidays
- Sports
- Castles, Fortresses, Empires
- Video Watching
- Environment, Climate, Sustainability
- Radio and Broadcasts
- Road Safety, Accidents
- Gardening and Outdoor Spaces
- Artificial Intelligence, Technology
- Dostoevsky, Greek Perspectives
- Friendship and Social Bonds
- Disability, Education, Inclusion
- Nutrition and Health
- Books and Authors
- Football, Championships, Goals
- Poetry
- Travel and Routes
- Music
- Cars and Vehicles
- Announcements
- Pets and Animals
- School Life
- Spanish Language
- Video Games, Gaming
- Women Stereotypes
- Multilingual Excerpts
- Peace and Non-Violence
- Memories and Nostalgia
- Stories, Narration, Events
- German Culture
- Mothers, Family, Parenting
- French Culture
- History
- Internet and Digital World
- Puberty
- Fashion and Style
- Composers, Singers, Performers
- Cities, Countries, Refugees
- Health, Community, COVID
- Fairy Tales and Food
Appendix C: DAG
- Read pending jobs from MySQL queue.
- Apply BEATs for audio classification.
- Generate transcript with FasterWhisperXXL (large-v2-ct2 + VAD).
- Refine transcript with Gemini-2.5-Flash (Vertex AI).
- Apply BART model for label classification.
- Generate embeddings with LaBSE (title, description, transcript).
- Upsert embeddings into Qdrant.
- Synchronize outputs with external APIs (Youth Radio / European School Radio).
- Mark jobs as executed.
- Read pending jobs from MySQL queue.
- Apply BART model for label classification.
- Apply BEATs for audio classification.
- Generate embeddings for title, description, transcript.
- Upsert embeddings into Qdrant.
- Synchronize outputs with external APIs.
- Mark jobs as executed.
- Update user embeddings for listeners of the updated podcast.
- Read pending jobs from MySQL queue.
- Delete embeddings from Qdrant.
- Mark jobs as executed.
- Update user embeddings for affected listeners.
- User Insert DAG
- Read pending jobs from database.
- Perform integrity check (ensure episode embeddings exist).
- Generate user embeddings from listening history or engagement.
- Finalize database updates and mark jobs as completed.
- Read pending jobs from database.
- Update user country flag in Qdrant metadata.
- Commit updates to database.
- Read pending jobs from database.
- Recompute embeddings to reflect attribute removals.
- Finalize updates and mark jobs as completed.
References
- Jones, R.; Zamani, H.; Schedl, M.; Chen, C.-W.; Reddy, S.; Clifton, A.; Karlgren, J.; Hashemi, H.; Pappu, A.; Nazari, Z.; Yang, L.; Semerci, O.; Bouchard, H.; Carterette, B. Current Challenges and Future Directions in Podcast Information Access. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’21), Virtual Event, 11–15 July 2021; ACM: New York, NY, USA, 2021; pp. 1554–1565. Available online: https://dl.acm.org/doi/10.1145/3404835.3462805 (Accessed on 4 October 2025).
- Jones, R.; Carterette, B.; Clifton, A.; Eskevich, M.; Jones, G. J. F.; Karlgren, J.; Pappu, A.; Reddy, S.; Yu, Y. TREC 2020 Podcasts Track Overview, arXiv, 2021. Available online: https://arxiv.org/abs/2103.15953 (Accessed on 4 October 2025).
- Ghinassi, I.; Wang, L.; Newell, C.; Purver, M. Multimodal Topic Segmentation of Podcast Shows with Pre-trained Neural Encoders. In Proceedings of the 2023 ACM International Conference on Multimedia Retrieval (ICMR ’23), Thessaloniki, Greece, 12–15 June 2023; Association for Computing Machinery: New York, NY, USA, 2023; pp. 602–606. Available online: https://dl.acm.org/doi/10.1145/3591106.3592270 (Accessed on 4 October 2025).
- Manakul, P.; Yang, M.; Gales, M. CUED_speech at TREC 2020 Podcast Summarisation Track. arXiv 2020, arXiv:2012.02535. Available online: https://arxiv.org/abs/2012.02535 (Accessed on 4 October 2025).
- Clifton, A.; Févotte, C.; Jones, R.; King, S.; Max, A.; Watanabe, S.; Zhang, Y.; Bell, P.; Gales, M.; Garimella, V.R.K.; et al. 100,000 Podcasts: A Spoken English Document Corpus. In Proceedings of the 28th International Conference on Computational Linguistics (COLING 2020), Barcelona, Spain (Online), 8–13 December 2020; International Committee on Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 5903–5917. Available online: https://aclanthology.org/2020.coling-main.519 (Accessed on 4 October 2025).
- Clifton, A.; Févotte, C.; Jones, R.; King, S.; Max, A.; Watanabe, S.; Zhang, Y.; Bell, P.; Gales, M.; Garimella, V.R.K.; et al. The Spotify Podcast Dataset. arXiv 2020, arXiv:2004.04270. Available online: https://arxiv.org/abs/2004.04270 (Accessed on 4 October 2025).
- Alexander, A.; Jones, R.; Eskevich, M.; Ferrom, S.; Pecina, P.; Jones, G.J.F. Audio Features, Precomputed for Podcast Retrieval and Information Access Experiments. In Experimental IR Meets Multilinguality, Multimodality, and Interaction (CLEF 2021); Springer: Cham, Switzerland, 2021; Lecture Notes in Computer Science, Vol. 12880; pp. 3–14. Available online: https://doi.org/10.1007/978-3-030-85251-1_1 (Accessed on 4 October 2025.
- Ghazimatin, A.; Garmash, E.; Penha, G.; Sheets, K.; Achenbach, M.; Semerci, O.; Galvez, R.; Tannenberg, M.; Mantravadi, S.; Narayanan, D. et al. PODTILE: Facilitating Podcast Episode Browsing with Auto-Generated Chapters. In Proceedings of the 2024 ACM International Conference on Information and Knowledge Management (CIKM ’24). Available online: https://arxiv.org/abs/2410.16148 (Accessed on 4 October 2025).
- Meggetto, F.; Moshfeghi, Y. Podify: A Podcast Streaming Platform with Automatic Logging of User Behaviour for Academic Research. In Proceedings of the 2023 ACM SIGIR Conference (SIGIR ’23 — Demos), 2023. Available online: https://dl.acm.org/doi/10.1145/3539618.3591824 (Accessed on 4 October 2025).
- Paraskevopoulos, G.; Tsoukala, C.; Katsamanis, A.; Katsouros, V. The Greek Podcast Corpus: Competitive Speech Models for Low-Resourced Languages with Weakly Supervised Data. In Proceedings of Interspeech 2024, pp. 3969–3973. Available online: https://arxiv.org/pdf/2406.15284v1 (Accessed on 4 October 2025).
- Kotliar, M.; Kartashov, A.V.; Barski, A. CWL-Airflow: A Lightweight Pipeline Manager Supporting Common Workflow Language. GigaScience 2019, 8(7), giz084. Available online: https://academic.oup.com/gigascience/article/8/7/giz084/5535758 (Accessed on 4 October 2025).
- Yasmin, J.; Wang, J.; Tian, Y.; Adams, B. An Empirical Study of Developers’ Challenges in Implementing Workflows as Code: A Case Study on Apache Airflow. arXiv 2024, arXiv:2406.00180. Available online: https://arxiv.org/abs/2406.00180 (Accessed on 4 October 2025).
- Ahmed, A. E.; Allen, J. M.; Bhat, T.; Burra, P.; Fliege, C. E.; Hart, S. N.; Heldenbrand, J. R.; Hudson, M. E.; Istanto, D. D.; Kalmbach, M. T., et al. Design Considerations for Workflow Management Systems in production genomics research and the clinic. Scientific Reports 2021, 11, 21680. Available online: https://www.nature.com/articles/s41598-021-99288-8 (Accessed on 4 October 2025).
- Purfview. whisper-standalone-win: Whisper & Faster-Whisper Standalone Executables. GitHub repository, 2024. Available online: https://github.com/Purfview/whisper-standalone-win (Accessed on 10 August 2025).
- Google DeepMind, "Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities," Available online: https://arxiv.org/html/2507.06261v1 (Accessed on 10 August 2025).
- Chen, S.; Wu, Y.; Wang, C.; Liu, S.; Tompkins, D.; Chen, Z.; Che, W.; Yu, X.; Wei, F. BEATs: Audio Pre-Training with Acoustic Tokenizers. In Proceedings of the 40th International Conference on Machine Learning (ICML 2023); Proceedings of Machine Learning Research; Volume 202; PMLR: Honolulu, HI, USA, 23–29 July 2023; pp. 5178–5193. Available online: https://proceedings.mlr.press/v202/chen23ag/chen23ag.pdf (Accessed on 4 October 2025).
- Grootendorst, M. BERTopic: Neural Topic Modeling with a Class-Based TF–IDF Procedure. arXiv 2022, arXiv:2203.05794. Available online: https://arxiv.org/abs/2203.05794 (Accessed on 4 October 2025).
- Song, K., Tan, X., Qin, T., Lu, J., & Liu, T.-Y. (2020). MPNet: Masked and Permuted Pre-training for Language Understanding. Available online: https://arxiv.org/abs/2004.09297 (Accessed on 4 October 2025).
- Feng, F.; Yang, Y.; Cer, D.; Arivazhagan, N.; Wang, W. Language-Agnostic BERT Sentence Embedding. arXiv 2020, arXiv:2007.01852. Available online: https://arxiv.org/abs/2007.01852 (Accessed on 4 October 2025).
- Radford, A.; Kim, J.W.; Xu, T.; Brockman, G.; McLeavey, C.; Sutskever, I. Robust Speech Recognition via Large-Scale Weak Supervision. arXiv 2022, arXiv:2212.04356. Available online: https://arxiv.org/abs/2212.04356 (Accessed on 4 October 2025).
- Ma, R.; Qian, M.; Gales, M.; Knill, K. ASR Error Correction Using Large Language Models. arXiv 2024, arXiv:2409.09554. Available online: https://arxiv.org/abs/2409.09554 (Accessed on 4 October 2025).
- Szymański, P.; Żelasko, P. WER We Are and WER We Think We Are: Skepticism on Very Low ASR Error Rates. Findings of the Association for Computational Linguistics: EMNLP 2020; Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 455–467. Available online: https://aclanthology.org/2020.findings-emnlp.295.pdf (Accessed on 4 October 2025).
- D. K. Thennal et al., “Advocating Character Error Rate for Multilingual ASR Evaluation,” *Findings of NAACL*, 2025. [Online]. Available: https://aclanthology.org/2025.findings-naacl.277.pdf.
- European Union. Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 Laying Down Harmonised Rules on Artificial Intelligence (Artificial Intelligence Act) and Amending Regulations (EC) No 300/2008, (EU) No 167/2013, (EU) No 168/2013, (EU) 2018/858, (EU) 2018/1139 and (EU) 2019/2144 and Directives 2014/90/EU, (EU) 2016/797 and (EU) 2016/798; Official Journal of the European Union, L 202, pp. 1–180, 12 July 2024. Available online: https://eur-lex.europa.eu/eli/reg/2024/1689/oj (Accessed on 4 October 2025).
- European Union. Regulation (EU) 2022/2065 of the European Parliament and of the Council of 19 October 2022 on a Single Market for Digital Services and Amending Directive 2000/31/EC (Digital Services Act); Official Journal of the European Union, L 277, pp. 1–102, 27 October 2022. Available online: https://eur-lex.europa.eu/eli/reg/2022/2065/oj (Accessed on 4 October 2025).
- European Union. Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 on the Protection of Natural Persons with Regard to the Processing of Personal Data and on the Free Movement of Such Data (General Data Protection Regulation); Official Journal of the European Union, L 119, pp. 1–88, 4 May 2016. Available online: https://eur-lex.europa.eu/eli/reg/2016/679/oj (Accessed on 4 October 2025).
- European Parliament and Council of the European Union, Regulation (EU) 2023/2854 of the European Parliament and of the Council of 13 December 2023 on harmonised rules on fair access to and use of data (Data Act), Official Journal of the European Union, L 2023/2854, Dec. 22, 2023.
- International Organization for Standardization (ISO). ISO 15836-1:2017 — Information and Documentation — The Dublin Core Metadata Element Set — Part 1: Core Elements; ISO: Geneva, Switzerland, 2017. Available online: https://www.iso.org/standard/71339.html (Accessed on 4 October 2025).







| Model | WER (%) ↓ | Human-rated Quality (1–5) | Average Latency (× real-time) | GPU Memory Usage (GB) |
|---|---|---|---|---|
| Whisper Medium | 17.8 | 3.3 | 0.85× | 7.2 |
| Whisper Large-v2 | 14.2 | 3.8 | 1.10× | 10.8 |
| Whisper Large-v3 | 12.6 | 4.1 | 1.25× | 12.6 |
| Whisper S2T | 13.5 | 3.9 | 1.05× | 9.4 |
| Whisper Turbo | 15.1 | 3.6 | 0.55× | 8.3 |
| FasterWhisperXXL | 11.0 | 4.3 | 0.90× | 11.5 |
| Model | Similarity ↓ | Description |
|---|---|---|
| Climate! Can we change? | 75.47% | The climate is changing throughout the world. We are the children of this world... |
| RADIO GIRLS-Climate change | 73.78% | Climate change: what can we do... |
| Music Radioactivity | 73.11% | A show dedicated to women... on the occasion of International Women's Day... |
| The climate...on your neck | 71.70% | The climate change is already a reality. The temperature is increasing... |
| Women's Diaries | 70.82% | Gender Equality, Sports, work, housekeeping and interpersonal relationships... |
| Model | Similarity ↓ | Description |
|---|---|---|
| M Squared | 69.27% | In today's show we will talk about the movie Good Will Hunting... |
| Music Radioactivity | 67.83% | A show dedicated to women... on the occasion of International Women's Day... |
| Women's Diaries | 65.03% | The show is dedicated to women. Through the narration... |
| Voices of History | 62.63% | The podcast “Voices of History” was created by our school’s radio team... |
| Women in education | 59.45% | Interviews with important women in education... |
| Component | Configuration | Rationale | Key Features |
|---|---|---|---|
| Episode-related DAGs | 30-minute execution cycles | Optimal balance between data freshness and resource efficiency | Batch-oriented processing, no real-time requirements, modular structure allows dynamic rescheduling |
| User-related DAGs | 5-minute execution cycles | Higher sensitivity to real-time user interactions | Near real-time adaptation to behavioral changes, maintains recommendation quality |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).