Karpov / Gerazov | Speech and Computer | Buch | 978-3-032-37869-9 | www.sack.de

Buch, Englisch, 474 Seiten, Format (B × H): 155 mm x 235 mm

Reihe: Lecture Notes in Artificial Intelligence

Karpov / Gerazov

Speech and Computer

28th International Conference, SPECOM 2026, Ohrid, North Macedonia, September 17-18, 2026, Proceedings, Part II
Erscheinungsjahr 2026
ISBN: 978-3-032-37869-9
Verlag: Springer Nature Switzerland AG

28th International Conference, SPECOM 2026, Ohrid, North Macedonia, September 17-18, 2026, Proceedings, Part II

Buch, Englisch, 474 Seiten, Format (B × H): 155 mm x 235 mm

Reihe: Lecture Notes in Artificial Intelligence

ISBN: 978-3-032-37869-9
Verlag: Springer Nature Switzerland AG


This two-volume set LNAI16934 - 16935 constitutes the refereed proceedings of the 28th International Conference on Speech and Computer SPECOM 2026, held in Ohrid, North Macedonia, during September 17–18, 2026.

The 65 full papers included in these volumes were carefully reviewed and selected from 99 submissions. They were organized in topical sections as follows:

Part I: Automatic Speech Recognition; Speech Processing for Healthcare; Dysartric Speech Analysis; Natural Language Processing; and Processing Under-Resourced Languages.

Part II: Computational Paralinguistics; Multimodal Analysis; Speech and Language Resources; and Audio Signal Processing.

Karpov / Gerazov Speech and Computer jetzt bestellen!

Zielgruppe


Research

Weitere Infos & Material


.- Computational Paralinguistics.
.- Detecting Sarcasm from Speech using Transformer-based and Traditional Methods.
.- AMYGDALA: A Speech Foundation Model Adaptation Framework for Robust Emotion Recognition in Low-Resource Scenarios.
.- An Audio-Gated Cascade for Resource-Efficient Bimodal Emotion Recognition.
.- Effect of the Amount of Training Data in Speech Emotion Recognition.
.- Automatic Classification vs. Human Annotation of Emotions in Everyday Spoken Russian: A Case Study of the ESC Corpus.
.- The Influence of Implicit Robotic Cues on the Efficiency of Cognitive Task Performance.
.- Multimodal Analysis.
.- Understanding Pose-based Sign Language Translation: Ablation Studies.
.- Investigating Transformers-based Feature Representations for Audio-Visual Speaker Verification.
.- Sound and Colour: Acoustic and Visual Features of Vowel Perception in Mongolian in Comparison with Russian.
.- A Unified Phoneme-to-Viseme Mapping for Serbian Based on Linguistic Knowledge, Human Perception, and AV-HuBERT Visual Embeddings.
.- Gaze, Speech and Gesture of L2 Learners as Mediated by a Discourse Type: Description, Narration and Instruction.
.- The Neural Signature of Dark Humor Memes: Beta Oscillations in the Prefrontal Cortex and Temporoparietal Junction.
.- A Lightweight LSTM-based iEEG-to-Speech Decoder with Event-Level Unseen Neural Pattern Detection.
.- Talking Faces in Serbian: From Speech and Text to Blendshape Animation.
.- Speech and Language Resources.
.- Scaling Conversational Hungarian ASR: The BEA-Dialogue+ Corpus.
.- UzEnParC 1.0: A Cross-Script, Metadata-Rich Uzbek-English Parallel Corpus for Low-Resource NLP.
.- Rhythm Contributes to Personality Identification: Gender and Age Features in Australian and New Zealand Corpora.
.- Predicting Changes in L1 Vowels of Late Czech and Italian Bilinguals with French as L2.
.- Rhythmic Accommodation in English as a Lingua Franca: A Case Study of German–Italian Dialogue.
.- Hybrid Algorithm for Automatic Annotation of Modified Multiword Expressions in Russian Colloquial Speech (Based on the ORD Corpus).
.- Computer-Aided Language Learning for Arabic-Speaking Children.
.- Chinese Speakers’ Adaptability to English Accents: a Perceptual Analysis of Word Stress Patterns.
.- Pilot Classification of Macedonian Regional Speech based on Acoustic Features and Machine Learning.
.- Audio Sign.
.- Towards Streaming Neural Speech Codecs through Time-Invariant Representations.
.- AccentFlow: Continuous Accent Transfer via Accent Delta Vectors in Speaker Embedding Space.
.- StreamSV: End-to-End Streaming Speaker Verification for Short Utterances.
.- When Pseudo-Labels Go Wrong: The Importance of Label Correction in Self-Supervised Speaker Verification.
.- Applicability and Generalization of Audio Fingerprinting Methods to Speech.
.- Restoring Bone-Conducted Whispered Speech using Self-Supervised Voice Conversion with Pseudo Bone-Conduction Data Augmentation.
.- A Lightweight Voice Activity Detection Model for Scene-Aware Cochlear Implant Processing.
.- Multiple Noise Suppression via Spatiotemporal Spectral Image Inpainting.
.- Multilingual Voicemail and Answering Machine Detection via Pretrained Audio Embeddings and Gradient Boosting Machines.
.- Robust Speech Enhancement via Extended LSTM with Perceptual Contrast Stretching and Monte Carlo Dropout.



Ihre Fragen, Wünsche oder Anmerkungen
Vorname*
Nachname*
Ihre E-Mail-Adresse*
Kundennr.
Ihre Nachricht*
Lediglich mit * gekennzeichnete Felder sind Pflichtfelder.
Wenn Sie die im Kontaktformular eingegebenen Daten durch Klick auf den nachfolgenden Button übersenden, erklären Sie sich damit einverstanden, dass wir Ihr Angaben für die Beantwortung Ihrer Anfrage verwenden. Selbstverständlich werden Ihre Daten vertraulich behandelt und nicht an Dritte weitergegeben. Sie können der Verwendung Ihrer Daten jederzeit widersprechen. Das Datenhandling bei Sack Fachmedien erklären wir Ihnen in unserer Datenschutzerklärung.