Tag
GigaChat Audio 10B is an audio-native LLM built on GigaChat 3.1 Lightning, integrating a Conformer speech encoder and modality adapter for audio question answering, temporal grounding, and tool use.
Details a method to run a 13 million parameter ASR Conformer model directly on a microcontroller, highlighting advances in edge AI deployment.
This paper introduces GigaAM Multilingual, a Conformer encoder pre-trained on 2M hours of audio with a HuBERT-style objective, addressing data scarcity for Central Asian languages (Kazakh, Kyrgyz, Uzbek). It employs cluster-level data balancing and domain-aware sampling to outperform strong baselines like Whisper Large v3 on target languages.
This paper proposes a multimodal framework that jointly improves Automatic Speech Recognition (ASR) and Dialect Identification (DID) for Indian languages, using a Bottleneck Encoder and RoBERTa with a gating mechanism. Evaluated on eight languages with 33 dialects, it achieves 81.63% DID accuracy and reduces CER/WER to 4.65%/17.73%.
This paper introduces HRVConformer, a hybrid Convolution-Transformer architecture for classifying neonatal hypoxic-ischemic encephalopathy directly from raw heart rate signals, achieving an AUC of 83.23% and outperforming baseline models like ResNet50 and Transformer.
This paper proposes an end-to-end Conformer-based neural decoder for intracortical speech decoding from a participant with ALS, achieving a 23.80% character error rate without any external language model. It demonstrates that meaningful character-level decoding is possible in a fully end-to-end framework.