Preprint
Article

This version is not peer-reviewed.

Does a Faster Transcript Make a Better Signing Avatar? Streaming Speech Recognition for a Live International Sign Avatar

Submitted:

06 October 2026

Posted:

08 October 2026

You are already at the latest version

Abstract
Real-time speech-to-sign systems such as INTERACT pass a live speech recognition transcript to a 3D avatar that plays International Sign (IS) clips from a dictionary and fingerspells words without an entry. Streaming recognition policies are compared on word error rate (WER) and latency, but their effect on the signing has, to our knowledge, not been measured. We built an open benchmark that couples streaming Whisper policies with the INTERACT text-to-sign mapping and a sequential avatar model, and ran it on the AMI meeting corpus. Even with a perfect transcript of 9.06 hours of meetings, only 23.2\% of word tokens are in the lexicon, and the deployed mapping asks for 6.36 seconds of signing per second of meeting. The demand falls below one only when unknown words are left to captions and clips play at 1.5 times speed or faster. On 16 three-minute excerpts with Whisper small, the deployed fixed 1\,s chunks released words soonest (median 0.55\,s) but had the highest WER (39.7\% against 25.9 to 27.2\%) and 2.4 times the spurious fingerspelling of the other policies, including a ``thank you'' hallucinated in pauses. Pause-based utterances were accurate but late and bursty, so bounded-lag playback dropped 54\% of their signs. With load reduction and playback at 1.5 times speed within a 4\,s lag limit, AlignAtt delivered as many correct signs as the fixed chunks, 74\% of the number with a perfect transcript, with about half as many wrong signs. WER tracked the wrong signs but not the timely delivery of correct ones. The avatar's throughput is the binding constraint, and within it an early-committing incremental policy with voice activity detection suits the avatar best.
Keywords: 
;  ;  ;  ;  ;  ;  ;  
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.