AESSI: An Around-Ear Silent Speech Interface for Cross-Day Online Reuse without Test-Day Calibration
Silent speech interfaces (SSIs) enable private communication without audible speech and may support people with post-stroke dysarthria. Everyday reuse requires articulation-related representations that generalize across days despite sensor repositioning and physiological changes. We present AESSI, an around-ear SSI using masked-context representation pretraining (MCRP): a student predicts teacher representations from masked time-frequency inputs to encourage robustness to recording variability. We collected 44 electrophysiological recordings from 24 participants for 25 everyday Mandarin sentences. With six participants' complete final recordings held out, AESSI achieved 92.24 percent mean accuracy without test-day calibration. AESSI exceeded the best adapted baseline by 50.57 percentage points. At least 21 days after each participant's last recording, five participants each completed 50 independently randomized online tests without calibration, achieving 98.0 percent overall accuracy. Median preprocessing and inference time in CPU replay was 31.70 ms. These results demonstrate end-to-end system operation and support online reuse without test-day calibration. Demo is included with the paper.