People often recognize spoken words more accurately and more quickly than written words. From the perspective of active inference under the free-energy principle, the present study develops a formal account of this speaking-over-writing advantage in production-modality retrieval. We modeled retrieval under spoken and written production-modality states, together with the selection of an early or late response policy according to its expected sensory consequences. University students spoke or wrote words presented on a computer screen and later judged the modality in which each word had been produced. Responses were followed by explicit visual feedback indicating whether each judgment was correct or incorrect, thereby providing an observable sensory consequence of the recognition response. Response times were implemented as early and late policies within a partially observable Markov decision process (POMDP). The model included modality-specific likelihood parameters linking response timing to feedback outcomes. Spoken words were identified more quickly and accurately than written words, and early responding was more strongly associated with correct feedback under the speaking state than under the writing state. Bayesian model selection favored the POMDP over a simpler linear model estimated using Variational Laplace, indicating that the POMDP provided a better balance of fit and complexity. These findings show how production modality, response timing, and expected sensory consequences can be represented within a common generative framework. The framework provides a basis for future extensions that model uncertainty about production modality, preserve continuous response timing, incorporate learning across trials, and compare active-inference accounts with alternative models of recognition memory.