Preprint
Article

This version is not peer-reviewed.

Understanding Media Sentiments Towards Central Bank Digital Currency: Evidence from India

Submitted:

29 June 2026

Posted:

01 July 2026

You are already at the latest version

Abstract
This study investigates how the media portrays India’s Central Bank Digital Currency (CBDC) using sentiment analysis and topic modelling techniques. Using a dataset of 915 media articles published between October 2022 and January 2026, we apply a natural language processing (NLP) model to classify sentence-level sentiment and then use topic modelling to identify favourable and unfavourable themes of discussion. The results suggest that sentiment towards CBDC adoption is generally positive in the Indian media, with coverage focusing on financial inclusion, cross-border payments, and the modernization of digital currencies. The narratives expressing concern are few and focus on macroeconomic and currency risks related to CBDCs.
Keywords: 
;  ;  ;  

1. Introduction

The rapid development of digital financial technologies has transformed how economies conduct transactions, store money, and perceive it. As part of this broader digital shift, central banks worldwide are increasingly assessing the advantages and disadvantages of Central Bank Digital Currencies (CBDCs). CBDC is generally regarded as a digital version of sovereign-backed money, issued and regulated by the central bank, with the same legal standing as physical currency but existing solely in electronic form.
Globally, around 137 central banks are exploring CBDCs.1 CBDCs have been identified by the International Monetary Fund (IMF) and the Bank for International Settlements (BIS) as a potential tool to promote financial inclusion, enhance the robustness of payments systems, and strengthen monetary sovereignty amid digital challenges. Several advanced economies, including Sweden and the Euro Area, have launched pilot projects, while countries like the Bahamas and Nigeria have already rolled out operational CBDCs. The People’s Bank of China has taken the lead in one of the few major CBDC projects worldwide with its e-CNY initiative.
The Reserve Bank of India (RBI) announced a pilot for the digital rupee, with the first wholesale pilot launched in November 2022 and a retail pilot in December 2022. India has also established one of the most rapidly expanding digital payment systems globally, with the Unified Payments Interface allowing low-cost, real-time retail transactions. While cash continues to play an important role in the payment system, its share of total payments has been gradually moderating alongside the rapid expansion of digital payment options.2 The value of advancing digitalization while also addressing the costs and inefficiencies associated with cash is a strong rationale for CBDC. In addition, policymakers see the digital rupee improving payment efficiency and lowering transaction costs, increasing financial inclusion, and facilitating cross-border transactions on the lines of stablecoins through bilateral or multilateral corridors, programmability, and protecting monetary sovereignty from private digital currencies.3
Societal acceptance of CBDCs is necessary for a successful rollout. The public’s trust and acceptance to a significant extent is shaped by discourse, media coverage, and public opinion. Therefore, understanding how the Indian media treats CBDCs can provide vital insights into how society perceives the transition from cash to digital currency. In addition, analysing the themes that are portrayed in positive or concerning terms can help identify the issues that receive greater public attention and may require policy focus. Such insights are important for policymakers because trust, transparency, and communication are key determinants of public willingness to adopt CBDCs.
To further throw light on this space, the current study conducts a systematic analysis of Indian media discourse on CBDCs over a four-year period. Using a dataset of media articles about CBDCs, the study employs Financial Bidirectional Encoder Representations from Transformers (FinBERT), a transformer-based language model fine-tuned for financial text. By applying this method, the study creates a monthly time series of average sentiment scores, enabling examination of trends and changes in media perceptions of CBDCs. Further, the study uses BERTopic (a topic modelling technique based on transformer embeddings and clustering) to find the major themes talked about CBDC in the public sphere. Together, these methods provide insights not only into the tone of media coverage but also into the specific issues and narratives shaping public discourse on CBDCs.
The remainder of the paper is organized as follows. Section II reviews the literature, Section III presents stylized facts, Section IV describes the data and methodology, Section V reports the results, and Section VI concludes.

2. Literature Review

CBDCs are a form of digital currency that, unlike privately issued cryptocurrencies, are issued and regulated by a central authority. While cryptocurrencies are known for their speculative price bubbles and high volatility (Bordo and Levin, 2017), CBDCs keep parity with the national currency. They are backed by central bank assets (Guley and Koldovskyi, 2023). Scholars have argued that CBDCs could reshape the financial system as highlighted by Buckley (2024) and McLaughlin (2021) as a potential to strengthen financial stability. Using a dynamic stochastic general equilibrium model, Tong and Jiayou (2021) found that CBDCs may reduce systemic risk while allowing the public to hold them without exposure to liquidity or credit risks. Feyen et al. (2021) and Auer et al. (2022) emphasize that central bank credibility—especially on issues such as privacy, inclusion, and technological efficiency—plays a key role in building citizen trust.
With the rise of NLP, researchers have begun to examine how CBDCs are discussed in both official and public communications. Scharnowski (2022) analyses central bank speeches primarily from the United States and the euro area, showing that central bank communication significantly influences investor reactions in financial markets. Meanwhile, Alonso-Robisco and Santiago Carbó (2023) employ language models such as BERT and ChatGPT to evaluate sentiment in policy communications across multiple regions, including Europe and North America, and find that the models’ outputs closely align with expert-coded sentiment. In addition, Hofmann et al. (2025) uses large language models to construct indices of central bank and media sentiment toward CBDCs across 15 major economies, including India. Their findings reveal a notable divergence between central bank sentiment and media coverage, with central bank communication exerting a stronger influence on media framing than the reverse.
Advances in sentiment analysis tools have also made domain-specific evaluations more precise. For example, FinBERT has been shown to outperform general-purpose models such as VADER at detecting negative sentiment in financial texts (Kim et al., 2023). Advances in sentiment analysis tools have also made domain-specific evaluations more precise. FinBERT, for example, has been shown to outperform general-purpose models such as VADER in financial sentiment analysis by better capturing the contextual nuances of financial language across all sentiments (Bansal et al., 2025). Topic modelling methods have also advanced, with BERTopic offering greater flexibility and customization than traditional methods such as Latent Dirichlet Allocation (Sy et al., 2024). Studies have demonstrated that this approach tends to generate more coherent, meaningful, and detailed topics in news or social media collections compared to LDA (most commonly used topic modelling methods) (Grootendorst, 2022; Chen et al., 2023).
Despite growing global interest, research on how the Indian media covers CBDCs remains limited. Most existing studies on India focus on technical design, pilot projects, and public opinion surveys (RBI, 2022; Auer et al., 2020; Kappal et al., 2025; Eichengreen et al., 2022; Kaur et al., 2025). They lack a comprehensive analysis of news content at a large scale. This lack of detailed media analysis makes it hard to understand the narratives and sentiment patterns that shape public discussion in India.
To address this research gap, the study combines FinBERT-based sentiment analysis with BERTopic-based topic modelling to conduct a theme-level analysis of CBDC coverage in Indian news media. To our knowledge, this is among the early empirical studies examining Indian media sentiment on CBDCs at a theme-specific level. It offers both thematic and longitudinal insights that can help shape policy communication strategies and enhance public engagement. This article contributes in two ways. First, it quantifies media sentiments and identifies major themes on CBDCs in India, serving as a companion to the technical design and policy. Second, this paper showcases the value of using advanced NLP methods, such as FinBERT and BERTopic, for analysing financial discourse and enhances the debate in economic and policy literature on the use of computational linguistics.

3. Stylised Facts

To motivate the discussion, we use publicly available RBI data on the volume and value of CBDC transactions to highlight key stylized facts on the evolving use of CBDCs in India.
Based on weekly statistics published by the Reserve Bank of India on the amount of CBDC in circulation, the retail pilot shows a steady expansion over time. In December 2023, the total amount of CBDC-R in circulation stood at about ₹1.03 lakh crore, increasing to around ₹13.49 lakh crore by August 2025 (Figure 1), reflecting the gradual expansion of user access and merchant participation during the pilot phase. The fall in circulation since September 2025 largely reflects operational and technical adjustments associated with the migration to an updated system architecture.
Furthermore, many initiatives have been taken from time to time with regards to CBDC such as launch of the pilot, its distribution through UPI ecosystem and its use-case in food subsidy. Key policy announcements and pilot developments that shaped the evolution of India’s digital rupee initiatives are summarised in Table 1. These milestones provide important context for understanding the phases of CBDC rollout and the corresponding developments in circulation and policy experimentation.
Taken together, these patterns point to a clear growth in both the volume and value of retail CBDC transactions, signalling an encouraging trajectory of public adoption and operational expansion.

4. Data and Methodology

We collected news articles from major Indian financial newspapers published between October 2022 and January 20264 from the ProQuest database using keywords: ‘CBDC’, ‘central bank digital currency’, ‘e₹’, ‘digital rupee’, ‘retail CBDC’, or ‘wholesale CBDC’ in the heading of the articles. As a result, 915 relevant articles were obtained, and each article was linked to a structured data frame containing a unique ID and related metadata. Standard pre-processing steps were undertaken (for further information, please refer to Annexure A2).

4.1. Sentiment Classification Framework

To evaluate the sentiment of media discussions surrounding CBDC, this study employs a transformer-based language model, FinBERT which is developed by Araci (2019). It is a domain-specific adaptation of BERT (Bidirectional Encoder Representations from Transformers) trained on financial corpora such as Reuters TRC2 (Text Research Collection) and the Financial PhraseBank (Malo et al., 2014). FinBERT effectively captures subtle sentiment variations in economic and financial narratives. The model is implemented using the Hugging Face Transformers library (model identifier: yiyanghkust/finbert-tone). (For further information on the FinBERT architecture, please refer to Appendix B1).
Sentiment labels are assigned according to the highest predicted probability obtained from the FinBERT logit functions following Kirtac and Germano (2025), such that:
Label = argmax{ P(Neu), P(Pos), P(Neg),}
This classification represents a qualitative interpretation of sentiment, as each sentence is assigned to the category with the highest predicted probability.5
While labelling provides a discrete classification, sentiment scoring offers a continuous quantification of sentiment intensity. Following Zhang (2005), sentiment scores are derived to capture the direction and intensity of each sentence’s sentiment. This score is computed as the difference between the model’s predicted probabilities for positive and negative sentiments:
s i t = P i , t ( P o s ) P i . t ( N e g )   [ 1,1 ]
where s i t is the sentiment score for the sentence i on day t is. A positive score indicates a higher likelihood of positive sentiment relative to negative sentiment, while a negative score reflects a stronger negative sentiment. A value close to zero suggests either neutral sentiment or a balance between positive and negative tones. To construct a time-series measure of media sentiment index, sentence-level continuous scores are averaged at the monthly frequency as
  S e n t i m e n t   I n d e x m = 1 N m i = 1 N m s i , m
where:
Sentiment Index = Average continuous sentiment score for month m
Nm = Total number of sentences in month m
si,m = Continuous sentiment score of sentence i occurring in month m

4.2. BERTopic-Derived Themes Framework

To identify the main themes in Indian media discussions on CBDCs, we used an unsupervised topic modelling technique. We applied BERTopic to both positive and negative sentences. BERTopic is a transformer-based model that generates topic descriptions using textual embeddings and cluster-based classification (Halimeh et al., 2023). Given the diverse language used in CBDC-related news coverage, BERTopic is well-suited for capturing distinct themes and the underlying drivers of media narratives.
Additionally, we visualise the inter-topic distance map. It represents the interconnections among topic clusters generated by BERTopic modeling (for reference on the BERTopic architecture, see Appendix B2). Each circle represents a topic, with its spatial arrangement reflecting the degree of similarity. Themes represented by circles that are close together show greater similarity, while those farther apart indicate less similarity.

5. Empirical Results and Discussion

The findings from text mining and sentiment analysis reveal a clear temporal pattern in Indian media coverage of CBDC. As shown in Figure 2, the volume of published articles increased sharply in late 2022, with coverage rising from 39 articles in October 2022 to 79 in November and 76 in December 2022. This surge coincides with major policy announcements, including the release of the RBI’s CBDC Concept Note on 7 October 2022, followed by the launch of the Digital Rupee–Wholesale (e₹-W) pilot on 1 November 2022 and the Digital Rupee–Retail (e₹-R) pilot on 1 December 2022 (Table 1). After this initial surge, media coverage gradually declined through 2023, falling to around 14–21 articles per month during much of the year, and remained relatively subdued during most of 2024. However, attention began to rise again toward late 2024 and throughout 2025, with coverage increasing from 34 articles in January 2025 to 54 articles by December 2025. This renewed interest coincided with evolving discourse and parallel policy developments, including the announcement of CBDC distribution through non-bank payment applications in April 2024, the launch of subsidy-related pilots in 2025, and the introduction of the CBDC Retail Sandbox on 8 October 2025, among other things (Table 1).

5.1. Sentiment Index

Using the FinBERT model, we conducted sentiment analysis on the collected corpus to examine the tone of CBDC-related media discourse. The dataset comprised 23,829 sentences extracted from 915 unique articles, each classified as positive, neutral, or negative as discussed above. The results indicate that a substantial share of sentences are positive in tone (around 53%), focusing on objective reporting while also highlighting potential benefits such as improvements in payment efficiency, financial inclusion, and technological innovation. Neutral sentiment accounts for approximately 29% of the sentences, suggesting that a considerable portion of media coverage remains informational. The remaining 18% of sentences exhibit negative sentiment, reflecting concerns about implementation challenges, privacy issues, and other risks associated with digital currency adoption. The temporal evolution of these sentiments is illustrated in Figure 3, which presents the monthly proportions of positive, neutral, and negative sentences from October 2022 to June 2025. Although the total volume of CBDC-related news items declined after peaking in early 2023, the relative distribution of sentiment categories remained broadly stable. Positive sentiment consistently dominates across most months, with occasional fluctuations. Neutral sentiment remains at moderate levels throughout, with intermittent spikes likely corresponding to phases of heightened information reporting or policy clarification. In contrast, negative sentiment stays relatively low and stable over time, with only minor increases in certain periods, indicating limited scepticism or critical discourse in media narratives.
The sentiment dynamics of CBDC-related news coverage exhibit a relatively stable but moderately positive pattern over time. As shown in Figure 4, the monthly average continuous sentiment score (defined above) generally remains above 0.6 throughout the sample period, indicating an overall favourable tone in media discussions of CBDC. Despite some short-term variations, the dotted trend line indicates a slight upward trajectory in sentiment over time, implying that media narratives surrounding CBDC have gradually become more positive as the pilot programmes expanded, and policy developments progressed.

5.2. Topics of Relevance for CBDC

Of the 12,629 positive-labelled sentences, BERTopic identified 27 topics based on statistical patterns in the data (Figure C1 in Appendix C). These topics represent the model’s data-driven clustering rather than a predefined conceptual framework. For interpretability, only the most coherent and substantively meaningful topics were retained6 and manually grouped into 10 broader themes (Table 2). The results indicate that news articles generally approach CBDCs positively across several dimensions, including adoption, public trust, improvements in cross-border payments, financial inclusion, economic digitization, and environmental sustainability.
Figure 5 presents a network representation of the relationships among the key CBDC-related themes identified in the analysis. Each node represents a theme, while the connecting edges indicate instances within the same sentence or article. The thickness of the edges reflects the number of shared mentions between themes, with thicker edges representing stronger overlaps and more frequent co-occurrence in the media discourse7. The network structure shows that economic digitization and interoperability occupy a central position, with strong connections to themes such as improvements in cross-border payments, financial inclusion, and consumer empowerment. This suggests that media narratives often frame CBDC as part of a broader digital transformation of the financial system that simultaneously improves payment infrastructure and expands financial access. In addition, the relationship among innovation and efficiency, monetary sovereignty, and technological advancement in payment systems suggests that discussions of technological advancement in payment systems are often connected to broader strategic considerations related to national monetary autonomy. Overall, the network shows that CBDC-related themes tend to cluster rather than appear in isolation, highlighting the multidimensional nature of policy discussions surrounding the digital rupee.
Similarly, BERTopic was applied to the 3,574 negative sentences identified earlier. The model detected nine topics (Figure C2 in Appendix C), reflecting data-driven clustering rather than a predefined conceptual structure. For interpretability, only the substantively meaningful and coherent topics were retained and manually grouped into five broader themes based on relevance (Table 3). The results highlight concerns about systemic risks in crypto, technological disruption from artificial intelligence, vulnerabilities in monetary systems, and the evolving position of the US dollar.
Figure 6 presents the co-occurrence network of negative themes associated with CBDC discussions. The network shows that macroeconomic risks occupies a central position, forming strong connections with safe-haven Assets, geopolitical instability, and currency dynamics. This suggests that media discussions of CBDC-related risks are often embedded within broader narratives of global economic uncertainty and financial market instability. In addition, technological disruption is closely connected with crypto-related systemic threats, reflecting how emerging technologies and developments in the digital asset ecosystem are often framed as potential sources of systemic financial risk. The connection between macroeconomic risks and technological disruption further suggests that technological change is sometimes discussed in the context of broader macro-financial vulnerabilities.
Beyond media narratives, the findings also carry broader macroeconomic implications. Predominantly positive media discourse around CBDC can accelerate public acceptance of the digital rupee, increasing its usage as a medium of exchange and potentially reducing reliance on cash. This shift may enhance the effectiveness of monetary policy transmission by improving the speed and completeness with which policy signals such as changes in interest rates or liquidity conditions are reflected in household and firm behaviour. Additionally, greater adoption of CBDC can strengthen financial formalization which would expand the observable economic base and thereby improving the efficiency in resource allocation. Positive narratives around financial inclusion and cross-border payment efficiency suggest that CBDCs are increasingly framed as instruments supporting the broader digital transformation of the financial system. Furthermore, widespread trust in CBDC can help preserve monetary sovereignty in the face of competing private digital currencies, reinforcing the role of the central bank in anchoring inflation expectations, and maintaining macroeconomic stability. At the same time, concerns about technological disruption and financial stability underscore the importance of carefully managing macro-financial risks in the evolving digital currency ecosystem.

6. Conclusions

The results show that Indian media discourse has been predominantly positive towards CBDCs during the study period. Key themes linked to positive sentiment include financial inclusion, smoother cross-border payments, and innovation. It emphasizes optimism about CBDCs as a tool to democratize access to finance and improve efficiency in international transactions. Conversely, pessimistic sentiments are associated with concerns about technological disruptions and the possible decline of traditional safe-haven assets. These findings indicate that although the overall narrative supports CBDCs, the discourse is divided, reflecting both the opportunities and the risks. The findings suggest that policymakers should continue to emphasize transparency, regulatory clarity, and technological safeguards in the development of CBDCs to maintain public trust and address concerns about financial stability and technological disruption. Strengthening communication around the objectives and safeguards of CBDC initiatives may further support informed public discourse and facilitate smoother adoption.
However, the analysis is limited to English-language coverage and formal news-based sources and therefore reflects media narratives in that setting rather than the full spectrum of public opinion or policy outcomes. Future research could extend this analysis by incorporating additional languages and other data sources, such as social media discussions, to better understand how CBDC narratives evolve.

Author Contributions

Conceptualization, Paras; methodology, Shreya Gupta; software, Shreya Gupta; validation, Shreya Gupta and Paras; formal analysis, Shreya Gupta; investigation, Shreya Gupta; resources, Paras; data curation, Shreya Gupta; writing—original draft preparation, Shreya Gupta and Paras; writing—review and editing, Shreya Gupta and Paras; visualization, Shreya Gupta; supervision, Paras; project administration, Paras. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Acknowledgments

The authors are grateful to colleagues for their helpful comments and suggestions on earlier drafts of this paper. The authors also acknowledge the permission granted by the RBI for publication of this research. Any remaining errors are the responsibility of the authors.

Conflicts of Interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Abbreviations

The following abbreviations are used in this manuscript:
CBDC Central Bank Digital Currency
RBI Reserve Bank of India
NLP Natural Language Processing
FinBERT Financial Bidirectional Encoder Representations from Transformers
BERTopic Bidirectional Encoder Representations from Transformers for Topic Modeling
e₹-W Digital Rupee–Wholesale
e₹-R Digital Rupee–Retail

Appendix A. Data Sources and Pre-Processing

Appendix A.1. Newspaper Sources

The dataset comprises articles from major Indian financial newspapers, providing extensive coverage of economic policy, financial markets, and developments in digital finance. The newspapers included in the sample are Business Standard, Financial Express, Mint, The Economic Times, The Hindu Business Line, and The Times of India. These outlets were selected due to their consistent reporting on monetary policy, banking, and financial technology developments relevant to CBDCs. All articles were retrieved from the ProQuest database using keyword-based searches.

Appendix A.2. Pre-Processing Strategy

We develop a comprehensive text pre-processing pipeline that cleans and standardizes data to enhance readability and remove unnecessary clutter. It includes methods such as removing irrelevant elements—hyperlinks, HTML tags, emojis, acronyms, non-text symbols, extra whitespace, and duplicate articles — and removing stop words using the stopword package from Python’s Natural Language Toolkit. A corpus is a systematically collected body of text used for linguistic and computational analysis, like a dataset in empirical research. In this study, the corpus consists of newspaper articles. Each article serves as a document, and each document is composed of sentences and words, which are processed to extract themes, sentiment, and insights into the discussions surrounding CBDCs. Additionally, we combine article headlines with their body text, then split the combined text into individual sentences. After tokenizing the articles into sentences, we eliminate noisy ones—those with fewer than three words. Finally, words within each sentence are lemmatized to their root forms.

Appendix B. FinBERT and BERTopic

The models are applied at the sentence level rather than the article level, as financial news articles often contain multiple, sometimes conflicting sentiments, especially in complex policy areas such as CBDC adoption. Sentence-level analysis allows for more precise detection of sentiment variations and reduces the risk of sentiment dilution that may arise when analysing entire articles. Prior studies (Vaswani et al., 2023; Jehnen et al., 2025) also show that sentence-level classification improves sentiment-detection accuracy in financial texts, where tone can vary across segments.

Appendix B.1. FinBERT Architecture

Each cleaned sentence in the dataset is tokenized using the FinBERT tokenizer, which converts words or subwords into token identifiers from a fixed vocabulary. These tokens are then encoded as 768-dimensional contextual vectors suitable for transformer-based processing. Special tokens—[CLS]8 at the beginning and [SEP]9 at the end of each vector are added to delineate input token embeddings. Each token embedding is the sum of its token, positional, and segment embeddings, thereby incorporating both semantic and positional information. These embeddings are passed through FinBERT’s 12-layer transformer encoder, which resembles the architecture of BERT-base. Each encoder block consists of two sub-layers: a multi-head self-attention mechanism and a feed-forward neural network, ensuring stable learning. The multi-headed self-attention mechanism enables each token embedding to attend to all other embeddings in the sequence, thereby capturing semantic dependencies. At the same time, the feed-forward network applies non-linear transformations that enhance contextual understanding. Lower encoder blocks (1–4) capture lexical patterns, middle encoder blocks (5–8) capture semantic dependencies, and higher encoder blocks (9–12) capture abstract features relevant to sentiment prediction.
After processing all 12 encoder layers, the contextual embedding of the [CLS] token is used as the sentence representation. This token vector is passed through a fully connected linear layer (multiplied by a weight matrix (W) and adds a bias (b)) to produce three unnormalized sentiment scores (logits) corresponding to positive, neutral, and negative classes, as expressed by the relation:
l o g i t s = W . ( h [ C L S ] ) + b  
where W R 3×768 b R 3
Initial values of W and b are small random values that are learned during training by optimizing against labelled data with known sentiment (positive, negative, or neutral). The logit values are then transformed into normalized class probabilities using the softmax function:
P ( c l a s s i ) = e l o g i t i j e l o g i t j   ,   i { p o s i t i v e ,   n e g a t i v e ,   n e u t a l }
yielding the probability for each sentiment category— P(Neg) : Probability of negative sentiment, P(Neu): Probability of neutral sentiment and P(Pos) : Probability of positive sentiment

Appendix B.2. BERTopic Architecture

The detailed process involves: (1) Embedding sentences using the MiniLM-L12-v2 Sentence-BERT model to generate dense semantic representations (768-dimension vectors for BERT); (2) reducing high dimensional embeddings with Uniform Manifold Approximation and Projection to maintain semantic structure and improve clustering efficiency; (3) clustering using Hierarchical Density-Based Spatial Clustering of Applications with Noise, which automatically detects topic groups and isolates noise (groups into clusters that will become topics), thereby enhancing the representational power of document features. CountVectorizer is applied with an n-gram range of (1, 3), allowing the model to capture individual terms and short phrases; and (4) the final step in BERTopic is extracting keywords for each of the clusters/class/topics. To do this, BERTopic uses class-based Term Frequency–Inverse Document Frequency (c-TF-IDF) to extract distinctive keywords for each topic.
All sentences in a cluster are concatenated, and the term frequency (TF), which measures how often a term appears in a cluster/class, is computed. The inverse class frequency (ICF), which penalizes standard terms across the entire corpus, is then determined by dividing the total number of clusters by the number of clusters where the term appeared, followed by logarithmic scaling. This weighting method ensures that terms that occur frequently within a single cluster but rarely in others receive greater weight, while standard terms shared across multiple clusters are down-weighted. Conceptually, this implies that a term’s importance is inversely related to its frequency across clusters, ensuring that only topic-specific terms stand out in the final representation.
The c-TF-IDF for a term t in class c is given by:
TF(t,c)=Number of times term t appears in class c
DF(t) = Number of classes containing term t
Given a term t and the total number of classes N, the inverse class frequency ICF(t) is calculated as:
I C F ( t ) = log ( N 1 + D F ( t ) )
Finally, the c-TF-IDF weight of a term t in class c is expressed as:
c-TF-IDF(t,c) = TF(t,c)×ICF(t)
The computed c-TF-IDF values indicate the relative importance of a term t within a topic cluster compared to all other clusters. A higher value indicates that the term more uniquely represents that cluster, whereas lower values suggest the term is more common across multiple clusters and less informative (Grootendorst, 2022).

Appendix C

Figure C1. Positive Topic Distance Visualisation (displaying only relevant topics).
Figure C1. Positive Topic Distance Visualisation (displaying only relevant topics).
Preprints 220754 g0a1
Figure C2. Negative Topic Distance Visualisation (displaying only relevant topics).
Figure C2. Negative Topic Distance Visualisation (displaying only relevant topics).
Preprints 220754 g0a2

Notes

1
Atlantic Council Central Bank Digital Currency Tracker. Available online: https://www.atlanticcouncil.org/cbdctracker/ (accessed on 01-01-2026).
2
Reserve Bank of India. (2025, June). Payment Systems Report. Available online: https://rbi.org.in/Scripts/PublicationsView.aspx?id=23436 (accessed on 01-01-2026).
3
Sankar, T.R. (2025). Stablecoins – Do They Have a Role in the Financial System’. Keynote address delivered at the Mint Annual BFSI Conclave 2025 on December 12, 2025, in Mumbai.
4
The complete list of sources is provided in Appendix A1.
5
For instance, if the predicted probabilities are 0.76 (positive), 0.21 (neutral), and 0.03 (negative), the sentence is classified as positive. In cases where two classes receive equal highest probabilities (e.g., 0.40 for positive and 0.40 for negative, with 0.20 for neutral), the assignment follows the argmax rule based on the predefined label ordering, which, in the Hugging Face Transformers implementation, is specified as neutral (0), positive (1), and negative (2). In the rare instance where probabilities are equal across all classes, the assignment follows the argmax rule, which selects the first class according to the predefined label ordering; in this implementation, this corresponds to the neutral category.
6
Not all extracted topics are reported, as some consist primarily of stopwords or connectors (e.g., “a”, “an”, “the”) and do not convey meaningful semantic content.
7
For example, if a sentence contains keywords such as “access” (linked to financial inclusion) and “data” (linked to privacy), the sentence is assigned to both themes, creating a connection between them. When such keyword-based co-occurrences appear repeatedly across multiple sentences or articles, the link between the corresponding themes becomes stronger.
8
Classify token [CLS] is a special token used in NLP and machine learning (ML) models, particularly those based on the Transformer architecture.
9
A separate token [SEP] is a special token used in NLP and machine learning ML models to mark the separation between different segments of text.

References

  1. Alo.Alonso-Robisco, A.; Carbó, J. M. Analysis of CBDC narrative by central banks using large language models. Finance Research Letters 2023, 58, 104643. [Google Scholar] [CrossRef]
  2. Araci, D. Finbert: Financial sentiment analysis with pre-trained language models. arXiv 2019, arXiv:1908.10063. [Google Scholar]
  3. Auer, R.; Böhme, R. (2020). The technology of retail central bank digital currency. BIS Quarterly Review, March.
  4. Auer, R.; Frost, J.; Gambacorta, L.; Monnet, C.; Rice, T.; Shin, H. S. Central bank digital currencies: motives, economic implications, and the research frontier. Annual review of economics 2022, 14(1), 697–721. [Google Scholar] [CrossRef]
  5. Bansal, S.; Singh, B. K.; Jain, M. K. An analysis of different sentiment analysis models on financial text using transformer. In Proceedings of the International Conference on Artificial Intelligence and Emerging Human Systems (ICAIEHS); Atlantis Pres, 2025. [Google Scholar]
  6. Bordo, M. D.; Levin, A. T. (2017). Central bank digital currency and the future of monetary policy. National Bureau of Economic Research (No. w23711). [CrossRef]
  7. Buckley, R. P. Implications for the dollar of central bank digital currencies. Law and Contemporary Problems 2024, 87(2), 69–89. [Google Scholar]
  8. Chen, W.; Rabhi, F.; Liao, W.; Al-Qudah, I. Leveraging State-of-the-Art Topic Modeling for News Impact Analysis on Financial Markets: A Comparative Study. Electronics 2023, 12(12), 2605. [Google Scholar]
  9. Eichengreen, B.; Gupta, P.; Marple, T. (2022). A central bank digital currency for India? National Council of Applied Economic Research, Working Paper No.: 138.
  10. Feyen, E.; Frost, J.; Natarajan, H.; Rice, T. (2021). What does digital money mean for emerging market and developing economies?. In The Palgrave Handbook of Technological Finance. 217-241. Cham: Springer International Publishing.
  11. Grootendorst, M. BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv 2022, arXiv:2203.05794. [Google Scholar]
  12. Guley, A.; Koldovskyi, A. Digital Currencies of Central Banks (CBDC): Advantages and Disadvantages. Financial Markets, Institutions and Risks 2023, 7(4), 54–66. [Google Scholar] [CrossRef]
  13. Halimeh, H.; Caron, M.; Müller, O. (2023). Early depression detection with Transformer Models: Analyzing the relationship between Linguistic and Psychology-based features. Proceedings of the 56th Hawaii International Conference on System Science, 3377-3386. [CrossRef]
  14. Hofmann, B.; Tang, X.; Zhu, F. (2025). Central bank and media sentiment on central bank digital currency: an international perspective (No. 1279). Bank for International Settlements.
  15. Jehnen, S.; Ordieres-Meré, J.; Villalba-Díez, J. FinTextSim: Enhancing Financial Text Analysis with BERTopic. arXiv 2025, arXiv:2504.15683. [Google Scholar]
  16. Kappal, J. M.; Parashar, N.; Mujumdar, M.; Sharma, R.; Nair, P. G. An empirical study on the acceptance of CBDC by Indian bankers: A structural equation modelling approach. Asian Economic and Financial Review 2025, 15(4), 630–647. [Google Scholar] [CrossRef]
  17. Kaur, H.; Mehta, K.; Mago, M. Understanding factors influencing CBDC usage intentions among Indian households: Applying an extended UTAUT model. NMIMS Management Review 2025, 33(2), 87–102. [Google Scholar] [CrossRef]
  18. Kim, W.; Spörer, J. F.; Handschuh, S. Analyzing FOMC minutes: Accuracy and constraints of language models. arXiv 2023, arXiv:2304.10164. [Google Scholar]
  19. Kirtac, K.; Germano, G. (2025). Large language models in finance: Estimating financial sentiment for stock prediction. Available at SSRN 5166656.
  20. Malo, P.; Sinha, A.; Korhonen, P.; Wallenius, J.; Takala, P. Good debt or bad debt: Detecting semantic orientations in economic texts. Journal of the Association for Information Science and Technology 2014, 65(4), 782–796. [Google Scholar]
  21. McLaughlin, T. Two paths to tomorrow’s money. Journal of Payments Strategy and Systems 2021, 15(1), 23–36. [Google Scholar] [CrossRef]
  22. Reserve Bank of India. Annual Report of the RBI for the Year 2024-25. 2025. Available online: https://www.rbi.org.in/Scripts/AnnualReportPublications.aspx?year=2025 (accessed on 02-02-2026).
  23. Reserve Bank of India. Annual Report of the RBI for the Year 2023-24. 2024. Available online: https://www.rbi.org.in/Scripts/AnnualReportPublications.aspx?year=2024 (accessed on 02-02-2026).
  24. RBI. Concept Note on Central Bank Digital Currency. October 2022. Available online: https://rbidocs.rbi.org.in/rdocs/PublicationReport/Pdfs/CONCEPTNOTEACB531172E0B4DFC9A6E506C2C24FFB6.PDF (accessed on 02-02-2026).
  25. Scharnowski, S. Central bank speeches and digital currency competition. Finance Research Letters 2022, 49, 103072. [Google Scholar] [CrossRef]
  26. Sy, C. Y.; Maceda, L. L.; Flores, N. M.; Abisado, M. B. Unsupervised machine learning approaches in NLP: a comparative study of topic modeling with BERTopic and LDA. International Journal of Intelligent Systems and Applications in Engineering 2024, 12(21s), 3276–3283. [Google Scholar]
  27. Tong, W.; Jiayou, C. A study of the economic impact of Central Bank Digital Currency under global competition. China Economic Journal 2021, 14(1), 78–101. [Google Scholar] [CrossRef]
  28. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. (2017). Attention is All You Need. Neural Information Processing Systems. Advances in neural information processing systems, 30.
  29. Zhang, C. Dynamic Asset Pricing: Integrating FinBERT-Based Sentiment Quantification with the Fama--French Five-Factor Model. arXiv 2025, arXiv:2505.01432. [Google Scholar]
Figure 1. CBDC-R India (Amount outstanding in ₹ crore). Source: Author’s calculations based on RBI data (Currency in Circulation - Historical Data. Available online: https://rbi.org.in/Scripts/BS_CurrencyCirculationExtractDetails.aspx (accessed.
Figure 1. CBDC-R India (Amount outstanding in ₹ crore). Source: Author’s calculations based on RBI data (Currency in Circulation - Historical Data. Available online: https://rbi.org.in/Scripts/BS_CurrencyCirculationExtractDetails.aspx (accessed.
Preprints 220754 g001
Figure 2. Article Corpus Distribution (Frequency). Source: Author’s calculations.
Figure 2. Article Corpus Distribution (Frequency). Source: Author’s calculations.
Preprints 220754 g002
Figure 3. Sentiments in CBDC News Coverage (Per cent). Source: Author’s calculations.
Figure 3. Sentiments in CBDC News Coverage (Per cent). Source: Author’s calculations.
Preprints 220754 g003
Figure 4. Monthly Average Continuous Sentiment Score (Index). Source: Author’s calculations.
Figure 4. Monthly Average Continuous Sentiment Score (Index). Source: Author’s calculations.
Preprints 220754 g004
Figure 5. Positive Theme Co-occurrence Network. Source: Author’s calculations.
Figure 5. Positive Theme Co-occurrence Network. Source: Author’s calculations.
Preprints 220754 g005
Figure 6. Negative Theme Co-occurrence Network. Source: Author’s calculations.
Figure 6. Negative Theme Co-occurrence Network. Source: Author’s calculations.
Preprints 220754 g006
Table 1. Key CBDC Announcements in India.
Table 1. Key CBDC Announcements in India.
Date Event
1st Feb 2022 Announcement of Central Bank Digital Currency (CBDC) in the Union Budget 2022–23.
7th Oct 2022 Release of the CBDC Concept Note explaining design, objectives, and architecture.
1st Nov 2022 Launch of Digital Rupee – Wholesale (e₹-W) Pilot for government securities settlement.
1st Dec 2022 Launch of Digital Rupee – Retail (e₹-R) Pilot in select cities and banks.
5th April 2024 RBI announces plan to expand e₹ distribution through non-bank payment apps (UPI ecosystem).
8th October 2025 RBI launches CBDC Retail Sandbox for fintech innovation and experimentation.
15th February 2026 CBDC-based food subsidy (Public Distribution System) pilot launched in Gujarat.
Source: RBI; and Press Information Bureau.
Table 2. Topic Clusters and Associated Positive Themes in CBDC News Coverage.
Table 2. Topic Clusters and Associated Positive Themes in CBDC News Coverage.
Topic No. Themes Keyword example
0,16,19 Public Trust, Acceptance, and Institutional Preparedness adoption, acceptance, fintech, future pilots
1,13 Legal and Regulatory Framework regulation, regulatory, compliance, trust, safety
3,22 Cross-Border Payments Improvement cross-border payments, settlement, global, payments
4, 6, 14 Economic Digitization and Interoperability digital economy, digital infrastructure, modernization, interoperability
5, 17, 23, 24 Financial Inclusion and Consumer Empowerment unbanked, financial inclusion, promoting financial inclusion, access
7,8,9, Innovation and efficiency innovation, progress, fintech
10, 25 Competition Competition, digital competition, digital competition act,
11 Illicit Activity Reduction anti-money laundering, AML, KYC
12 Monetary Sovereignty assets, stablecoin, monetary sovereignty
26 Environmental Sustainability Climate, risk, sustainable finance, climate change mitigation
Source: Author’s calculations.
Table 3. Topic Clusters and Associated Negative Themes in CBDC News Coverage.
Table 3. Topic Clusters and Associated Negative Themes in CBDC News Coverage.
Topic No. Themes Keyword example
1 Geopolitical Instability and Currency Dynamics Ukraine, war, geopolitical, China, dollar
2 Macroeconomic Risks Inflation, recession
4 Technological Disruption Generative AI, artificial intelligence, privacy risk
7 Safe-Haven Assets Digital gold, investing, and emerging markets
8 Crypto-related Systemic Threats Crypto, risk, threat, macroeconomic risk
Source: Author’s calculations.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings