Submitted:
21 July 2023
Posted:
24 July 2023
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Related Work
3. Methodology
3.1. Data Collection and Preparation
3.1.1. Corpus Collection
3.1.2. Corpus Cleaning
3.1.3. Data Preprocessing
3.1.4. Manual Annotation
3.1.5. Feature Extraction
- N-grams: These are fundamental elements in detection problems, where a sequence of n words can be denoted as a unigram (single word), bigram (two words), trigram (three words), and so on, depending on the value of n.
- BERT (Bidirectional Encoder Representations from Transformers): BERT is a vectorization method that generates high-quality vector representations of words or sentences. When used for vectorization, BERT takes in a text input and generates a fixed-size vector representation of the input that can be used for downstream tasks such as text classification or sentiment analysis. BERT embeddings have been shown to outperform traditional methods like TF-IDF and word2vec in a variety of natural language processing tasks.
- Sentiment Analysis: Flair sentiment analysis is a powerful tool for analyzing the sentiment of text data and can be useful in a wide range of applications, from social media monitoring to customer feedback analysis. We used a famous flair model to classify the sentiment , it's known for its good accuracy on general and specific domains.
- NER (Named Entity Recognition): NER is a natural language processing (NLP) task that involves identifying and classifying named entities in a text. Named entities are typically proper nouns that refer to specific people, places, organizations, dates, times, or other types of entities.
- POS (Part of Speech): POS refers to the grammatical category to which a word belongs based on its function and its relationship with other words in a sentence. There are eight main parts of speech in English: noun, pronoun, adjective, verb, adverb, preposition, conjunction, and interjection. Each part of speech has its own set of rules and characteristics that dictate how it can be used in a sentence. By understanding the different parts of speech, one can improve their ability to communicate effectively and accurately in written and spoken language.
3.2. Exploratory Data Analysis (EDA)
3.3. Classification Models
3.3.1. Baseline Models
-
SVM (Support Vector Machine)The SVM is a supervised learning model that distinguishes between two distinct classes within a high-dimensional space. This model boasts several advantages, including high speed, scalability, and the ability to identify intrusions in real-time and update training patterns dynamically. There are also variations of SVM such as kernel SVM, which allows the hyperplane to be non-linear, and nu-SVM, which allows for more flexibility in controlling the number of support vectors.
-
XGBoost (eXtreme Gradient Boosting)XGBoost is a popular and powerful machine learning algorithm used for supervised learning tasks such as regression, classification, and ranking problems. It is an implementation of the gradient boosting decision tree algorithm. Moreover, it is particularly effective when working with large datasets, and can handle missing values and noisy data. Some of the key features of XGBoost include: regularization techniques to prevent overfitting and improve generalization performance and ability to handle both sparse and dense data.
-
RF (Random Forest)The RF is an ensemble learning method that combines multiple decision trees and uses their collective output to make predictions. The name "random" forest comes from the fact that the algorithm creates a forest of decision trees, where each tree is constructed using a random subset of the training data and a random subset of the features. The general idea behind Random Forest is to use the diversity of the trees in the forest to reduce overfitting, which is a common problem in decision tree algorithms, the classifier achieves high prediction accuracy.
-
NB (Naive Bayes)NB is a classification algorithm that is based on Bayes' theorem. Despite its simplification, Naive Bayes is a powerful and popular algorithm, especially in natural language processing and text classification tasks. The Naive Bayes algorithm is relatively fast and can handle high-dimensional data with ease. However, it may not perform as well as more complex algorithms when the independence assumption is not met, or when the data has a lot of noise or missing values.
-
LSTM (Long Short-Term Memory)LSTM stands for Long Short-Term Memory, which is a type of recurrent neural network (RNN) architecture that was introduced to overcome the problem of vanishing gradients in traditional RNNs. LSTM networks have been widely used for sequence modeling tasks, such as speech recognition, natural language processing, and time series prediction. They are particularly useful when dealing with long-term dependencies in sequential data, where traditional RNNs struggle due to the vanishing gradient problem.
-
BERT (Bidirectional Encoder Representations from Transformers)BERT model is a deep learning model for natural language processing (NLP) tasks. It is a transformer-based model that uses a technique called self-attention to capture the context of a given word or phrase within a sentence or document. BERT is pre-trained on large amounts of text data using a masked language modeling (MLM) task and a next sentence prediction (NSP) task, which allows it to learn the relationships between words and phrases in a given text. BERT has achieved state-of-the-art performance in many NLP tasks and has become a standard model for natural language processing. Its ability to generate highly informative embeddings of text has made it a popular choice for a wide range of applications in industry and academia.
3.3.2. Ensemble Method
- Tokenization: The input texts are tokenized using the BERT (Bidirectional Encoder Representations from Transformers) tokenizer. This step breaks down the texts into smaller units called tokens, which are later used as input for the BERT model.
- BERT Model: The tokenized texts are passed through the BERT model. BERT is a pre-trained language model that captures contextual information from the input text. It encodes the tokens by considering their surrounding context, which helps in understanding the meaning of the text.
- Dropout Layer: The output from the BERT model is passed through a dropout layer. Dropout is a regularization technique that randomly deactivates some of the neural network units during training. This helps prevent overfitting by reducing the reliance on specific features.
- Feature FC Layer: The additional features associated with the text samples are processed through a fully connected (FC) layer. This layer applies linear transformations to the feature tensor to extract relevant information.
- Concatenation: The outputs from the BERT model and the feature FC layer are concatenated. This merging of information allows the model to incorporate both the contextualized representations from BERT and the additional features.
- Ensemble FC Layer: The concatenated output is passed through another FC layer specifically designed for the ensemble. This layer further processes the concatenated information, enabling the model to capture complex relationships between the BERT output, categorical and numerical features.
- Sigmoid Activation: Finally, a sigmoid activation function is applied to the output of the ensemble FC layer. This activation function maps the value to a probability between 0 and 1, indicating the likelihood of the input text being drug-related or not.
3.4. Performance Evaluation
4. Experiments
4.1. Data Collection
4.1.1. Data Cleaning
4.1.2. Data Preprocessing
- Map the target classes to "drug use" and "non-drug use".
- Introduce a new feature that captures the length and word count of each tweet text.
- Remove irrelevant columns such as ConversationId, Coordinates, Tcooutlinks, etc.
- Check for any missing values in the dataset.
- Remove of all tweets less than 10 words.
- Removal of all emojis and mentions.
- Removal of all whitespaces and cleaning punctuation.
- Splitting attached words: after removal of punctuation or white spaces, words can be attached. This happens especially when deleting the periods at the end of the sentences. The corpus might look like: “I need another drugdealer show”. So, there is a need to split “drugdealer” into two separate words.
- Convert the text to lowercase and remove stop words: stop words are basically a set of commonly used words in any language. By removing the words that are very commonly used in each language, we can focus only on the important words instead, and improve the accuracy of the text processing.
- Lemmatization for all words to reduce inflectional word forms to linguistically valid lemmas.
- Removing short words, where the length of the word less than 3 characters.
- Tokenize the text and pull out only the verbs, nouns and adjectives with using the part of speech tagging (POS_tag) with using python Natural Language Toolkit (NLTK) library.
- Stemming — reducing words to their root form.
4.1.3. Data Annotation
4.1.4. Feature Extraction
4.2. Exploratory Data Analysis (EDA)
4.2.1. Metadata Analysis
4.2.2. NLP Analysis
4.3. Classification Models
4.3.1. Baselines Models
4.3.2. Ensemble Method
4.4. Results
5. Conclusions
Author Contributions
Funding
Conflicts of Interest
References
- T. Alsulimani, “Social Media and Drug Smuggling in Saudi Arabia,” Journal of Civil & Legal Sciences, vol. 07, no. 02, 2018. [CrossRef]
- D. Chaffey, “Global social media statistics research summary 2023,” Smart Insights, Jun. 07, 2023.
- R. Prieto Curiel, S. Cresci, C. I. Muntean, and S. R. Bishop, “Crime and its fear in social media,” Palgrave Commun, vol. 6, no. 1, Dec. 2020. [CrossRef]
- M. Al-Otaibi, “8 held for drug dealing through social media,” Saudi Gazetti, Mar. 13, 2018.
- “UNODC/WHO Program on Drug Dependence Treatment and Care,” World Health Organization. https://www.who.int/initiatives/joint-unodc-who-programme-on-drug-dependence-treatment-and-care (accessed Jul. 07, 2023).
- E. Bigeard, N. Grabar, and F. Thiessard, “Detection and analysis of drug misuses. A study based on social media messages,” Front Pharmacol, vol. 9, no. JUL, Jul. 2018. [CrossRef]
- “Most popular social networks worldwide as of January 2023, ranked by number of monthly active users,” Statista, Jan. 2023.
- “Kaggle.” https://www.kaggle.com/ (accessed Jul. 07, 2023).
- “Substance Abuse and Mental Health Services Administration.” https://www.samhsa.gov/ (accessed Jul. 07, 2023).
- “Google Dataset Search.” https://datasetsearch.research.google.com/ (accessed Jul. 07, 2023).
- “An official website of the United States government.” https://catalog.data.gov/dataset (accessed Jul. 07, 2023).
- “UC Irvine Machine Learning Repository”, Accessed: Jul. 07, 2023. [Online]. Available: https://archive.ics.uci.edu/datasets.
- “IEEE Data Port”, Accessed: Jul. 07, 2023. [Online]. Available: https://ieee-dataport.org/.
- “Harvard Dataverse”, Accessed: Jul. 07, 2023. [Online]. Available: https://dataverse.harvard.edu/.
- “University of California, Riverside (UCR) Library Search”, Accessed: Jul. 07, 2023. [Online]. Available: https://search.library.ucr.edu/discovery/search?vid=01CDL_RIV_INST:UCR.
- “THE DATALAB”, Accessed: Jul. 07, 2023. [Online]. Available: https://thedatalab.com/.
- Sarker and G. Gonzalez, “A corpus for mining drug-related knowledge from Twitter chatter: Language models and their utilities,” Data Brief, vol. 10, pp. 122–131, Feb. 2017. [CrossRef]
- H. W. Meng, S. Kath, D. Li, and Q. C. Nguyen, “National substance use patterns on Twitter,” PLoS One, vol. 12, no. 11, Nov. 2017. [CrossRef]
- U. Lokala, R. Daniulaityte, R. Carlson, F. Lamy, and A. Sheth, “Social Media data for exploring the association between Cannabis use and Depression.” figshare, 2021.
- J. Tassone, P. Yan, M. Simpson, C. Mendhe, V. Mago, and S. Choudhury, “Utilizing deep learning and graph mining to identify drug use on Twitter data,” BMC Med Inform Decis Mak, vol. 20, Dec. 2020. [CrossRef]
- T. Nasralah, O. El-Gayar, and Y. Wang, “Social Media Text Mining Framework for Drug Abuse: An Opioid Crisis Case Analysis,” 2020. [CrossRef]
- S. J. Kim, L. A. Marsch, J. T. Hancock, and A. K. Das, “Scaling up research on drug abuse and addiction through social media big data,” J Med Internet Res, vol. 19, no. 10, Oct. 2017. [CrossRef]
- J. Xie, Z. Zhang, X. Liu, and D. Zeng, “Unveiling the Hidden Truth of Drug Addiction: A Social Media Approach Using Similarity Network-Based Deep Learning,” Journal of Management Information Systems, vol. 38, no. 1, pp. 166–195, 2021. [CrossRef]
- Roy, A. Paul, H. Pirsiavash, and S. Pan, “Automated Detection of Substance Use-Related Social Media Posts Based on Image and Text Analysis,” 2017. [Online]. Available: https://www.drugabuse.gov/drugs-abuse/commonly-abused-drugs-charts.
- F. Jenhani, M. S. Gouider, and L. Ben Said, “Hybrid system for information extraction from social media text: Drug abuse case study,” in Procedia Computer Science, Elsevier B.V., 2019, pp. 688–697. [CrossRef]
- H. Hu et al., “An ensemble deep learning model for drug abuse detection in sparse twitter-sphere,” in Studies in Health Technology and Informatics, IOS Press, Aug. 2019, pp. 163–167. [CrossRef]
- F. C. Tsai, M. C. Hsu, C. T. Chen, and D. Y. Kao, “Exploring drug-related crimes with social network analysis,” in Procedia Computer Science, Elsevier B.V., 2019, pp. 1907–1917. [CrossRef]
- S. J. Fodeh, M. Al-Garadi, O. Elsankary, J. Perrone, W. Becker, and A. Sarker, “Utilizing a multi-class classification approach to detect therapeutic and recreational misuse of opioids on Twitter,” Comput Biol Med, vol. 129, Feb. 2021. [CrossRef]
- N. Phan, M. Bhole, S. Ae Chun, and J. Geller, “Enabling real-Time drug abuse detection in tweets,” in Proceedings - International Conference on Data Engineering, IEEE Computer Society, May 2017, pp. 1510–1514. [CrossRef]
- H. Hu, P. Moturu, K. N. Dharan, J. Geller, S. Di Iorio, and H. Phan, “Deep learning model for classifying drug abuse risk behavior in tweets,” in Proceedings - 2018 IEEE International Conference on Healthcare Informatics, ICHI 2018, Institute of Electrical and Electronics Engineers Inc., Jul. 2018, pp. 386–387. [CrossRef]
- M. A. Al-Garadi et al., “Text classification models for the automatic detection of nonmedical prescription medication use from social media,” BMC Med Inform Decis Mak, vol. 21, no. 1, Dec. 2021. [CrossRef]
- H. Hu et al., “An insight analysis and detection of drug-abuse risk behavior on Twitter with self-taught deep learning,” Comput Soc Netw, vol. 6, no. 1, Dec. 2019. [CrossRef]
- T. K. Mackey and J. Kalyanam, “Detection of illicit online sales of fentanyls via Twitter,” F1000Res, vol. 6, 2017. [CrossRef]
- T. K. Mackey, J. Kalyanam, T. Katsuki, and G. Lanckriet, “Twitter-based detection of illegal online sale of prescription opioid,” Am J Public Health, vol. 107, no. 12, pp. 1910–1915, Dec. 2017. [CrossRef]
- Y. Fan, Y. Zhang, Y. Ye, X. Li, and W. Zheng, “Social media for opioid addiction epidemiology: Automatic detection of opioid addicts from Twitter and case studies,” in International Conference on Information and Knowledge Management, Proceedings, Association for Computing Machinery, Nov. 2017, pp. 1259–1267. [CrossRef]
- Safaa. S Al Dhanhani, “Framework for Analyzing Twitter to Detect Community Suspicious Crime Activity,” Academy and Industry Research Collaboration Center (AIRCC), Jan. 2018, pp. 41–60. [CrossRef]
- T. Ding, W. K. Bickel, and S. Pan, “Multi-View Unsupervised User Feature Embedding for Social Media-based Substance Use Prediction,” 2017.
- J. Li, Q. Xu, N. Shah, and T. K. Mackey, “A machine learning approach for the detection and characterization of illicit drug dealers on instagram: Model evaluation study,” J Med Internet Res, vol. 21, no. 6, Jun. 2019. [CrossRef]
- J. Tassone, P. Yan, M. Simpson, C. Mendhe, V. Mago, and S. Choudhury, “Utilizing deep learning and graph mining to identify drug use on Twitter data,” BMC Med Inform Decis Mak, vol. 20, Dec. 2020. [CrossRef]
- S. Al Amin et al., “Data Driven Classification of Opioid Patients Using Machine Learning-An Investigation,” IEEE Access, vol. 11, pp. 396–409, 2023. [CrossRef]
- J. T. Prieto et al., “The detection of opioid misuse and heroin use from paramedic response documentation: Machine learning for improved surveillance,” J Med Internet Res, vol. 22, no. 1, 2020. [CrossRef]
- Aubeer Smith, “23 essential Twitter statistics to guide your strategy in 2023,” Feb. 2023.
- “NLTK Library.” https://www.nltk.org/index.html (accessed Jul. 07, 2023).
- K. Sahoo, A. K. Samal, J. Pramanik, and S. K. Pani, “Exploratory data analysis using python,” International Journal of Innovative Technology and Exploring Engineering, vol. 8, no. 12, pp. 4727–4735, Oct. 2019. [CrossRef]
- Kulkarni and A. Shivananda, Natural Language Processing Recipes. Apress, 2019. [CrossRef]
- “Wosom.” https://wosom.ai/ (accessed May 26, 2023).









| Ref | Year | Algorithm | Feature Selection | SN | Dataset Size | Performance Metric |
|---|---|---|---|---|---|---|
| [29] | 2017 | DT, RF SVMs, NB | String2WordVector, TF-IDF | 300 tweets | P= 0.748 R= 0.757 F= 0.746 | |
| [24] | 2017 | CNN | Image feature learning with CNN, Textual feature learningwith Doc2Vec | 100,500 posts | Acc= 0.9 F= 0.75 | |
| [30] | 2018 | SVM, CNN | Word2Vec | 3M tweets | Acc= 0.865 R= 0.886 F1= 0.866 | |
| [36] | 2018 | NA | Content features, sentiment analysis, user profile | 10% of random tweets | ||
| [6] | 2018 | NB, RF, Simple Logistic | Brown clustering, Word2Vec | Doctissimo website Forum | 119,562 messages | P= 0.778 R= 0.772 F= 0.773 |
| [25] | 2019 | J48, LR, Libsvm for SVM and NB | named entity (NE), semantic links (SL) and lexical features (LF) |
1M tweets | P= 0.95 | |
| [31] | 2021 | SVM, RF, Gaussian NB, Shallow NN, KNN, CNN, BiLSTM | BERT | 16,443 tweets | F1= 0.95 | |
| [26] | 2019 | CNN, SVM, RF, NB | Word2Vec, Glove | NA | Acc= 0.857 (ML) P= 0.846 (CNN) R=0.891 (ML) F1= 0.862 (ML) | |
| [38] | 2019 | DT, RF, SVM, RNN-LSTM | Text | 12,857 posts | F1= 0.95 | |
| [33] | 2017 | Biterm Topic Model (BTM) | Text | 28,711 tweets | ||
| [32] | 2019 | SVM, Naive Bayes, CNN, LSTM |
Tf, Tf-idf, Word2Vec |
1,794 tweets | Acc= 0.865 R= 0.886 F1= 0.866 | |
| [34] | 2017 | BTM | URL | 619,937 tweets | ||
| [21] | 2020 | Text mining | TF-IDF | 10,000 tweets | P= 0.941 R= 0.966 F= 0.953 Acc= 0.928 | |
| [37] | 2017 | SVD, LDA, D-DM, D-DBOW | User feature embedding | 22M posts | AUC=0.86 for predicting tobacco use, AUC=0.81 for alcohol use and AUC=0.84 for illicit drug use |
|
| [35] | 2017 | LLGC | BOW, users’ profiles | 19,722 tweets, 2,312 users | Acc= 0.8336 F1= 0.8215 | |
| [39] | 2020 | SVM, XGBoost and CNN-based classifier | word2vec embedding | 3,696,150 tweets | Acc= 0.823 P= 0.893 Recall= 0.784 F1= 0.835 AUC=0.91 | |
| [17] | 2017 | data and language models | word representation, n-gram | 267,215Twitter posts | ||
| [40] | 2022 | AdaBoost, LR, SVM, XGB, RF, LSTM, ANN and CNN | Dataset attributes, tabular data | Database | 37,127 distinct cases | |
| [41] | 2020 | RF,KNN, SVM and L1-regularized LR | NLP features | free-text narratives, impressions, list of medications |
54,359 trip reports | AUC=0.94 |
| [23] | 2021 | Similarity Network-based Deep Learning (SINDEL) | word embedding and a network of words | Drugs-Forum | 27,154 posts | F1=0.767 |
| [28] | 2020 | RF, SVM, BiLSTM, BERT | Woed2Vec | 5,523,588 tweets | F1=0.71 for the Pain-misuse class, and 0.79 for the Recreational-misuse class | |
| [27] | 2019 | DT, NB, LR, SVM, RF, k-NN | Personal Feature Set and Social feature Set | Criminal Warehouse | 5,780 records with 4,561 unique individuals | F1=0.622 |
| [18] | 2017 | Text classifier and analytical approach | sentiment scores, substanceuse variables, and underage variables | 79,848,992 tweets |
| No. | Algorithm | Features | Metrics | ||||
|---|---|---|---|---|---|---|---|
| F1-Score | Accuracy | Precision | Recall | AUC | |||
| 1 | SVM | Textual | 0.9017 | 0.8995 | 0.8819 | 0.9225 | 0.8995 |
| 2 | XGBoost | Textual | 0.9000 | 0.8982 | 0.8803 | 0.9206 | 0.8984 |
| 3 | RF | Textual | 0.8693 | 0.8662 | 0.8459 | 0.8940 | 0.8664 |
| 4 | NB | Textual | 0.8216 | 0.8087 | 0.7641 | 0.8886 | 0.8094 |
| 5 | LSTM | Textual | 0.8939 | 0.9074 | 0.8790 | 0.9094 | 0.9077 |
| 6 | BERT | Textual | 0.9005 | 0.9018 | 0.8922 | 0.9133 | 0.9041 |
| 7 | Ensemble method | Textual Categorical Numerical |
0.9112 | 0.9166 | 0.8953 | 0.9341 | 0.9133 |
| Precision | Recall | F1-score | Support | |
|---|---|---|---|---|
| Non-drug use | 0.92 | 0.88 | 0.90 | 12794 |
| Drug use | 0.88 | 0.92 | 0.90 | 9618 |
| accuracy | 0.90 | 22412 | ||
| macro avg | 0.90 | 0.90 | 0.90 | 22412 |
| weighted avg | 0.90 | 0.90 | 0.90 | 22412 |
| Precision | Recall | F1-score | Support | |
|---|---|---|---|---|
| Non-drug use | 0.93 | 0.89 | 0.91 | 2028 |
| Drug use | 0.89 | 0.93 | 0.91 | 1972 |
| accuracy | 0.91 | 4000 | ||
| macro avg | 0.91 | 0.91 | 0.91 | 4000 |
| weighted avg | 0.91 | 0.91 | 0.91 | 4000 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).