Submitted:
27 April 2023
Posted:
02 May 2023
You are already at the latest version
Abstract
Keywords:
I. Introduction
II. Article Database and Sub Databases Based on Article Database

III. Prompt-Based Search

IV. Data-Based Search

V. Loop of Excellence

VI. Conclusion
Appendix A. Methods
Appendix A.1. Creating Text-Based Article Database
Appendix A.2. Creating Figure-Based Article Database
References
- C. Angermueller, T. Pärnamaa, L. Parts, and O. Stegle, “Deep learning for computational biology,” Molecular Systems Biology, vol. 12, no. 7, p. 878, July 2016. [Online]. Available: https://doi.org/10.15252/msb.20156651. [CrossRef]
- F. Jiang, Y. Jiang, H. Zhi, Y. Dong, H. Li, S. Ma, Y. Wang, Q. Dong, H. Shen, and Y. Wang, “Artificial intelligence in healthcare: past, present and future,” Stroke and Vascular Neurology, vol. 2, no. 4, pp. 230–243, June 2017. [Online]. Available: https://doi.org/10.1136/svn-2017-000101. [CrossRef]
- I. Kononenko, “Machine learning for medical diagnosis: history, state of the art and perspective,” Artificial Intelligence in Medicine, vol. 23, no. 1, pp. 89–109, 2001. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S093336570100077X. [CrossRef]
- R. Roscher, B. Bohn, M. F. Duarte, and J. Garcke, “Explainable machine learning for scientific insights and discoveries,” IEEE Access, vol. 8, pp. 42 200–42 216, 2020. [Online]. Available: https://doi.org/10.1109/access.2020.2976199. [CrossRef]
- W. Li, T. Yang, C. Liu, Y. Huang, C. Chen, H. Pan, G. Xie, H. Tai, Y. Jiang, Y. Wu, Z. Kang, L.-Q. Chen, Y. Su, and Z. Hong, “Optimizing piezoelectric nanocomposites by high-throughput phase-field simulation and machine learning,” Advanced Science, vol. 9, no. 13, p. 2105550, Mar. 2022. [Online]. Available: https://doi.org/10.1002/advs.202105550. [CrossRef]
- Y. N. Yohei Nakajima, “Yoheinakajima/babyagi.” [Online]. Available: https://github.com/yoheinakajima/babyagi.
- Significant-Gravitas, “Significant-gravitas/auto-gpt: An experimental open-source attempt to make gpt-4 fully autonomous.” [Online]. Available: https://github.com/Significant-Gravitas/Auto-GPT.
- A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention Is All You Need,” arXiv e-prints, p. arXiv:1706.03762, June 2017. [CrossRef]
- J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” arXiv e-prints, p. arXiv:1810.04805, Oct. 2018. [CrossRef]
- A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems, F. Pereira, C. Burges, L. Bottou, and K. Weinberger, Eds., vol. 25. Curran Associates, Inc., 2012. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf.
- J. Redmon and A. Farhadi, “Yolov3: An incremental improvement,” arXiv preprint arXiv:1804.02767, 2018. [CrossRef]
- I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative Adversarial Networks,” arXiv e-prints, p. arXiv:1406.2661, June 2014.
- M. Fire and C. Guestrin, “Over-optimization of academic publishing metrics: observing goodhart’s law in action,” GigaScience, vol. 8, no. 6, May 2019. [Online]. Available: https://doi.org/10.1093/gigascience/giz053. [CrossRef]
- J. M. Gomez-Perez and R. Ortega, “Look, read and enrich. learning from scientific figures and their captions,” arXiv e-prints, vol. arXiv:1909.09070, 2019. [Online]. Available: https://ui.adsabs.harvard.edu/abs/2019arXiv190909070G. [CrossRef]
- R. Greene, T. Sanders, L. Weng, and A. Neelakantan, “New and improved embedding model,” Retrieved from https://openai.com/blog/new-and-improved-embedding-model, 2022.
- D. Kozlowski, J. Dusdal, J. Pang, and A. Zilian, “Semantic and relational spaces in science of science: Deep learning models for article vectorisation,” Scientometrics, vol. 126, no. 7, p. 5881–5910, 2021. [CrossRef]
- E. A. van Dis, J. Bollen, W. Zuidema, R. van Rooij, and C. L. Bockting, “Chatgpt: five priorities for research,” Nature, vol. 614, no. 7947, pp. 224–226, 2023. [CrossRef]
- B. Gordijn and H. t. Have, “Chatgpt: evolution or revolution?” Medicine, Health Care and Philosophy, pp. 1–2, 2023. [CrossRef]
- K. Singhal, S. Azizi, T. Tu, S. S. Mahdavi, J. Wei, H. W. Chung, N. Scales, A. Tanwani, H. Cole-Lewis, S. Pfohl, P. Payne, M. Seneviratne, P. Gamble, C. Kelly, N. Scharli, A. Chowdhery, P. Mansfield, B. Aguera y Arcas, D. Webster, G. S. Corrado, Y. Matias, K. Chou, J. Gottweis, N. Tomasev, Y. Liu, A. Rajkomar, J. Barral, C. Semturs, A. Karthikesalingam, and V. Natarajan, “Large Language Models Encode Clinical Knowledge,” arXiv e-prints, p. arXiv:2212.13138, Dec. 2022. [CrossRef]
- OpenAI, “Gpt-4 technical report,” 2023. [CrossRef]
- M. Reichstein, G. Camps-Valls, B. Stevens, M. Jung, J. Denzler, and N. Carvalhais, “Deep learning and process understanding for data-driven earth system science,” Nature, vol. 566, no. 7743, pp. 195–204, 2019. [CrossRef]
- K. T. Butler, D. W. Davies, H. Cartwright, O. Isayev, and A. Walsh, “Machine learning for molecular and materials science,” Nature, vol. 559, no. 7715, pp. 547–555, 2018. [CrossRef]
- T. Hey, K. Butler, S. Jackson, and J. Thiyagalingam, “Machine learning and big scientific data,” Philosophical Transactions of the Royal Society A, vol. 378, no. 2166, p. 20190054, 2020. [CrossRef]
- F. J. Sigworth, “A maximum-likelihood approach to single-particle image refinement,” Journal of structural biology, vol. 122, no. 3, pp. 328–339, 1998. [CrossRef]
- Q. Chen and V. Koltun, “Photographic Image Synthesis with Cascaded Refinement Networks,” arXiv e-prints, p. arXiv:1707.09405, July 2017. [CrossRef]
- C. Wang, F. Jing, L. Zhang, and H.-J. Zhang, “Content-based image annotation refinement,” in 2007 IEEE Conference on Computer Vision and Pattern Recognition, 2007, pp. 1–8. [CrossRef]
- R. Kraft and J. Zien, “Mining anchor text for query refinement,” in Proceedings of the 13th international conference on World Wide Web, 2004, pp. 666–674. [CrossRef]
- M. Zhu, P. Pan, W. Chen, and Y. Yang, “Dm-gan: Dynamic memory generative adversarial networks for text-to-image synthesis,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2019. [CrossRef]
- N. Shinn, B. Labash, and A. Gopinath, “Reflexion: an autonomous agent with dynamic memory and self-reflection,” arXiv preprint arXiv:2303.11366, 2023. [CrossRef]
- C. Lin, P.-H. Wang, Y. Hsiao, Y.-T. Chan, A. C. Engler, J. W. Pitera, D. P. Sanders, J. Cheng, and Y. J. Tseng, “Essential step toward mining big polymer data: Polyname2structure, mapping polymer names to structures,” ACS Applied Polymer Materials, vol. 2, no. 8, pp. 3107–3113, 2020. [CrossRef]
- M. Manica, C. Auer, V. Weber, F. Zipoli, M. Dolfi, P. Staar, T. Laino, C. Bekas, A. Fujita, H. Toda, S. Hirose, and Y. Orii, “An information extraction and knowledge graph platform for accelerating biochemical discoveries,” arXiv e-prints, p. arXiv:1907.08400, 2019. [CrossRef]
- P. L. Dognin, I. Melnyk, I. Padhi, C. Nogueira dos Santos, and P. Das, “DualTKB: A dual learning bridge between text and knowledge base,” arXiv e-prints, p. arXiv:2010.14660, 2020. [CrossRef]
- P. W. Staar, M. Dolfi, C. Auer, and C. Bekas, “Corpus conversion service,” in Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2018. [CrossRef]
- N. Siegel, N. Lourie, R. Power, and W. Ammar, “Extracting scientific figures with distantly supervised neural networks,” arXiv e-prints, p. arXiv:1804.02445, 2018. [CrossRef]
- W. Ammar, D. Groeneveld, C. Bhagavatula, I. Beltagy, M. Crawford, D. Downey, J. Dunkelberger, A. Elgohary, S. Feldman, V. Ha, R. Kinney, S. Kohlmeier, K. Lo, T. Murray, H.-H. Ooi, M. Peters, J. Power, S. Skjonsberg, L. L. Wang, C. Wilhelm, Z. Yuan, M. van Zuylen, and O. Etzioni, “Construction of the literature graph in semantic scholar,” arXiv e-prints, May 2018. [CrossRef]
- J. Bhatt, K. A. A. Hashmi, M. Z. Afzal, and D. Stricker, “A survey of graphical page object detection with deep neural networks,” Applied Sciences, vol. 11, no. 12, p. 5344, 2021. [CrossRef]
- Q. Wang, M. Li, X. Wang, N. Parulian, G. Han, J. Ma, J. Tu, Y. Lin, H. Zhang, W. Liu, A. Chauhan, Y. Guan, B. Li, R. Li, X. Song, Y. R. Fung, H. Ji, J. Han, S.-F. Chang, J. Pustejovsky, J. Rah, D. Liem, A. Elsayed, M. Palmer, C. Voss, C. Schneider, and B. Onyshkevych, “Covid-19 literature knowledge graph construction and drug repurposing report generation,” arXiv e-prints, Jul 2020. [CrossRef]
- M. Li, L. Cui, S. Huang, F. Wei, M. Zhou, and Z. Li, “Tablebank: A benchmark dataset for table detection and recognition,” arXiv e-prints, 2019.
- X. Zhong, E. ShafieiBavani, and A. J. Yepes, “Image-based table recognition: data, model, and evaluation,” arXiv e-prints, 2019.
- M. Li, Y. Xu, L. Cui, S. Huang, F. Wei, Z. Li, and M. Zhou, “Docbank: A benchmark dataset for document layout analysis,” arXiv e-prints, 2020. [Online]. Available: https://ui.adsabs.harvard.edu/abs/2020arXiv200601038L. [CrossRef]
- E. Schwenker, W. Jiang, T. Spreadbury, N. Ferrier, O. Cossairt, and M. K. Y. Chan, “Exsclaim!–an automated pipeline for the construction of labeled materials imaging datasets from literature,” arXiv e-prints, vol. arXiv:2103.10631, 2021. [CrossRef]
- C. Marzahl, M. Aubreville, C. A. Bertram, J. Maier, C. Bergler, C. Kröger, J. Voigt, K. Breininger, R. Klopfleisch, and A. Maier, “Exact: a collaboration toolset for algorithm-aided annotation of images with annotation version control,” Scientific Reports, vol. 11, p. 4343, 2021. [CrossRef]
- J. P. Greco and S. Danieli, “Artpop: A stellar population and image simulation python package,” The Astrophysical Journal, vol. 941, no. 1, p. 26, 2022. [CrossRef]
- S. Subramanian, L. L. Wang, S. Mehta, B. Bogin, M. van Zuylen, S. Parasa, S. Singh, M. Gardner, and H. Hajishirzi, “Medicat: A dataset of medical images, captions, and textual references,” arXiv e-prints, vol. arXiv:2010.06000, 2020. [Online]. Available: https://doi.org/10.48550/arXiv.2010.06000. [CrossRef]
- W. Hu, Y. Yang, Z. Cheng, C. Yang, and X. Ren, “Time-series event prediction with evolutionary state graph,” 2020. [CrossRef]
- X. Zhang, M. Zeman, T. Tsiligkaridis, and M. Zitnik, “Graph-guided network for irregularly sampled multivariate time series,” 2022. [CrossRef]
- Y. Anand, Z. Nussbaum, B. Duderstadt, B. Schmidt, and A. Mulyar, “Gpt4all: Training an assistant-style chatbot with large scale data distillation from gpt-3.5-turbo,” https://github.com/nomic-ai/gpt4all, 2023.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).