علوم رایانشی

علوم رایانشی

از ماشین تورینگ تا مبدّل‌ها: مروری سیستماتیک بر تکامل هوش مصنوعی و بُن‌انگاره‌های هوش مصنوعی مولد و مبدّل

نوع مقاله : مروری

نویسندگان
1 پژوهشکده علوم کامپیوتر، پژوهشگاه دانش های بنیادی IPM ، تهران، ایران
2 دانشجوی کارشناسی ارشد، دانشکده مهندسی برق وکامپیوتر، دانشگاه صنعتی قم، قم، ایران
3 دانشجوی کارشناسی ارشد، دانشکده مهندسی برق وکامپیوتر، واحد آشتیان، دانشگاه آزاد اسلامی، آشتیان، ایران
10.22034/csj.2025.547262.1125
چکیده
اصطلاح «هوش مصنوعی» اولین بار توسط جان مک‌کارتی در دهه ۱۹۶۰ ابداع شد و در دهه ۱۹۷۰ به عنوان شاخه‌ای مستقل از علوم کامپیوتر تثبیت گردید. در مراحل اولیه، هدف اصلی هوش مصنوعی توسعه سیستم‌هایی بود که قادر به تقلید وظایف روزمره انجام ‌شده توسط انسان‌های عادی (غیرمتخصص) بدون برنامه‌نویسی صریح باشند. از جمله نخستین تجلیات عملی این ایده می‌توان به سیستم‌های بازی مانند الگوریتم‌های شطرنج اشاره کرد. از منظر علمی، هوش مصنوعی را می‌توان ترکیبی پیشرفته از اصول بیونیک[1] (مطالعه سیستم‌های زنده برای الهام‌گیری مهندسی) و سایبرنتیک[2] (مطالعه سیستم‌های کنترل و ارتباطات) دانست. پیشرفت‌های انقلابی در قدرت محاسباتی طی سه دهه گذشته، توسعه نظریه‌های پیچیده و طیف گسترده‌ای از کاربردهای عملی در این حوزه را ممکن ساخته است. ظهور معماری مبدل‌ها[3] و مدل‌های زبانی بزرگ[4] در سال‌های اخیر نقطه عطفی در تاریخ هوش مصنوعی محسوب می‌شود. این دستاوردها نه تنها دامنه کاربردهای عملی را به شدت گسترش داده‌اند، بلکه منجر به ایجاد شکاف فزاینده‌ای بین درک عمومی از هوش مصنوعی (به عنوان فناوری‌های گفتگوی هوشمند) و تعاریف فنی آن شده‌اند. این مقاله به تحلیل محتوای پژوهش‌های معتبر بین‌المللی، سه حوزه اصلی شامل مبانی نظری و چارچوب‌های مفهومی هوش مصنوعی؛ کاربردهای مدرن و دستاوردهای کلیدی؛ چشم انداز و چالش‌های فعلی و راهکارهای پیشنهادی می‌پردازد. تحلیل‌های ارائه‌ شده در این مقاله می‌تواند سهم قابل‌توجهی در درک عمیق‌تر از وضعیت فعلی و مسیر آینده توسعه هوش مصنوعی ایفا نمایید.



1.     Bionics


2.     Cybernetics


3.     Transformer


4.     Large language model (LLM)
کلیدواژه‌ها

1.         Asimov, I. (1964). The human brain, its capacity and functions.
2.         Russell, S. J., & Norvig, P. (1995). Artificial intelligence: A modern approach;[the intelligent agent book] (pp. I-XXVIII). Prentice hall.
3.         Jouppi, N. P., Young, C., Patil, N., Patterson, D., Agrawal, G., Bajwa, R., ... & Yoon, D. H. (2017, June). In-datacenter performance analysis of a tensor processing unit. In Proceedings of the 44th annual international symposium on computer architecture (pp. 1-12).
4.         Bommasani, R. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.
5.         Chen, M. (2021). Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
6.         McCulloch, W. S., & Pitts, W. (1943). A logical calculus of the ideas immanent in nervous activity. The bulletin of mathematical biophysics, 5(4), 115-133..
7.         McCulloch, W. S., & Pitts, W. (1990). A logical calculus of the ideas immanent in nervous activity. Bulletin of mathematical biology, 52(1), 99-115.
8.         Turing, A. M. (2007). Computing machinery and intelligence. In Parsing the Turing test: Philosophical and methodological issues in the quest for the thinking computer (pp. 23-65). Dordrecht: Springer Netherlands.
9.         Turing, A. M. (1950). Mind. Mind, 59(236), 433-460.
10.     Weizenbaum, J. (1966). ELIZA—a computer program for the study of natural language communication between man and machine. Communications of the ACM, 9(1), 36-45.
11.     Nilsson, N. J. (1984). Shakey the robot.
12.     Lighthill, J. (1973). Artificial intelligence: a general survey. Science Research Council. Science Research Council (SRC), Government Report.
13.     Pearl, J. (1988). Evidential reasoning under uncertainty. In Exploring Artificial Intelligence (pp. 381-418). Morgan Kaufmann.
14.     Crevier, D. (1993). AI: the tumultuous history of the search for artificial intelligence. Basic Books, Inc..
15.     Breiman, L. (2001). Random forests. Machine learning, 45(1), 5-32.
16.     LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. nature, 521(7553), 436-444.
17.     Amodei, D., Ananthanarayanan, S., Anubhai, R., Bai, J., Battenberg, E., Case, C., ... & Zhu, Z. (2016, June). Deep speech 2: End-to-end speech recognition in english and mandarin. In International conference on machine learning (pp. 173-182). PMLR.
18.     Covington, P., Adams, J., & Sargin, E. (2016, September). Deep neural networks for youtube recommendations. In Proceedings of the 10th ACM conference on recommender systems (pp. 191-198).
19.     Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30.
20.     Thirunavukarasu, A. J., Ting, D. S. J., Elangovan, K., Gutierrez, L., Tan, T. F., & Ting, D. S. W. (2023). Large language models in medicine. Nature medicine, 29(8), 1930-1940.
21.     Bommasani, R. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.
22.     NVIDIA, N. (2023). H100 Tensor Core GPU Architecture: EXCEPTIONAL PERFORMANCE. SCALABILITY, AND SECURITY FOR THE DATA CENTER v1, 4.
23.     Bommasani, R. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.
24.     Garey, M. R., & Johnson, D. S. (1990). A Guide to the Theory of NP-Completeness. Computers and intractability, 37-79.
25.     Cormen, T. H., Leiserson, C. E., Rivest, R. L., & Stein, C. (2022). Introduction to algorithms. MIT press.
26.     Krolak, P., Felts, W., & Marble, G. (1971). A man-machine approach toward solving the traveling salesman problem. Communications of the ACM, 14(5), 327-334.
27.     Salkin, H. M., & De Kluyver, C. A. (1975). The knapsack problem: a survey. Naval Research Logistics Quarterly, 22(1), 127-144.
28.     Kershner, R. (1939). The number of circles covering a set. American Journal of mathematics, 61(3), 665-671.
29.     May, K. O. (1965). The origin of the four-color conjecture. Isis, 56(3), 346-348.
30.     McQuillan, J. M. (1974). Adaptive Routing Algorithms for Distributed Computer Networks (No. BBN2831).
31.     Robbins, S. M., & Murphy, T. E. (1949). Economics of Scheduling for Industrial Mobilization. Journal of Political Economy, 57(1), 30-45.
32.     Haessler, R. W. (1971). A heuristic programming solution to a nonlinear cutting stock problem. Management Science, 17(12), B-793.
33.     Norman, R. Z., & Rabin, M. O. (1959). An algorithm for a minimum cover of a graph. Proceedings of the American Mathematical Society, 10(2), 315-319.
34.     Mitchell, T. M. (1997). Does machine learning really work?. AI magazine, 18(3), 11-11.
35.     Bishop, C. M., & Nasrabadi, N. M. (2006). Pattern recognition and machine learning (Vol. 4, No. 4, p. 738). New York: springer.
36.     Goodfellow, I. (2016). Deep learning.
37.     Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., ... & Amodei, D. (2020). Language models are few-shot learners. Advances in neural information processing systems, 33, 1877-1901.
38.     Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019, June). Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers) (pp. 4171-4186).
39.     Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., ... & Liu, P. J. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140), 1-67.
40.     Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., ... & Scialom, T. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
41.     Manning, C. D. (2022). Human language understanding & reasoning. Daedalus, 151(2), 127-138.
42.     Jouppi, N., Kurian, G., Li, S., Ma, P., Nagarajan, R., Nai, L., ... & Patterson, D. A. (2023, June). Tpu v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings. In Proceedings of the 50th annual international symposium on computer architecture (pp. 1-14).
43.     von Zitzewitz, V. L. C. (2024). NVIDIA´ s Bet on Artificial Intelligence (Master's thesis, Universidade NOVA de Lisboa (Portugal)).
44.     Bringas, P. G., García, H. P., de Pisón, F. J. M., Flecha, J. R. V., Lora, A. T., de la Cal, E. A., ... & Corchado, E. (Eds.). (2022). Hybrid Artificial Intelligent Systems: 17th International Conference, HAIS 2022, Salamanca, Spain, September 5–7, 2022, Proceedings (Vol. 13469). Springer Nature.
45.     Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., ... & Amodei, D. (2020). Language models are few-shot learners. Advances in neural information processing systems, 33, 1877-1901.
46.     Liu, A., Feng, B., Xue, B., Wang, B., Wu, B., Lu, C., ... & Ruan, C. (1). others. 2024. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437.
47.     Liu, A., Feng, B., Xue, B., Wang, B., Wu, B., Lu, C., ... & Piao, Y. (2024). Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437.
48.     Calegari, R., Ciatto, G., Denti, E., & Omicini, A. (2020). Logic-based technologies for intelligent systems: State of the art and perspectives. Information, 11(3), 167.
49.     Davis, M., Andrews, P., Bibel, W., Robinson, A., & Siekmann, J. (2001). The Early History of Automated. Handbook of Automated Reasoning, 1, 1.
50.     Haack, S. (1974). Deviant logic: Some philosophical issues. CUP Archive.
51.     Zadeh, L. A. (1965). Fuzzy sets. Information and control, 8(3), 338-353.
52.     Rudolph, L. (2013). Qualitative mathematics for the social sciences: Mathematical models for research on cultural dynamics. Routledge.
53.     Lukasiewicz, J. (1920). On three-valued logic. Ruch filozoficzny, 5(170-171).
54.     Bayes, T. (1958). Studies in the history of probability and statistics: Ix. Thomas Bayes' essay towards solving a problem in the doctrine of chances. Biometrika, 45(3), 296-315.
55.     Robinson, J. A. (1965). A machine-oriented logic based on the resolution principle. Journal of the ACM (JACM), 12(1), 23-41.
56.     Shapiro, S. (1991). Foundations without foundationalism: A case for second-order logic (Vol. 17). Clarendon Press.
57.     Zadeh, L. A. (1965). Fuzzy sets. Information and control, 8(3), 338-353.
58.     Hegel, F. G. W. (2020). Wissenschaft der logik. BoD–Books on Demand.
59.     Hart, P. E., Nilsson, N. J., & Raphael, B. (1968). A formal basis for the heuristic determination of minimum cost paths. IEEE transactions on Systems Science and Cybernetics, 4(2), 100-107.
60.     Goldberg, A. V., & Harrelson, C. (2005, January). Computing the shortest path: A search meets graph theory. In SODA (Vol. 5, pp. 156-165).
61.     Salton, G. (1989). Automatic text processing: The transformation, analysis, and retrieval of. Reading: Addison-Wesley, 169.
62.     Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781.
63.     Han, J., Kamber, M., & Pei, J. (2012). Data mining: Concepts and. Techniques, Waltham: Morgan Kaufmann Publishers.
64.     Johnson, J., Douze, M., & Jégou, H. (2019). Billion-scale similarity search with GPUs. IEEE Transactions on Big Data, 7(3), 535-547.
65.     Ramos, J. (2003, December). Using tf-idf to determine word relevance in document queries. In Proceedings of the first instructional conference on machine learning (Vol. 242, No. 1, pp. 29-48).
66.     Domingos, P. (2012). A few useful things to know about machine learning. Communications of the ACM, 55(10), 78-87.
67.     Shi, X. H., Liang, Y. C., Lee, H. P., Lu, C., & Wang, L. M. (2005). An improved GA and a novel PSO-GA-based hybrid algorithm. Information processing letters, 93(5), 255-261.
68.     Reeves, C. R., & Rowe, J. E. (2002). Genetic algorithms—principles and perspectives: a guide to GA theory. Boston, MA: Springer US.
69.     Stützle, T., & Dorigo, M. (1999). ACO algorithms for the traveling salesman problem. Evolutionary algorithms in engineering and computer science, 4, 163-183.
70.     Teodorović, D. (2009). Bee colony optimization (BCO). In Innovations in swarm intelligence (pp. 39-60). Berlin, Heidelberg: Springer Berlin Heidelberg.
71.     Negi, G., Kumar, A., Pant, S., & Ram, M. (2021). GWO: a review and applications. International Journal of System Assurance Engineering and Management, 12(1), 1-8.
72.     Yang, X. S., & He, X. (2013). Bat algorithm: literature review and applications. International Journal of Bio-inspired computation, 5(3), 141-149.
73.     Mirjalili, S., Mirjalili, S. M., & Yang, X. S. (2014). Binary bat algorithm. Neural Computing and Applications, 25(3), 663-681.
74.     Kaveh, A., & Farhoudi, N. (2013). A new optimization method: Dolphin echolocation. Advances in Engineering Software, 59, 53-70.
75.     Salem, S. A. (2012, October). BOA: A novel optimization algorithm. In 2012 international conference on engineering and technology (ICET) (pp. 1-5). IEEE.
76.     Nasiri, J., & Khiyabani, F. M. (2018). A whale optimization algorithm (WOA) approach for clustering. Cogent Mathematics & Statistics, 5(1), 1483565.
77.     Hassanzadeh, T., & Kanan, H. R. (2014). Fuzzy FA: a modified firefly algorithm. Applied Artificial Intelligence, 28(1), 47-65.
78.     Zomaya, A. Y. (Ed.). (2006). Handbook of nature-inspired and innovative computing: integrating classical models with emerging technologies. Springer Science & Business Media.
79.     Kachenoura, A., Albera, L., Senhadji, L., & Comon, P. (2008). ICA: a potential tool for BCI systems. IEEE Signal Processing Magazine, 25(1), 57-68.
80.     Bai, Y. Y., Xiao, S., Liu, C., & Wang, B. Z. (2012). A hybrid IWO/PSO algorithm for pattern synthesis of conformal phased arrays. IEEE Transactions on antennas and Propagation, 61(4), 2328-2332.
81.     Yazdani, M., & Jolai, F. (2016). Lion optimization algorithm (LOA): a nature-inspired metaheuristic algorithm. Journal of computational design and engineering, 3(1), 24-36.
82.     Jouppi, N. P., Young, C., Patil, N., Patterson, D., Agrawal, G., Bajwa, R., ... & Yoon, D. H. (2017, June). In-datacenter performance analysis of a tensor processing unit. In Proceedings of the 44th annual international symposium on computer architecture (pp. 1-12).
83.     Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., ... & Zheng, X. (2016). {TensorFlow}: a system for {Large-Scale} machine learning. In 12th USENIX symposium on operating systems design and implementation (OSDI 16) (pp. 265-283).
84.     Zaharia, M., Xin, R. S., Wendell, P., Das, T., Armbrust, M., Dave, A., ... & Stoica, I. (2016). Apache spark: a unified engine for big data processing. Communications of the ACM, 59(11), 56-65.
85.     Jia, J. S., Lu, X., Yuan, Y., Xu, G., Jia, J., & Christakis, N. A. (2020). Population flow drives spatio-temporal distribution of COVID-19 in China. Nature, 582(7812), 389-394.
86.     Fingerhut, L. C., Miller, D. J., Strugnell, J. M., Daly, N. L., & Cooke, I. R. (2020). ampir: an R package for fast genome-wide prediction of antimicrobial peptides. Bioinformatics, 36(21), 5262-5263.
87.     Nagel, S. (2021). Report of the Editor of" The Journal of Finance" for the Year 2020. The Journal of Finance, 1019-1028.
88.     Al Banna, M. H., Taher, K. A., Kaiser, M. S., Mahmud, M., Rahman, M. S., Hosen, A. S., & Cho, G. H. (2020). Application of artificial intelligence in predicting earthquakes: state-of-the-art and future challenges. IEEE Access, 8, 192880-192923.
89.     Zeinalnezhad, M., Chofreh, A. G., Goni, F. A., Hashemi, L. S., & Klemeš, J. J. (2021). A hybrid risk analysis model for wind farms using Coloured Petri Nets and interpretive structural modelling. Energy, 229, 120696.
90.     McKinney, S. M., Sieniek, M., Godbole, V., Godwin, J., Antropova, N., Ashrafian, H., ... & Shetty, S. (2020). International evaluation of an AI system for breast cancer screening. Nature, 577(7788), 89-94.
91.     Gulshan, V., Peng, L., Coram, M., Stumpe, M. C., Wu, D., Narayanaswamy, A., ... & Webster, D. R. (2016). Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. jama, 316(22), 2402-2410.
92.     Campanella, G., Hanna, M. G., Geneslaw, L., Miraflor, A., Werneck Krauss Silva, V., Busam, K. J., ... & Fuchs, T. J. (2019). Clinical-grade computational pathology using weakly supervised deep learning on whole slide images. Nature medicine, 25(8), 1301-1309.
93.     Yang, J., Li, S., Wang, Z., Dong, H., Wang, J., & Tang, S. (2020). Using deep learning to detect defects in manufacturing: a comprehensive survey and current challenges. Materials, 13(24), 5755.
94.     Martinelli, F., Scalenghe, R., Davino, S., Panno, S., Scuderi, G., Ruisi, P., ... & Dandekar, A. M. (2015). Advanced methods of plant disease detection. A review. Agronomy for sustainable development, 35(1), 1-25.
95.     Escobar-Grisales, D., Ríos-Urrego, C. D., & Orozco-Arroyave, J. R. (2023). Deep learning and artificial intelligence applied to model speech and language in Parkinson’s disease. Diagnostics, 13(13), 2163.
96.     Du, W., & Han, Q. (2021, November). Research on application of artificial intelligence in movie industry. In 2021 International Conference on Image, Video Processing, and Artificial Intelligence (Vol. 12076, pp. 265-270). SPIE.
97.     Covington, P., Adams, J., & Sargin, E. (2016, September). Deep neural networks for youtube recommendations. In Proceedings of the 10th ACM conference on recommender systems (pp. 191-198).
98.     Lian, D., Wu, Y., Ge, Y., Xie, X., & Chen, E. (2020, August). Geography-aware sequential location recommendation. In Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining (pp. 2009-2019).
99.     Steck, H. (2013, October). Evaluation of recommendations: rating-prediction and ranking. In Proceedings of the 7th ACM conference on Recommender systems (pp. 213-220).
100.    Raca, M., Tormey, R., & Dillenbourg, P. (2014, March). Sleepers' lag-study on motion and attention. In Proceedings of the fourth international conference on learning analytics and knowledge (pp. 36-43).
101.    Boukouvala, F., Misener, R., & Floudas, C. A. (2016). Global optimization advances in mixed-integer nonlinear programming, MINLP, and constrained derivative-free optimization, CDFO. European Journal of Operational Research, 252(3), 701-727.
102.    Revanna, J. K. C., & Al-Nakash, N. Y. B. (2023). Metaheuristic link prediction (MLP) using AI based ACO-GA optimization model for solving vehicle routing problem. International Journal of Information Technology, 15(7), 3425-3439.
103.    Bazzi, S., & Sternad, D. (2020). Robustness in human manipulation of dynamically complex objects through control contraction metrics. IEEE robotics and automation letters, 5(2), 2578-2585.
104.    Markowitz, H. M. (2008). Portfolio selection: efficient diversification of investments. Yale university press.
105.    Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., ... & Hassabis, D. (2017). Mastering the game of go without human knowledge. nature, 550(7676), 354-359.
106.    Romanchenko, D., Odenberger, M., Göransson, L., & Johnsson, F. (2017). Impact of electricity price fluctuations on the operation of district heating systems: A case study of district heating in Göteborg, Sweden. Applied Energy, 204, 16-30.
107.    Chen, W., Wang, P., Ren, H., Sun, L., Li, Q., Yuan, Y., & Li, X. (2024, October). Medical image synthesis via fine-grained image-text alignment and anatomy-pathology prompting. In International conference on medical image computing and computer-assisted intervention (pp. 240-250). Cham: Springer Nature Switzerland.
108.    Chen, W., Wang, P., Ren, H., Sun, L., Li, Q., Yuan, Y., & Li, X. (2024, October). Medical image synthesis via fine-grained image-text alignment and anatomy-pathology prompting. In International conference on medical image computing and computer-assisted intervention (pp. 240-250). Cham: Springer Nature Switzerland.
109.    Pasquier, P., Ens, J., Fradet, N., Triana, P., Rizzotti, D., Rolland, J. B., & Safi, M. (2025, April). MIDI-GPT: A controllable generative model for computer-assisted multitrack music composition. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 39, No. 2, pp. 1474-1482).
110.    Tang, J., Han, X., Pan, J., Jia, K., & Tong, X. (2019). A skeleton-bridged deep learning approach for generating meshes of complex topologies from single rgb images. In Proceedings of the ieee/cvf conference on computer vision and pattern recognition (pp. 4541-4550).
111.    Maimone, A., Georgiou, A., & Kollin, J. S. (2017). Holographic near-eye displays for virtual and augmented reality. ACM Transactions on Graphics (Tog), 36(4), 1-16.
112.    Narasimhan, A., & Rao, K. P. A. V. (2021). CGEMs: A metric model for automatic code generation using GPT-3. arXiv preprint arXiv:2108.10168.
113.    Farghal, T., Shraideh, K., & Al-Omari, A. M. (2025). Evaluating Free Legal Translation Tools between Arabic and English: A Comparative Study of Google Translate, ChatGPT, and Gemini. International Journal for the Semiotics of Law-Revue internationale de Sémiotique juridique, 1-28.
114.    Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019, June). Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers) (pp. 4171-4186).
115.    Espinoza, J., Crown, K., & Kulkarni, O. (2020). A guide to chatbots for COVID-19 screening at pediatric health care facilities. JMIR public health and surveillance, 6(2), e18808.
116.    Min, M., Chen, X. B., Wang, P., Landeck, L., Chen, J. Q., Li, W., ... & Man, X. Y. (2017). Role of keratin 24 in human epidermal keratinocytes. PLoS One, 12(3), e0174626.
117.    Krishna, K., Bigham, J. P., & Lipton, Z. C. (2021, November). Does pretraining for summarization require knowledge transfer?. In Findings of the Association for Computational Linguistics: EMNLP 2021 (pp. 3178-3189).
118.    Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., ... & Amodei, D. (2020). Language models are few-shot learners. Advances in neural information processing systems, 33, 1877-1901.
119.    Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., ... & Amodei, D. (2020). Language models are few-shot learners. Advances in neural information processing systems, 33, 1877-1901.
120.    Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., ... & Zheng, X. (2016). {TensorFlow}: a system for {Large-Scale} machine learning. In 12th USENIX symposium on operating systems design and implementation (OSDI 16) (pp. 265-283).
121.    Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., ... & Chintala, S. (2019). Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32.
122.    Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., & Sutskever, I. (2023, July). Robust speech recognition via large-scale weak supervision. In International conference on machine learning (pp. 28492-28518). PMLR.
123.    Liang, Y., Wu, C., Song, T., Wu, W., Xia, Y., Liu, Y., ... & Duan, N. (2024). Taskmatrix. ai: Completing tasks by connecting foundation models with millions of apis. Intelligent Computing, 3, 0063.
124.    Kamath, U., Keenan, K., Somers, G., & Sorenson, S. (2024). Large language models: A deep dive. Springer Nature, DOI, 10, 978-3.
125.    Howard, J., & Ruder, S. (2018). Universal language model fine-tuning for text classification. arXiv preprint arXiv:1801.06146.
126.    Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., ... & Zheng, X. (2016). {TensorFlow}: a system for {Large-Scale} machine learning. In 12th USENIX symposium on operating systems design and implementation (OSDI 16) (pp. 265-283).
127.    Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., ... & Chintala, S. (2019). Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32.