Introduction

The incorporation of cutting-edge technologies into the realm of second/foreign language (L2/FL) education has seen remarkable growth over the past few decades1,2. Among these innovative tools, artificial intelligence (AI) has emerged as a promising instrument for FL learners to enhance academic achievements3,4. Markedly upgraded from traditional AI, generative AI (GenAI) underpins Large Language Models and machine learning to precisely understand context, recognize errors, provide real-time feedback, and engage in creative communication, thereby rendering highly personalized and anthropomorphized services5,6. The adaptability, inventiveness, and natural language interaction attributes of GenAI satisfy language learners’ developmental needs and learning preferences7.

Given these advantages, GenAI-driven technologies have permeated multiple aspects of FL pedagogy, including essay scoring, grammar checking, and language practice8. It has garnered considerable interest from scholars and researchers, encouraging more research to understand the effectiveness of GenAI on FL learning9,10. In response, numerous empirical studies have demonstrated GenAI as a promising tool for enhancing L2 learners’ motivation11, and improving their academic performance in speaking12,13 and writing6,14. However, focusing dominantly on GenAI effectiveness cannot answer a paramount question proposed by Granić and Marangunić15—how FL learners accept and use GenAI tools. Among the existing research, there is a dearth of studies specifically centering on student users’ acceptance of GenAI integration in L2 learning, among which exploring its potential determinants is even few and far between. This research gap needs to be addressed as unveiling learners’ acceptance status of GenAI is foundational for gaining insights into its potential influence on language educational practices, and thus in turn, benefits other stakeholders such as scholars, educators, technology developers and policymakers16,17,18.

Focusing on Chinese FL students in the tertiary setting, this study adopts a verified effective model, the Unified Theory of Acceptance and Use of Technology (UTAUT19), as the theoretical foundation to systematically probe into the status quo of GenAI acceptance and utilization20,21. Moreover, the present study incorporates potential variables that may influence GenAI acceptance to complement the theoretical gaps and practical phenomena, including users’ demographic backgrounds, emotions, AI literacy, and AI self-efficacy. Theoretically, the original UTAUT model does not incorporate affective and cognitive constructs. The interplay between variables such as emotions and UTAUT’s core determinants are underexplored, leaving the psychological and cognitive mechanism of AI acceptance partially explained. Practically, the disconnect between GenAI tool developers and the educational end-users tends to reveal a clear applied gap. The findings of this study will generate a profile of target user portrait, offering actionable intelligence for both EdTech companies in product positioning and for universities in selecting and implementing GenAI tools that truly meet their students’ psychological and pedagogical needs. Thus, through innovative exploration of these factors, the study may contribute to the expansion of the UTAUT model into a broader context, investigating if and to what degree users’ demographic backgrounds, emotions, AI literacy, and AI self-efficacy impact the acceptance of GenAI among Chinese FL students in higher education. In this regard, the findings of the study enrich the existing body of literature on GenAI acceptance and offer valuable empirical insights into its mechanism.

Literature review

Theoretical background

The study proposed the research model by using the UTAUT model as the theoretical basis, as it is considered a verified and widespread model for evaluating users’ acceptance of new technologies20,21. Furthermore, emotions, AI literacy and AI self-efficacy were incorporated as external variables, since numerous studies validated the extensibility of the UTAUT model with these three individual elements, which can significantly shape individuals’ perceptions of technology22,23,24. Another reason for choosing these three as independent variables is that this research tends to complement the UTAUT model by capturing the affective reactions (emotions), necessary cognitive foundations (AI literacy), and individual capacity beliefs (AI self-efficacy), thereby offering a more holistic prediction of L2 student GenAI acceptance. Nevertheless, comprehensive studies that concurrently examine the interrelationships among technology acceptance, emotions, AI literacy and AI self-efficacy, along with their potential implications, are still scarce, let alone in the context of FL learners. This study seeks to bridge this gap by expanding the UTAUT model into a bigger picture, and exploring its relationship with other constructs.

GenAI acceptance in FL education

Technology application in language learning offers a new field of study, especially in relation to how learners’ attitudes and acceptance of technology affect the process and experiences of language acquisition25. GenAI acceptance in FL refers to the extent to which learners and educators adopt and effectively use the technologies in the L2 learning context26,27. Empirical studies showed that the acceptance and utilization of GenAI can substantially impact students’ psychological factors, such as motivation, confidence, well-being and learning engagement, and their intention to use these tools for English learning26,28,29.

The present study adopts the UTAUT model as the theoretical framework19. The UTAUT has been considered a highly effective and reliable framework as it synthesizes eight existing user technology acceptance models: the Innovation Diffusion Theory30; the Theory of Reasoned Action31; the Social Cognitive Theory32; the Technology Acceptance Model33; the Theory of Planned Behavior34; the Model of PC Utilization35; the Motivational Model36; a model Combining the Technology Acceptance Model and the Theory of Planned Behavior37. Within this model, a given technology’s Actual Usage (AU) may be predicted by its Behavioral Intention (BI; the extent to which a user’s intend to utilize the system) and the existence of Facilitating Conditions (FC; the extent to which an individual believes that technical and organizational assistance for system use19). Additionally, BI could be further explained by several factors, including Performance Expectancy (PE; the extent to which an individual’s faith in the system’s effectiveness to improve work performance), Effort Expectancy (EE; the extent of ease associated with utilizing the system) and Social Influence (SI; the extent to which a person believes that significant others or important individuals support the implementation of the new system19). Besides, technology acceptance and utilization would be moderated by user demographics (e.g., gender, age), experience, and cultural factors18,38.

As the application of the UTAUT model has been instrumental in comprehending technology acceptance in educational settings, its scope is expanding to explore students’ responses to GenAI applications21,29,39. Based on this theory, some research has examined the willingness and behavioral intention to use GenAI among students from different fields, such as medicine40 and business41, but few related to FL students. Therefore, the current study aims to embed the UTAUT model in the FL context, and proposes the following hypotheses based on prior research:

H1a

Performance Expectation positively predicts Behavioral Intention.

H1b

Effort Expectancy positively predicts Behavioral Intention.

H1c

Social Influence positively predicts Behavioral Intention.

H1d

Facilitating Condition positively predicts Actual Usage.

H1e

Behavior Intention positively predicts Actual Usage.

Emotions in FL education

The relationship between emotions and L2 education has long been a crucial field of study that aims to comprehend the emotional and engagement aspects of language acquisition42. Positive emotions such as enjoyment are found to be associated with better academic performance, learning engagement and more additional learning opportunities43,44,45, while negative emotions such as anxiety can hinder learners’ engagement, reduce motivation and affect learning effectiveness46,47,48. Similar to traditional L2 education, incorporating technology into English classrooms also evokes students’ emotions2,49,50. Despite negative emotions such as fear, embarrassment, and boredom were also commonly reported in the AI-mediated L2 classes50, most studies have asserted that innovative technology could arouse L2 learners’ positive emotions such as enjoyment, excitement and motivation51,52,53. Owing to its personalized and adaptive features, AI technologies can mitigate anxiety and pressure, especially for speaking performance54,55.

In addition, previous research has shown that emotions affect users’ technology acceptance, beliefs and behavior2,56,57,58. Likewise, positive and negative emotions have distinctly different impacts on technology usage. Positive emotions such as pleasure and happiness are reported to be positively correlated with use intention36, attitude59, and acceptance60. Conversely, attitude57 and intention to utilize technology19 are adversely correlated with unpleasant emotions such as anxiety. Hence, the following hypotheses are proposed based on the previous research:

H2

Emotions serve as a predictor of FL students’ Actual Usage.

AI literacy in FL education

As an important branch of digital literacy, AI literacy refers to “the capability to accurately identify, effectively utilize, and critically assess AI-related products while adhering to ethical standards”61,62. Various academics have relatively different perspectives on what AI literacy is. Ng et al. (2021) suggest four aspects, namely use and apply, know and understand, detect AI, and AI ethics; while Wang et al.’s61 four-dimensional theoretical framework includes usage, awareness, evaluation, and ethics. In the setting of language acquisition, users with proficient AI literacy exhibit better learning outcomes and are more willing to communicate63,64. It also showed that FL learners’ AI literacy is closely associated with their attitudes towards AI-assisted learning. The degree to which a student understands and uses information technology is strongly correlated with their attitude toward it, whereas their degree of dread of it is inversely correlated65.

Research on the interplay between AI literacy and acceptance is nascent and marked by divergent findings. While some studies affirm that AI literacy enhances perceptions of usefulness and acceptance66, others report that lower AI literacy was associated with a higher propensity to use AI for academic tasks67, challenging a straightforward positive correlation. Moreover, within the specific domain of foreign language (FL) education, little research has examined this interplay from the student perspective. Existing research demonstrated that AI literacy might influence instructors’ adoption of technology68. A more optimistic attitude and a greater willingness to accept AI products may result from increased AI literacy66,69. It may be inferred that the greater a user’s level of GenAI literacy, the more willing they will be to accept it. Therefore, the study proposes the following hypothesis:

H3

AI literacy is positively associated with FL students’ Actual Usage.

AI self-efficacy in FL education

Self-efficacy is a prominent concept in education, including L2 learning26,70,71. Instead of a fixed trait, self-efficacy is a dynamic construct that varies across contexts72,73. This concept has been transformed into a term known as AI self-efficacy, which describes an individual’s overall confidence in their ability to use and interact with AI2. Generally, scholars examine the acceptance or self-efficacy of other information technology applications as references to evaluate users’ AI self-efficacy. Examples include computer self-efficacy74, Internet self-efficacy75, robot use self-efficacy76 and technology acceptance19. As an established construct in educational technology, AI self-efficacy demonstrates a significant positive correlation with students’ learning experiences, academic performance, and level of satisfaction61,77. However, there is a dearth of studies specifically centering FL learners’ AI self-efficacy. Most studies focus on the overall concept of self-efficacy in L2 learning. Numerous research found that high self-efficacy contributes to building foundational language skills and confidence, motivating learners’ willingness to communicate and engage in language practice, and enhancing language production and usage78,79,80.

Self-efficacy is a frequently integrated external variable in technology acceptance frameworks, such as the Technology Acceptance Model81, and many studies have demonstrated a close connection between these two factors82,83. Concretely, technology self-efficacy produces an effect on key antecedents of technology acceptance, such as perceived cognitive effect (ease of use84), remote work effectiveness85, perceived usefulness86,87, perceived ease of use88, and actual usage89. Nevertheless, most studies have predominantly emphasized computer self-efficacy, rather than keeping pace with newer technologies such as (Gen)AI84,90, let alone in the FL context. This represents a significant gap that the current study aims to address. Based on the prior literature, H4 is formulated as follows:

H4

AI self-efficacy is positively associated with FL students’ Actual Usage.

Furthermore, previous research indicates that respondents’ gender, study level, university level, region, and language may impact their acceptance of technology and its association with emotions, AI literacy, and self-efficacy18,19,89. As such, this study proposed the following hypothesis:

H5a

Gender serves as a moderator in the above-mentioned H1–H4 relationships.

H5b

Study level serves as a moderator in the above-mentioned H1–H4 relationships.

H5c

University level serves as a moderator in the above-mentioned H1–H4 relationships.

H5d

Region serves as a moderator in the above-mentioned H1–H4 relationships.

H5e

Major serves as a moderate in the above-mentioned H1–H4 relationships.

Given all this, the theoretical framework examined in this study, along with the relevant hypotheses, is depicted in Fig. 1.

Fig. 1
Fig. 1
Full size image

The theoretical framework.

Methodology

Participants

FL students, both undergraduate and postgraduate, from various parts of China, are recruited for this study to guarantee the representativeness of the data. After receiving ethical permission from the local institution (the academic committee of School of Foreign Studies, Zhongnan University of Economics and Law, approved number ZUEL/SFS/2024/058), the online questionnaire was distributed between December 2024 and January 2025, and all methods were performed in accordance with the relevant guidelines and regulations afterwards. Additionally, prior to conducting the pragmatic study, informed consent was obtained from all the participants. Every participant was informed that the survey was anonymous and that the information gathered would only be utilized for the study.

Instruments

A questionnaire consisting of 5 sections was tailored to gather empirical data for testing the conceptual model. Specifically, the first 4 parts assess participants’ GenAI acceptance, AI literacy, AI self-efficacy, and emotions respectively. At the conclusion of the questionnaire, demographic information (i.e. gender, study level, university level, region, and language) was collected to comprehend their cohort backgrounds. Likert scales with five points, ranging from 1 (strongly disagree) to 5 (strongly agree), were used to rate most of the items. For an item measuring AU (i.e. “your usage frequency for GenAI products/applications”) and items assessing emotions, five numerical points were converted to represent frequency (1 = never, 2 = rarely, 3 = sometimes, 4 = often, 5 = very often). Appendix A (see Supplementary Material) shows the list of items used in the questionnaire.

GenAI Acceptance Scale, developed by Yilmaz et al.18, is the assessment employed to measure participants’ acceptance of GenAI. It includes 20 items, divided into 4 categories to assess the 4 aspects of the UTAUT model: PE (using items such as ‘The use of GenAI applications increase my chances of achieving the things that are important to me.’), EE (Learning how to use GenAI applications is easy for me.’), SI (using items such as ‘The people I model my behavior on think I should use GenAI applications.’) and FC (using items such as ‘My interaction with GenAI applications is clear and understandable.’). Besides, the other two constructs of the UTAUT model, BI and AU, are assessed by 2 items (i.e. ‘I am likely to recommend GenAI systems to people I know.’; ‘I am interested in using GenAI systems.’) and 1 item (i.e. ‘Please choose your usage frequency for GenAI’), respectively. These items are adopted from the UTAUT291 and other well-established studies that have applied UTAUT to technology usage and acceptance66,90.

Generative Artificial Intelligence Literacy Scale is the instrument assessing participants’ GenAI literacy, which encompasses 15 items. The scale is adapted by Wang, Wang, Li, et al.62 based on the existing AI literacy scales. It comprises 4 dimensions, namely the awareness of GenAI (measuring the capacity to identify and comprehend GenAI in interactions via items such as “I understand how GenAI products process images to achieve visual recognition functionality.”), the usage of GenAI (assessing the capacity to effectively utilize GenAI to accomplish tasks via items such as “I can use GenAI applications or products to enhance my work efficiency.”), the evaluation of GenAI (assessing the capacity to critically evaluate GenAI and their outcomes via items such as “I can choose the appropriate solution from the various solutions provided by GenAI-related applications and products.”), and the ethics of GenAI (assessing the capacity to identify risks and recognize responsibilities associated with GenAI usage via items such as “I am always vigilant about the misuse of GenAI.”).

Self-efficacy items are chosen from the UTAUT model19 to measure FL students’ GenAI self-efficacy. 4 of the original 8 items were finalized by Venkatesh et al.19 based on their performance in the preliminary test in terms of validity and reliability. To fit the context of this study, the original wording “system” has been modified to “GenAI”. Hence, items such as “I could complete a job or task using GenAI applications, if I had a lot of time to complete the job for which the software was provided” were stated.

International Positive and Negative Affect Schedule Short Form (I-PANAS-SF92), a short version of a 20-item Positive and Negative Affect Schedule (PANAS93) was used in this study. It is a self-report questionnaire containing 10 items, 5 positive affect (alert; inspired; determined; attentive; active) and 5 negative (upset; hostile; ashamed; nervous; afraid92). The abbreviated form was chosen for the following reasons. First, given the numerous variables intended to be assessed in the study, a quick emotional assessment could prevent respondents from becoming weary or scatterbrained during the drawn-out survey process. Second, despite its briefness and conciseness, I-PANAS-SF was analyzed and demonstrated sound performance in a number of qualitative assessments, including convergent and criterion-related validities, internal reliability, and temporal stability92. Third, the form was developed with a cross-cultural focus, involving a diverse group of participants from over 10 countries, including China. Consequently, the I-PANAS-SF is suitable for this study to investigate the emotions of Chinese FL students towards GenAI.

Procedure

Prior to formally gathering participant data, preparatory work is required, including creating test materials, conducting pilot research, and modifying experiment specifics based on participant comments. Firstly, two certified individuals who passed the China Accreditation Test for Translators and Interpreters (CATTI) were invited to translate all of the materials from English to Chinese independently. After that, they contrasted their translations before finalizing the questionnaire. Two bilingual researchers were then asked to check for ambiguity and offer workable suggestions for revision. The final form was decided upon after much discussion and uploaded to an online data collection tool, Wenjuanxin.

A pilot study was conducted to identify any issues that the researcher might have overlooked, with a representative sample of 35 students. The researcher closely examined these students’ behavior to spot any indications of fatigue or restlessness while they were completing the tests, keeping account of how long they spent on each activity. The pilot subjects were then encouraged to offer comments and express their own feelings. In order to get helpful criticism, the author proactively inquired if the wording of the questionnaires caused any confusion and if there were any suggestions for improvement. Based on the thorough feedback that was received, the author went forward to improve the experimental design. For instance, the attention check item (i.e. “I would be grateful if you could select Strongly Agree.”) was inserted midway through the questionnaire.

Data analysis

To verify the research hypotheses, statistics obtained from the questionnaire were analyzed using SPSS (version 26.0) and Amos (version 28.0). The first procedure is descriptive data analysis, which examines the distribution of participants’ demographic details including gender, study level, university level, region and language. Secondly, tests for reliability (as determined by Cronbach’s alpha coefficients and composite reliability [CR]) and validity (as determined by convergent and discriminant validity) were applied to evaluate the measurement scales62,94. In this regard, internal consistency was evaluated using Cronbach’s alpha, where values greater than 0.70 signify ‘good’ reliability95. CR was calculated to further assess the construct’s overall reliability, with values higher than 0.70 were considered acceptable94. In addition, to assess convergent validity, average variance extracted (AVE) was computed, with a recommended threshold of greater than 0.5096. Meanwhile, discriminant validity is proven when the square root of the AVE for each concept is higher than its correlations with any other construct96.

After ensuring that the validity and reliability results satisfied the criteria, the Structural Equation Modeling (SEM) fit indices were measured to check whether the hypotheses are valid. All path coefficients were then estimated by SEM analysis97. According to Hair et al.98, it is composed of the Comparative Fit Index (CMID), Degrees of Freedom (DF), Goodness of Fit Index (GFI), Comparative Fit Index (CFI), Tucker-Lewis Index (TLI), and Root Mean Square Error of Approximation (RMSEA). The threshold value of CMID/DF, GFI, CFI, TLI and RMSEA is 5, 0.8, 0.8, 0.8 and 0.08, respectively99,100,101. Lastly, to examine if path coefficients vary by demographic information (e.g., gender, study level, university level, region and language), moderating effect analysis was conducted after the aforementioned study.

Results

Demographic analysis

Adopting the cross-sectional survey design102, this study collected quantitative data from a total of 527 responses from different universities across China to examine FL students’ acceptance of GenAI. Stringent procedures were followed to ensure the quality of the sample, in accordance with the methodological guidelines proposed by Hair et al.98, as detailed below. First, participants who failed to respond to the attention-check item were flagged as invalid. Second, the pilot test revealed that it should take at least two minutes to complete the whole questionnaire. Because of this, replies from participants who completed the questionnaire in an abnormally short amount of time (less than 100 s) were excluded as invalid in that it was assumed that they had not taken the task seriously. A valid sample of 409 replies (valid response rate: 77.6%), including 105 males and 304 females, was collected for the study following rigorous screening. Table 1 further reports the participants’ demographic information.

Table 1 Demographic profile (n = 409).

Reliability and validity results

Prior to examining the reliability (as measured by Cronbach’s alpha coefficients and CR) and the validity (as determined by convergent validity and discriminant validity), items from the Generative Artificial Intelligence Literacy Scale62 and I-PANAS-SF92 were parcelled for further analysis. Compared to item-level data, models utilizing parceling technique demonstrate following psychometric and estimation advantages: (a) enhanced parsimony, with fewer parameters estimated at both the construct and global model levels; (b) a minimized risk of correlated residuals or cross-loadings, owing to the use of fewer indicators with smaller unique variances; and (c) a reduction in various sources of sampling error103. A domain-representative parceling approach was utilized in this research. This technique explicitly addresses construct multidimensionality by ensuring that the parcels reflect not only the shared common variance but also the reliable unique variance from each dimension, thereby preserving the multifaceted nature of the constructs in the analysis104,105. Via this approach, each parcel includes all aspects, or dimensions, that are represented in the set of indicators105. In practice, the 15-item Generative Artificial Intelligence Literacy Scale was divided into 3 parcels, each comprising 2 awareness items, 1 usage item, 1 evaluation item, and 1 ethics item. Similarly, the 10-item I-PANAS-SF was divided into 5 parcels, each containing 1 positive and 1 negative item.

Regarding reliability, the Cronbach’s alpha coefficient (α) of each sub-scale ranges between 0.792 and 0.955 (see Table 2). These Cronbach’s α-values are all above the benchmark of 0.6106, with a coefficient of 0.7 or above being considered ‘good’95. Hence, the sound inter-correlation of indicators from the same construct has been demonstrated. In addition, factor loading, CR, and AVE were computed in turn for a more thorough assessment. The recommended threshold values are 0.60107, 0.7094 and 0.5096 for factor loading, CR, and AVE, respectively. According to some academics, the construct’s convergent validity is still adequate if AVE is less than 0.5 but CR is greater than 0.696. In this study, all factor loading, CR, and AVE values reached the standard, affirming a satisfactory level of internal consistency and acceptable reliability across the measured items and constructs (see Table 2).

Table 2 Convergent validity.

In addition, the discriminant validity among variables (describing the distinctiveness of a construct from other constructs) was examined. This can be computed by contrasting the correlation coefficients of the variables with the square root of the AVE. A variable is deemed to have good discriminant validity if the correlation coefficient between it and the other variable is less than the square root of the AVE96. The current study shows that the AVE’s square root was higher than all other values it measures, indicating that the proposed model possessed appropriate discriminant validity (see Table 3).

Table 3 Discriminant validity.

SEM results

The subsequent step entails an assessment of the SEM fit indices to ensure that the hypothesized relationships were not spurious. It consists of CMID, DF, GFI, CFI, TLI, RMSEA in accordance with Hair et al.98. The criteria recommended by Bagozzi and Yi99, Hair et al.100 and Hayduk101 were used to evaluate the model fit. The fit indices for the measurement model were acceptable with χ2 = 1653.879, χ2/df = 3.035, CFI = 0.898, and RMSEA = 0.071. Improvements were made to the model to correlate errors between error for item BI1 and error for item BI2. The new fit indices for the improved measurement model were favorable (χ2 = 1525.157, χ2/df = 2.804, CFI = 0.910, RMSEA = 0.066), demonstrating the model’s suitability for integrating the collected data and confirming the correlations between the variables. Table 4 displays the corresponding numerical results and suggested values.

Table 4 Model fit indices.

The results of the hypothesis testing are provided in Table 5. Seven out of the eight hypotheses garnered support. More precisely, in terms of the UTAUT model, PE (β = 0.546, p < 0.001) and SI (β = 0.514, p < 0.001) exerted a significant positive influence on BI, and FC (β = 0.229, p = 0.034) performed the same on AU. However, the direct relationship from EE to BI was insignificant (β = 0.108, p = 0.083). Regarding other paths, BI (β = 0.683, p < 0.001), E (β = 0.165, p = 0.033), L (β = 0.220, p = 0.041) and SE (β = 0.448, p < 0.001) all had positive significant effects on AU. Hence, this study accepts H1a, H1c, H1d, H1e, H2, H3 and H4, while H1b fails to find support within the study confines.

Table 5 Structural model path analysis results.

Moderating effects

In the presence of moderation, the strength and direction of the interaction between two constructs are determined by a third construct, known as the moderator. The present study systematically introduced gender, study level, university level, area, and language in response to H5 to assess its moderating effects on the model. Of these 5 moderators, only 2 revealed a significant effect. Specifically, gender played a moderating role in the relationships between L to AU (βmale = 0.665, βfemale = 0.428, p = 0.007; see Table 6), and region significantly moderates the path between PE to BI (βNortheast = 0.539, βEastern = 0.686, βCentral = 0.671, βWestern = 0.680, p = 0.022; see Table 7). The other 3 moderators, study level, university level and language failed to play any moderating role (individual path analysis of these 3 moderators is listed in the Appendix B; see Supplementary Material). Therefore, part of the H5a and H5d are verified. The results of SEM and moderating effects are displayed in Fig. 2.

Table 6 Individual path analysis for moderation effect of gender.
Table 7 Individual path analysis for moderation effect of region.
Fig. 2
Fig. 2
Full size image

Results of the SEM and moderating effects. *p < 0.05; **p < 0.01; ***p < 0.001. Solid lines represent confirmed hypotheses; dash lines represent unsupported hypotheses.

Discussion

SEM analysis

The current study examined the UTAUT model in the context of GenAI for FL learners and investigated the role of demographics, emotions, AI literacy and self-efficacy in the proposed framework. Empirically, the SEM results supported all hypotheses except for the path from EE to BI. The substantial support for these hypotheses offers strong evidence for the overall validity of the model that integrates the UTAUT with other constructs. Taking it by and large, the observation is consistent with recent research29,108,109, highlighting the significance of UTAUT in understanding users’ acceptance and behavioral intentions toward GenAI tools in FL settings.

In reference to the UTAUT model path, the study found that 4 of the 5 paths were confirmed. Firstly, PE emerges as the most influential factor on BI, indicating that it can, to a great extent, positively predict FL learners’ intention to use GenAI. This finding aligns with previous studies targeting different groups of participants, including EFL learners29, Poland university students110 and so on. This suggests that FL learners are more likely to use GenAI applications when they regard it as a helpful tool for completing assignments and enhancing their study efficiency. Conversely, the second moderator EE was revealed to be non-significantly related to BI. This resonates with studies related to massive online courses for higher education students in Saudi Arabia111 and GenAI usage among Egyptian university students110, but contradicts findings regarding mobile learning adoption in the higher education context112 and humanoid robot assistance in academic writing for freshmen113, where EE was found to be significant. One plausible explanation may be rooted in the ease of use of the GenAI applications and products114. The user-friendly and easily navigable GenAI may weaken the influence of EE on BI. Thirdly, this research showed that SI positively predicts BI in agreement with studies on ChatGPT in higher education learning in both Indonesia115 and Poland116. The findings imply that external figures that FL learners admire or view as role models, such as peers, administrators, and supervisors, have an impact on their adoption and utilization of GenAI. Fourthly, FC has been found to have a positive and significant effect on AU. This finding is consistent with earlier research on learning management system117 and GenAI such as ChatGPT115. It suggests that that learners who find an administrative and technological framework in place to support the use of GenAI are more willing to employ it. Lastly, the study unveils that BI is also a strong positive predictor of AU, which is in line with prior GenAI research in various contexts29,110,118. It implies that GenAI is typically used more frequently by students who are more interested in and intend to use it.

Regarding extended relationships linked to the UTAUT model, emotions, AI literacy and AI self-efficacy all emerged as a crucial influencer of AU. Specifically, emotions were found to be positively and significantly associated with AU, albeit with a lesser effect compared to other factors. This corresponds with Ahmadipour119, Hilliard et al.120 and Wang, Wang, Pan, et al.2, who pinpointed that students’ emotions directly influence their technology acceptance. The result suggests that a positive mentality to use GenAI could enhance technology acceptance and usage. Likewise, AI literacy exerts a positive significant impact on AU. In line with previous studies on AI-based technology for EU citizens66 and GenAI educational resources for university instructors121, the findings empirically support the hypothesis that people who are more effectively able to recognize, identify, and evaluate GenAI-related products are more likely to adopt and utilize them. Furthermore, AI self-efficacy also demonstrates predictive power over AU. These findings are in agreement with research on GenAI-powered learning environments among EFL Chinese students2, university students in Germany122, and ASEAN member state students123. This suggests that students are more prone to use GenAI if they have more faith and confidence in their abilities to interact with the technology.

Moderating effects analysis

While assessing bivariate relationships between variables indicates possible causal effects, it does not explain how, why, or for whom these effects hold124. The current study further explores the potential moderating effects of demographic data, such as gender, study level, university level, region and language. First of all, gender plays a moderating role only in the relationship of L → AU, with males’ path coefficient significantly higher than that of females. This result is consistent with studies on GenAI acceptance in the Polish education setting110 and AI acceptance in the primary care context125, showing that gender cannot enhance or diminish the effect of relationships in the UTAUT model. However, it contrasts with research on GenAI in the Egyptian sample110, where gender exerts significant moderating effects in most UTAUT paths. Thus, the UTAUT by gender needs to be further scrutinized. For relationships outside the UTAUT model, the result aligns with most previous studies that recognize gender disparity (male students outperform female students) in AI-related education including AI literacy126,127. This can be explained by different gender-specific brain perceptions and processing modes from the point of cognitivists128, or being less exposure to computer science during childhood for females from the pedagogy perspective129.

Second, study level does not significantly influence the proposed model relationships, which indicates that GenAI is equally effective for both FL undergraduate and postgraduate learners. A related research reported a similar finding that education degrees are not a significant determinants of acceptance of AI-based technology125. Nevertheless, different results were achieved in the Polish and Egyptian samples that found undergraduates of different grades perform significantly differently in the UTAUT model110. The inconsistent finding could potentially be explained by disparate technology products, participants’ majors, cultural backgrounds, and education system, and thus need to be further explored.

With respect to the last three moderators (university level, region and language), their findings are scarcely comparable to those of earlier research given the current study responds innovatively to it. Specifically, the moderating effect of the university level was not found in the current study. This implies that as key university stereotypes diminish, GenAI proves to be indiscriminatingly efficient for FL learners across various Chinese university levels. In terms of region, the moderating effect was significant in the path from PE to BI. Learners’ GenAI acceptance in the eastern China Region are more likely to be influenced by their belief in the system’s effectiveness for enhancing performance, followed by users in western China, central China and northeast China. One possible explanation could originate from the regionalization method in this study, since the ranking coincides exactly with the number of provinces included in each region. Lastly, there is no evidence of learners’ L2 playing a moderating role throughout the entire model in the current study. This could be plausibly explained by the fact that GenAI can generate authentic and idiomatic texts in a variety of languages including English, French, German and so on130, and thus students majoring in different languages do not show significant different acceptance towards GenAI.

Implications, limitations and future study

Theoretically, this study is one of the few that has expanded the UTAUT model into a broader context, investigating whether and to what degree users’ demographic backgrounds, emotions, AI literacy, and AI self-efficacy affect the acceptance of GenAI by FL students. It has narrowed the gap between traditional technology theory and newborn favourite GenAI, enhancing the understanding of GenAI acceptance in the L2 context. The findings provide valuable insights into the key factors affecting the acceptance and integration of GenAI in higher education, especially regarding FL students. In addition, the newly proposed theoretical framework introduces and validates the affective and cognitive factors, providing a more nuanced theoretical explanation for the interplay between GenAI acceptance and other potential predictors (e.g., demographics, emotions, AI literacy and AI self-efficacy). The garnered insights can be harnessed by educators, researchers and policymakers in integrating GenAI tools into both FL and non-FL students’ academic processes. This cooperative effort might encourage fruitful interaction among the stakeholders and strengthen the deployment of GenAI applications.

Practically, the study reinforces the significant role PE, EE, SI, AI literacy, AI self-efficacy and demographic backgrounds play in shaping FL students’ willingness to adopt and use GenAI products. The findings of this study yield concrete implications for various stakeholders in educational practice. To further motivate students to embrace new technologies, the following strategies can be suggested. Firstly, universities could implement GenAI awareness campaigns, supplemented by workshops and training programs to help students build GenAI confidence, improve GenAI literacy and self-efficacy, acquire GenAI knowledge and encourage student engagement in insightful discussion. It is also worth mentioning that pedagogies should incorporate guidance on critical thinking by explicitly address AI ethics, limitations, and academic integrity, encouraging students to challenge and enhance GenAI outcomes rather than merely accept them. By placing critical thinking at the core, such programs can harness the positive aspects of AI literacy and self-efficacy while mitigating the resistance that may arise from a superficial understanding of its risks. Secondly, those whose viewpoints students respect, such as supervisors, instructors, and influential peers, ought to leverage their authority to foster the adoption of GenAI. Motivating instructors and fellow students who have had positive interactions with GenAI technologies would help achieve this goal. Thirdly, teachers should take good advantage of GenAI to create joyful immersive language experiences for FL students, fostering a more positive emotional connection to the technology-added learning process. Such interventions should not be one-size-fits-all. For students with high anxiety and low self-efficacy, providing scaffolded, low-stakes opportunities to interact with AI can build confidence and reduce fear. Conversely, for students with high AI literacy, instruction should focus on guiding them toward critical and creative uses of AI, moving beyond basic functionality. Lastly, for GenAI products’ designers and developers, they should put more emphasis on the demands of students by devising user-friendly interfaces and easy-to-navigate systems. Such initiatives aim to encourage students to engage with GenAI, avoiding disuse because of unfavourable user experiences.

Notwithstanding the findings, this study admittedly suffers its own limitations which need to be carried out in the future study. First and foremost, the cross-sectional nature of the study restricts its capacity to build causality, underscoring the need for future longitudinal research to more effectively understand the dynamic interactions among these variables over time. Another limitation lies in the unbalanced population distribution of gender, and a limited number of languages. More male participants and students who major in minority languages could be invited in the future study. Finally, this study relied exclusively on quantitative methods to process data. To provide a more comprehensive perspective, future research could incorporate both qualitative and quantitative approaches, allowing for a deeper understanding of the phenomenon.

Conclusion

This study enriches the literature by affirming and expanding the UTAUT model into a broader context, and highlights the moderating effects of demographics, emotions, AI literacy, and AI self-efficacy in the FL setting. In particular, the findings of this study unveil that PE and SI serve as significant predictors of FL students’ BI towards GenAI products. FC, BI, emotions, AI literacy and AI self-efficacy significantly predict learners’ actual usage of the technology. Beyond that, gender and region demonstrate a moderating effect in certain paths of the overall model.

This study holds significant implications for both theory and practice. It contributes to a better understanding of how various factors influence Chinese FL students’ acceptance of GenAI. The study also offers practical strategies for universities, faculty members, researchers, technology developers, and designers to actively promote the rational adoption of GenAI, providing valuable insights for future research.