Background

Over the past thirty years, empowerment and recovery have emerged as transformative paradigms in mental health care. These concepts signify a shift from traditional paternalistic models that often left users feeling dependent and powerless [1], towards approaches on health democracy and patient-centered care. This evolution places the focus on autonomy, resilience, and the meaningful participation of users in shaping their care [2]. Thus, recovery and empowerment are interconnected pillars driving this transformation, enabling individuals to actively engage in decisions that affect their lives [3]. As a result, this shift in mental health care requires a genuine focus on users’ empowerment and recovery to have real meaning and not be merely ostensible.

Definitions and overlap between the concepts

Empowerment is a concept whose definitions and conceptual frameworks vary widely in the literature [4, 5]. Originally, this notion refers to the reclaiming of power—by and for the people who have been deprived of it—in order to address issues that concern them at an individual, community, and organizational level [6]. According to the World Health Organization, “empowerment needs to take place simultaneously at the population and the individual levels. Empowerment is a multidimensional social process through which individuals and groups gain better understanding and control over their lives. As a consequence, they are enabled to change their social and political environment to improve their health-related life circumstances” [7].

Recovery is defined as a personal and nonlinearly process of living a satisfying, hopeful, and meaningful life despite the difficulties caused by mental illness [8, 9]. This concept was initially theorized by Deegan [8], a researcher who identify as a psychiatric survivor—a person who have experienced psychiatric treatment and now contribute to research from that perspective—before being appropriated and popularized by different professional healthcare researchers, notably by Anthony, a pioneer in the field of psychosocial rehabilitation [9]. Although the concept of recovery also lacks a consensual definition [10], the conceptual framework most commonly used in the literature today is known as Connectedness, Hope and Optimism, Identity, Meaning and Purpose, and Empowerment (CHIME) [11].

There is an overlap between the two concepts—empowerment and recovery—which usually show strong correlations [12]: empowerment is a key component of the recovery process [11, 13], while psychological coping—an aspect of empowerment—has a definition that closely resembles that of recovery. For instance, a review describes this aspect as “how patients framed their illness, made sense of their illness in the context of their lives, maintained hope, and reduced fears” [5]. Additionally, some aspects of the two concepts are quite similar, such as “self-esteem and sense of efficacy” in empowerment and “personal confidence” or “goal orientation” in recovery [14, 15]. The study will thereby explore how the concept of recovery can assist in evaluating the responsiveness of an empowerment measure.

The recovery paradigm in practice

The recovery paradigm has redefined mental health care worldwide, moving beyond symptom management to encompass holistic approaches that emphasize hope, autonomy, and meaningful social life [3]. Central to this paradigm is patient-centeredness, which challenges hierarchical relationships in care and promotes collaboration between users and healthcare providers [16]. Shared decision-making, peer support, and the integration of lived experience are integral components [17]. Empowerment, as both a process and an outcome, plays a critical role in facilitating recovery, enabling users to regain control over their lives and actively shape their care [18]. However, people living with mental disorders are particularly prone to losing their empowerment, even if they’re engaged on a recovery journey—partly due to their condition and partly due to psychiatric measures aimed at “correcting” their behavior [19]. These corrective measures manifest as formal constraints (involuntary treatment) as well as informal constraints such as disapproval, coercion, or active pursuit of medication compliance by healthcare professionals [19]. This approach is contrary to the idea of users being active participants in their own care. Thus, the concept of empowerment has become as significant in the mental health field: it is at the heart of the user movement [20], changes in professional practices [21], as well as healthcare policy reforms worldwide [22].

The French context for empowerment in mental health care

In France, the mental health care system has been shaped by a policy of psychiatric sectorization since the 1960s, designed to ensure geographically localized care and continuity of services [23]. This framework aimed to support deinstitutionalization by prioritizing community-based care. However, significant disparities in the allocation of resources across sectors and regions persist, complicating equitable access to mental health care [24]. Additionally, the growing focus on outpatient care and mobile mental health teams has significantly reduced hospitalizations [25], aligning with deinstitutionalization goals. Yet, this shift may unintentionally undermine peer support that often flourish in inpatient settings and contribute to collective empowerment.

France has also embraced the principles of “health democracy,” reinforced by the 2002 Patients’ Rights Act, which seeks to amplify the voices of users in shaping health policies [26]. This approach has been particularly emphasized in mental health, where initiatives such as Local Mental Health Councils (Conseils Locaux de Santé Mentale, CLSM) promote co-governance by bringing together professionals, users, and local authorities to address mental health needs collectively [27]. Additionally, the role of user participation is increasingly recognized as a means to challenge traditional power dynamics in psychiatry and foster recovery-oriented practices [28]. Despite these advances, tensions persist between the ideals of user participation and and the practical realities of psychiatric care, as highlighted by challenges in fully integrating lay expertise into professional practices [29]. These tensions underscore the necessity of tools to measure empowerment, not only to assess individual outcomes but also to evaluate how well services promote empowerment and recovery-oriented practices.

Need for having a french-language measure

Developing a French-language empowerment scale addresses these needs, providing a culturally adapted instrument to evaluate the alignment of mental health services with empowerment and recovery goals. Such a tool would help capture the nuances of user experiences and support ongoing transformations in the French mental health care system towards more inclusive and participatory practices.

The Empowerment Scale (ES) is the most widely used and translated tool in the field of mental health [30], although its psychometric properties are not satisfactory in terms of validity and internal consistency [31, 32]. However, the ES has never been evaluated in terms of responsiveness, and has never been translated and studied in French.

Aim of the study

The aim of this study was to translate the Empowerment Scale into French and to assess its internal consistency, validity and responsiveness.

Presentation of the scale

The ES, a 4-point Likert scale, was developed in 1997 with users of mental health services considered to be leaders of the user movement in the United States [14]. The scale consists of 28 items distributed across five dimensions: community activism and autonomy, self-esteem and effectiveness, optimism and control over the future, righteous anger, power and powerlessness. The original scale is available in Additional file 1.

Method

Population

This study focused on participants from a randomized controlled trial of the Psychiatric Advance Directives facilitated by Peer Workers (PW-PAD) [33] recruited from seven hospitals across three major French cities. The decision to draw the sample from a concurrent RCT was driven by methodological and practical considerations. This allowed access to a well-defined and characterized sample meeting strict inclusion criteria, with data collected by trained investigators using standardized protocols. Additionally, the study design enabled multiple assessment points, providing a robust framework for evaluating the scale’s sensitivity to change.

The inclusion criteria of this RCT were being aged 18 years or older, having a diagnosis of schizophrenia, bipolar I disorder, or schizoaffective disorder according to the DSM-5 criteria, experiencing involuntary hospitalization in the past 12 months, and demonstrating decision-making capacity as assessed by a psychiatrist using the MacArthur Competence Assessment Tool for Clinical Research (MacCAT-CR) [34].

Study design

To assess the psychometric properties of the French ES (F-ES), sociodemographic and clinical data from the PW-PAD study were used, including age, gender, education level, employment status, marital status, diagnosis, symptom severity (assessed using the Clinical Global Impression scale), and guardianship status. Additionally, subjective measures were used, such as the Recovery Assessment Scale (RAS, for which a higher score indicates better recovery) [35, 36], the 4-Point ordinal Alliance Self-report (4-PAS, for which a higher score indicates better therapeutic alliance) [37], the Schizophrenia Quality-of-Life scale (S-QoL, for which a higher score indicates better quality of life, and was validated for bipolar disorders as well) [38, 39], and the Modified Colorado Symptom Index (MCSI, for which a higher score indicates a higher level of perceived symptoms) [40].

User involvement and experiential knowledge in Research

The involvement of users with lived experience of mental illness in research remains limited [41], despite growing recognition that interventions are more effective when co-designed with those directly affected [42]. The design of the principal RCT was co-constructed with several users with lived experience, who were directly involved in the selection of the Empowerment Scale as a key measurement tool. Users also contributed to validating the translation of the Empowerment Scale, ensuring the tool was culturally and contextually adapted to the French setting.

In the present study, which focuses on the psychometric properties of the F-ES, the first author—a public health PhD student with lived experience of psychiatry—acted as a “peer researcher” [43]. This perspective guided all stages of the research process, from conceptualization to interpretation of findings.

The study also benefited from a multidisciplinary team, including public health researchers and experts in psychometry, recovery, and empowerment. Each team member contributed complementary expertise, ensuring the adaptation of the Empowerment Scale was scientifically rigorous and grounded in both academic and experiential insights. The collaboration also fostered a reciprocal exchange of knowledge, where experiential insights informed theoretical approaches, and academic perspectives enriched lived experiences [44]. This mutual learning process exemplifies the empowerment principle at the heart of this research.

French translation of the empowerment scale

The Empowerment Scale was translated into French following cross-cultural adaptation guidelines for self-assessment measures [45]. The translation committee, comprising four experts in mental health empowerment (two native French speakers and two native English speakers) and a professional translator, oversaw the translation process. Native French speakers conducted forward translation into French, which was then back-translated into English by the professional translator. The back-translation was reviewed by the translation committee to ensure consistency with the original scale, but it was not assessed directly by the original authors. This limitation is acknowledged and was addressed by a rigorous translation process involving bilingual experts. Any discrepancies were resolved through discussion with native English speakers, resulting in the final French version agreed upon by all committee members. The French translation is available in Additional file 2.

Statistical analysis

After verifying that the missing data, including sociodemographic and clinical data, were not completely at random (MCAR) using the Little MCAR test, we addressed them with multiple imputation by chained equations (MICE) [46]. This method, widely recommended in the literature to minimize bias, repeatedly replaces missing values to achieve more robust estimates [47]. Fifty multiple imputations were performed for each missing variable to ensure result stability. All subsequent statistical analyses were conducted on the complete dataset unless otherwise noted.

The psychometric properties of the instrument were analysed using classical test theory. A principal component analysis (PCA) with varimax rotation and a bootstrap method was initially conducted to determine the optimal number of dimensions. Exploratory factor analyses (EFA) with the WLSMV estimator and oblique rotation [48] were then conducted to identify the instrument’s underlying structure. The final number of dimensions was determined using the Kaiser rule (eigenvalues > 1) and the scree plot [49] while considering the interpretability of the dimensions. Items with factor loadings below 0.32 [50] or cross-loadings above 0.32 were removed [51].

The specified model fit was verified using confirmatory factor analysis (CFA) with various indices. A good fit was indicated by a comparative fit index (CFI), Tucker‒Lewis index (TLI), and adjusted goodness-of-fit index (AGFI) of ≥ 0.95, a standardized root mean square residual (SRMR) of ≤ 0.08, and a root mean square error of approximation (RMSEA) of ≤ 0.06 [32]. An RMSEA of ≤ 0.10 is acceptable if the upper bound of the 90% confidence interval is also below this threshold [52]. The df/χ² ratio was used for model comparison, with lower ratios indicating better fit.

Internal consistency was assessed using Cronbach’s alpha (values ≥ 0.7 indicate satisfactory reliability [31]). Floor and ceiling effects were measured by the proportion of respondents at response extremes. The acceptability of each dimension was evaluated by the rate of missing data before imputation. Unidimensionality was measured by Loevinger’s H coefficient (considered low if ≥ 0.3, medium if ≥ 0.4, and high if ≥ 0.5 [53]). The inlier-sensitive fit (INFIT) for each item was also calculated, with values between 0.7 and 1.3 indicating that the items measured the same concept [54].

Convergent and discriminant validity were assessed using internal consistency coefficients (ICCs) and discriminant validity coefficients (DVCs). The ICC corresponds to the correlation of each item with its own dimension score (calculated without the evaluated item), while the DVC corresponds to its correlation with other dimension scores. Overlap analysis between minimum ICCs and maximum DVCs determined whether items correlated more strongly with their intended dimension [55]. Concurrent validity was assessed by examining the relationships between the F-ES index and the RAS, 4-PAS, S-QOL, and MCSI scales. It was hypothesized that the F-ES would strongly correlate with the RAS (recovery), moderately with the 4-PAS (therapeutic alliance) and S-QOL (quality of life), and moderately negatively with the MCSI (symptomatology). The “self-esteem and self-efficacy” dimension was evaluated using the “self-esteem” dimension of the S-QOL and the “self-confidence and hope” and “goal orientation and achievements” dimensions of the RAS, with moderate to strong correlations. The “autonomy” dimension was expected to strongly correlate with the “autonomy” dimension of the S-QOL. Other dimensions were not evaluated due to the lack of theoretical relationships. Differences between subgroups were analysed using Wilcoxon-Mann‒Whitney, Kruskal‒Wallis, Dunn’s tests, and Spearman’s rho for correlations. Effect sizes ≥ 0.10 were considered small, ≥ 0.30 moderate, and ≥ 0.50 large [56]. Differential item functioning (DIF) could not be studied because three of the four dimensions contained fewer than four items.

To assess ES responsiveness to change, participants were grouped based on changes in their recovery scores from baseline to 6 months. Participants with a ≥ 10-point increase were categorized as “risen,” those with a ≥ 10-point decrease as “fallen,” and others as “stable.” The Kruskal‒Wallis test was used to compare the mean differences in recovery scores between baseline and 6 months, followed by Dunn’s test to identify differences between groups. A second Kruskal‒Wallis test was used to evaluate differences in empowerment scores across recovery levels (risen, stable, fallen). Effect sizes were measured using eta-squared: 0.01 to < 0.06 for small effects, 0.06 to < 0.14 for moderate effects, and ≥ 0.14 for large effects [57].

The data analysis was conducted using RStudio software version 2023.09.1 + 494.

Scoring

The reverse-scored items (indicated in Additional file 4) were recorded so that higher scores corresponded to higher levels of empowerment. For each participant, the score for each dimension was obtained by calculating the average of the item scores within that dimension. If at least half of the items were completed, their average was used to replace missing items. Dimension scores were then linearly transformed on a scale from 0 to 100. A score of 100 indicated the highest possible level of empowerment, and a score of 0 indicated the lowest.

Results

Sample characteristics

A total of 394 participants were included between January 2019 and June 2020. Patient characteristics are presented in Table 1. The mean number of lifetime hospitalizations was 8.79 ± 13.45, and the mean number of involuntary hospitalizations was 4.4 ± 7.25.

Table 1 Participant characteristics (n = 394)

Missing data

Missing data on the Empowerment Scale ranged from 3.30 to 14.72% per item. However, complete cases comprised only 52.80% of the sample (N = 208). Little’s MCAR test applied to the full 28-item scale did not reject the null hypothesis, with a p value above the significance threshold (0.05). Therefore, exploratory analyses (PCA and EFA) were conducted only on complete cases. Following these analyses, five items were removed, as described in Sect. 3.3 on factor structure. A new MCAR test was conducted on the reduced scale, and the MCAR assumption was rejected (p = 0.004). Prior to conducting various confirmatory factor analyses, missing data for the remaining 23 items of the instrument were imputed, enabling these analyses and subsequent analyses to be conducted on the entire sample of 394 participants.

Factor structure, internal consistency, and unidimensionality

PCA with bootstrapping revealed 8 factors with eigenvalues greater than 1. An examination of the scree plot revealed two elbows [see Additional file 3], one around the second component and another around the fourth component. This plot suggested that reducing to just two axes could result in a significant loss of information, as substantial informational contributions were still observed with four axes but became minimal when moving from four to five factors. Therefore, utilizing four to eight axes appears most suitable for representing these data. The first four factors explained 41.95% of the cumulative variance, while the first eight factors explained 59.68%.

Several exploratory factor analyses were conducted to determine the number of factors to retain, considering 8, 7, 6, 5, and 4 factors. The first three attempts resulted in unsatisfactory factor structures, characterized by one or two factors with a single factor loading of 0.32 or higher [50]. Models with 4 or 5 factors were further examined in detail.

The EFA with 5 factors identified interpretable dimensions: self-esteem and efficacy, autonomy, community activism, critical thinking skills, and righteous anger. The factor loadings were less than 0.32 for items 3, 7, and 8, while items 16 and 22 showed cross-loadings. These five items were removed to enhance the convergent and discriminant validity of the model [50]. The structure of the 5-factor scale was further analysed using CFA. The fit indices indicated good model fit (CFI = 0.95, TLI = 0.96, AGFI = 0.94, SRMR = 0.07, RMSEA = 0.09, 90% CI: 0.08–0.10), with a df/χ² ratio of 3.02.

The EFA with 4 factors identified interpretable dimensions: self-esteem and efficacy, autonomy, community activism, and critical thinking skills. The factor loadings were less than 0.32 for items 3, 7, 8, and 15, with items 16 and 22 showing cross-loadings. These six items were removed from the analysis. The fit indices for the 4-factor, 22-item scale were good (CFI = 0.97, TLI = 0.97, AGFI = 0.96, SRMR = 0.07, RMSEA = 0.08, 90% CI: 0.07–0.09), with a df/χ² ratio of 2.83.

The 4-factor model was preferred due to its slightly better results. Cronbach’s alpha showed good reliability for the overall index (α = 0.84) and the “self-esteem” dimension (α = 0.88), questionable reliability for “community activism” (α = 0.63), and unacceptable reliability for the other dimensions [58]. Removing item 2 from the “autonomy” dimension increased the alpha from 0.48 to 0.57. Removing items 4 and 10 from the “critical thinking” dimension increased the alpha from 0.45 to 0.52 and then to 0.58, respectively. The removal of these three items was aimed at slightly improving the reliability from unacceptable to poor levels [58]. Item 1, “I can pretty much determine what will happen in my life”, was removed from the “self-esteem” dimension because it was outside the INFIT range, with no impact on the alpha dimension. Similarly, removal of items 1 and 2 did not affect the index alpha.

The 4-factor, 18-item model showed slightly better results than did the 4-factor, 22-item model (df/χ² = 2.70, CFI = 0.97, TLI = 0.97, AGFI = 0.97, SRMR = 0.07, RMSEA = 0.07 with a 90% confidence interval from 0.06 to 0.07). None of the models explored demonstrated superior internal consistency or factorial structure compared to the selected model. The 4-factor, 18-item scale is available in Additional file 4.

Analysis of item-internal coefficients (IIC) and item-discriminant coefficients (IDC) confirmed internal validity. All the IIC values exceeded 0.39, indicating that the item correlations with their respective dimensions were greater than those with other factors. However, the ICC for the “autonomy” dimension could not be studied because it contained only 2 items.

None of the items in this model exhibited an INFIT outside the acceptable range. The H coefficients suggested strong unidimensionality for self-esteem, moderate unidimensionality for autonomy and activism, and weak unidimensionality for critical thinking.

Acceptability, floor and ceiling effects, and mean scores

Regarding questionnaire acceptability, the rate of missing data per dimension was low (between 7.29% and 9.46%). Floor effects ranged from 0.26 to 1.36%, and ceiling effects ranged from 1.05 to 12.39%. Finally, the mean scores obtained by participants were 59.08 out of 100 for the index, with dimension scores ranging from 53.51 to 73.15. Details on the dimensions and index characteristics are provided in Table 2.

Table 2 Dimension and index characteristics of the French ES (n = 394)

Responsiveness

Statistical analyses to assess clinical subgroup relevance indicated a significant difference among groups (risen, fallen, stable), with an eta-squared of 0.78, suggesting a substantial effect. The mean differences in recovery rates from 0 to 6 months were − 23.25 ± 11.62 for the “fallen” group, -0.04 ± 5.44 for the “stable” group, and 19.42 ± 8.09 for the “risen” group. Dunn’s test confirmed significant differences between each pair of groups.

Furthermore, as shown in Table 3, analyses assessing the responsiveness of the Empowerment Scale revealed a significant difference between groups, with an eta-squared of 0.05, indicating low instrument responsiveness. The “self-esteem” dimension demonstrated moderate responsiveness, while the “activism” dimension showed low responsiveness. The “autonomy” and “critical thinking” dimensions did not appear responsive to variations in recovery rates.

Table 3 Responsiveness tests and effect sizes of the F-ES

Concurrent validity

The RAS, S-QOL, MCSI, and 4-PAS scales used to assess concurrent validity appeared to have missing data that were completely random (MCAR) based on Little’s test. Therefore, missing data from these four scales were not imputed, and subsequent analyses were conducted on complete cases.

Table 4 Correlations between ES scores and RAS, MCSI, 4-PAS, and S-QOL scores

As detailed in Table 4, the Empowerment Scale index showed moderate correlations with recovery (RAS index, r = 0.47) and quality of life (S-QOL index, r = 0.28), weak correlations with therapeutic alliance (4-PAS, r = 0.23), and a weak negative correlation with symptomatology (MCSI, r = -0.11). The “self-esteem and efficacy” dimension of the F-ES exhibited strong correlations with the “self-confidence and hope” (r = 0.73) and “goal orientation and achievements” (r = 0.60) dimensions of the RAS, as well as with the “self-esteem” dimension of the S-QOL (r = 0.55). The “autonomy” dimension showed a weak positive correlation with a similar dimension of the S-QOL (r = 0.23).

Regarding sociodemographic and clinical factors, the “critical thinking” dimension of F-ES appeared to be associated with education level (with a significant difference between participants with at least two years of higher education and others), while “autonomy” seemed to be linked to diagnosis (with a significant difference between participants living with bipolar disorder and those living with schizophrenia). The significant relationships of the sociodemographic and clinical variables in Table 1 with the F-ES dimension are reported in Table 5.

Table 5 Significant relationships between F-ES dimensions and index and sociodemographic and clinical factors

Discussion

The initial research question addressed the psychometric properties of the French version of the ES. Therefore, it is pertinent to discuss the validity, reliability, and responsiveness of the instrument. During exploratory analyses, six items exhibited factor loadings below 0.32 or cross-loadings across multiple factors. After removal, the CFA results indicated a good fit between the theoretical model and the observed data [32]. The sample size was sufficiently large for reliable results [59, 60], although it would have been preferable to validate the factor structure in a different subsample than the one used for exploratory analyses. Other studies on the ES have removed items due to insufficient factor loadings [61, 62], but none have reported examining cross-loadings among factors. To our knowledge, none of the previous versions of the ES have shown such strong results regarding factor structure validity. Convergent and discriminant validity were also excellent. However, removing six items—and four more to enhance the reliability and unidimensionality of the model—may raise concerns about measurement validity, particularly regarding comparability with other versions of the ES. It is worth noting, however, that studies on the ES across various cultural contexts, including the original version, have consistently reported somewhat unsatisfactory results in terms of factor structure and internal consistency. These limitations suggest broader challenges with the measurement of empowerment itself, as the construct appears difficult to operationalize with a consistent psychometric structure. This systemic issue inherently affects the comparability of results, regardless of the specific adjustments made to the French version.

The F-ES introduced a new structure. The original “autonomy and community activism” dimension was split into two distinct dimensions, as suggested by the EFA results. However, the ‘just anger’ dimension disappeared in the four-factor model. Despite its importance in evaluating the concept it aimed to assess, this dimension has consistently shown poor results in terms of validity and reliability in previous studies. Concurrently, a new dimension emerged in exploratory analyses named “learning to think critically,” referring to empowerment’s fifteen attributes defined by the original scale’s user group [14]. Although these new dimensions are epistemologically interesting, their empirical validation requires rigorous analysis to confirm construct capture. The study design permitted only empirical evaluation of the self-esteem and autonomy dimensions and the overall ES index.

As expected, the index correlated positively with recovery, quality of life, and therapeutic alliance concepts—assessed by the RAS, S-QOL, and 4-PAS, respectively—and negatively with symptomatology, estimated by the MCSI. However, these correlations were lower than hypothesized. Only the self-esteem dimension showed significantly high correlations with similar constructs from other scales, confirming its concurrent validity. In contrast, autonomy exhibited only a moderate correlation with the related S-QOL dimension (“I am free to act” and “I am free to make decisions”), indicating a difference in measurement approach. Item analysis revealed that the F-ES dimension reflected broader opinions on autonomy (“People should try to live their lives as they want” and “People have the right to make their own decisions, even if they are not the right ones”), while S-QOL directly measured autonomy.

To confirm the concurrent validity of the other dimensions, community activism and critical thinking, they should be compared with established scales. However, “critical thinking” stood out due to significant differences in educational level, particularly between those with at least a Baccalauréat diploma and others. Studies have noted a link between critical thinking and education [63]. Our findings suggest that more educated psychiatric patients may find it easier to “learn to think critically; unlearn the conditioning; see things differently” [14], as defined by the user group who designed the ES. This dimension showed weak unidimensionality according to the H coefficient despite having only three items. The other dimensions had moderate unidimensionality, except for self-esteem, which showed strong unidimensionality. These results raise concerns about construct validity, which is crucial for score interpretation, and dimensionality’s impact on reliability [64], as dimensions measuring multiple constructs can lack precision.

Therefore, it is important to discuss the instrument’s reliability, assessed in this study by the internal consistency of each dimension and the index. The reliability was good for the index and the main dimension but poor to questionable for the other dimensions, similar to findings in the original instrument [61, 65] and its cultural adaptations in Sweden [66], the Netherlands [67], and Portugal [62]. In Japan, only index consistency data are available [68].

Cronbach’s alpha is sensitive to random errors, especially when dimensions contain few items. These errors may result from respondents’ misunderstanding of items [69]. Some ES items could be clearer, such as “You can’t fight City Hall”, which was translated as “Tu ne peux pas lutter contre l’autorité.” Therefore, the internal consistency in this study should be interpreted with caution. A reproducibility study would provide valuable insights into the instrument’s reliability.

Nevertheless, the inability to demonstrate the reliability of three out of four dimensions led us to discuss the instrument’s responsiveness. An imprecise measure may struggle to detect real changes, as random error introduces noise that can obscure these changes. The “self-esteem” dimension appears to be most responsive to variations in users’ recovery scores, unlike the “autonomy” and “critical thinking” dimensions, which show no responsiveness. An instrument meant to assess changes over time, such as the ES, must detect these changes.

Some measures are less reactive due to their design, particularly those with items unlikely to change in response to an intervention [14]. The self-esteem dimension of the F-ES includes factual, first-person items (“I am often able to overcome obstacles”), while other dimensions include opinion-based items (“By working together, people can make an impact on their community”). Opinions are likely less prone to change than behaviors, suggesting that the original design of the ES may affect its overall responsiveness. For a measure to be responsive, it must be reliable and include items addressing aspects of the construct that are likely to change [70].

Limitations

It is important to note that the sample is not representative of all mental health service users. Participants were voluntarily recruited from seven hospitals in three major French cities upon referral by their psychiatrist. While this ensured alignment with the primary objective of the RCT—evaluating the effectiveness of PADs facilitated by peer workers in reducing involuntary hospitalizations—it introduced selection bias. The inclusion criteria focused on individuals with diagnoses most frequently associated with coercive hospitalizations [71] and required a history of involuntary hospitalization within the past year. This targeted approach allowed the study to address a population with significant challenges in decision-making capacity but limits the generalizability of findings. Despite these limitations, the ES has been validated in diverse international contexts, including community self-support centers and day hospitals, which supports its broader applicability. Future research should examine the French scale’s performance in more heterogeneous populations and across a wider range of diagnostic groups to better understand its utility and adaptability.

Practical implications of the F-ES and perspectives

The potential applications of the French Empowerment Scale (F-ES) are particularly significant in the context of patient-centered and recovery-oriented mental health services. By capturing empowerment—a cornerstone of recovery—the scale offers a valuable tool for clinical decision-making, program evaluation, and policy development [61]. For instance, its dimensions addressing self-esteem and autonomy can guide personalized care plans, while its community activism dimension may inspire group interventions fostering social participation. In therapeutic contexts, the critical thinking dimension emphasizes the role of reflective practices and education in empowering service users [63].

In the French mental health care system, which increasingly emphasizes outpatient care and mobile mental health teams, the F-ES aligns with broader human rights principles such as autonomy, self-determination, and inclusion. By systematically assessing empowerment, mental health professionals can identify barriers to recovery and tailor interventions to address users’ unique needs, contributing to the democratization of mental health care [26]. Moreover, the F-ES can serve as a framework for evaluating the recovery orientation of mental health services, providing insights into the impact of service transformations on user empowerment.

Future research could explore the application of the F-ES in community settings, particularly its potential to assess multidisciplinary team interventions and peer-led recovery initiatives [72]. Such studies would help to further validate its utility in promoting recovery-oriented practices and aligning mental health care with international human rights standards.

Finally, this study holds significant interest for an international audience by demonstrating how adapting the ES with an improved factor structure yields better psychometric results compared to previous versions. These findings can inspire further international efforts to refine and enhance the scale, making it more applicable across diverse cultural and clinical contexts.

Conclusion

The initial question was about the psychometric validity of the French version of the ES. After removing the items, the tool showed a good factor structure. Concurrent validity and reliability are good for the “self-esteem” dimension but need further demonstration or improvement for the other dimensions. Responsiveness is moderate for “self-esteem” but low or nonexistent for the other dimensions.

The F-ES, with its 4 factors and 18 items, has partial psychometric validity. The main dimension shows robustness, while the others require further investigation. In summary, while a French version of the ES now exists, ongoing research is essential to enhance the measurement of this multifaceted concept.