Face validity of a dyslexia screening instrument for secondary students
Plos.org·July 20, 2026
AI Summary
This study examines whether a dyslexia screening tool accurately measures what it intends to assess in secondary school students. The research addresses the gap in dyslexia identification at the secondary level, where many cases go undiagnosed due to insufficient screening practices.
Dyslexia is a specific learning difficulty marked by persistent challenges in reading accuracy, fluency, and word recognition despite adequate intelligence and instruction. In secondary schools, many cases remain unidentified due to the absence of screening instruments designed for older students. Early identification is essential for timely intervention, which requires valid and reliable tools. This study evaluated the face validity of a newly developed dyslexia screening instrument for secondary students through expert and test-taker assessments. The 28-item instrument, which focused on cognitive manifestations of dyslexia, was reviewed by ten experts and ten students. Items were rated on a 4-point Likert scale assessing clarity and comprehensibility, with analyses conducted using the Item-Level Face Validity Index (I-FVI), Scale-Level Face Validity Index by Average Agreement (S-FVI/AVE), and Scale-Level Face Validity Index by Universal Agreement (S-FVI/UA). Expert evaluation demonstrated that 85.7% of items achieved full agreement (I-FVI = 1.00), while scale-level indices exceeded recommended benchmarks (S-FVI/AVE = 0.96; S-FVI/UA = 0.86). Student evaluation also yielded high scale-level agreement (S-FVI/AVE = 0.92; S-FVI/UA = 0.82). A small number of items required linguistic refinement. Consequently, the findings indicate high agreement on the instrument’s clarity, comprehensibility, and acceptability, supporting face validity as a preliminary validation step for subsequent psychometric validation.
Citation: N. Mohd Nabil NZ-I, Mohd Matore MEE, Zainal MS (2026) Face validity of a dyslexia screening instrument for secondary students. PLoS One 21(7): e0353781. https://doi.org/10.1371/journal.pone.0353781
Editor: Fucai Lin, Minnan Normal University, CHINA
Received: January 13, 2026; Accepted: June 28, 2026; Published: July 20, 2026
Data Availability: All relevant data are within the manuscript and its Supporting information files.
Funding: This work was supported by the Faculty of Education, The National University of Malaysia under research grant number GP-K021854. The funders had role in study design, data collection and analysis, and decision to publish.
Competing interests: The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Early identification of dyslexia is crucial in both educational and clinical settings [1–3], as it enables timely intervention that can significantly enhance academic outcomes and cognitive development. Research has revealed that early screening, particularly during preschool years [1,4,5], can effectively identify children at risk of dyslexia, allowing for targeted interventions that improve literacy skills and long-term educational success [6,7]. Without early detection, children with dyslexia often experience persistent reading difficulties, leading to academic underachievement and negative socio-emotional consequences [8–11]. This issue is directly linked to Sustainable Development Goal (SDG) 4: Quality Education, which emphasizes inclusive and equitable education, ensuring that all learners, regardless of learning difficulties, have access to appropriate support mechanisms [12]. Developing reliable and valid screening tools is therefore essential in achieving this global objective. Given the profound impact of early screening, the development of reliable and valid dyslexia screening tools remains a priority in educational psychology and special education [13,14].
Despite the availability of various dyslexia screening instruments, several challenges remain. Existing tools may demonstrate inconsistencies in identifying students at risk across different linguistic and cultural contexts. [15,16]. In addition, misconceptions about dyslexia among educators and practitioners may contribute to misidentification and ineffective intervention strategies. Additionally, research indicates that educators and professionals may hold misconceptions about dyslexia, potentially leading to misdiagnosis and ineffective intervention strategies. [17–19]. Recent narrative reviews of school-based dyslexia screening tools have also highlighted considerable variation in the instruments used across educational contexts, with no clear consensus on the most appropriate screening approach. The diversity of tools, as well as differences in reported sensitivity and specificity, indicates the need for contextually appropriate screening instruments tailored to specific educational populations [20].
To address this gap, the Face Validity Index (FVI) provides a structured quantitative approach for evaluating item clarity and comprehensibility through systematic expert and user appraisal [21]. Instruments with high face validity are more likely to be trusted, accepted, and effectively implemented in real-world educational and clinical settings [22]. Although the FVI has been increasingly applied in health and educational measurement studies, its application in the development of dyslexia screening instruments, especially for secondary school populations, remains limited. Most existing dyslexia screening tools focus primarily on early childhood or primary school populations, leaving a gap in instruments specifically designed for adolescents whose reading difficulties may have been previously unidentified. In addition, few studies have incorporated both expert judgement and test-taker perspectives within the face validity evaluation process, despite recommendations advocating stakeholder involvement during early stages of instrument development [23]. Therefore, this study contributes to the literature by developing and evaluating the face validity of a dyslexia screening instrument tailored for secondary school students, addressing an important gap in current screening practices.
Accordingly, this study focuses on face validity as a foundational phase of instrument development, conducted prior to subsequent construct and reliability validation. The study aims to evaluate the face validity of a newly adapted dyslexia screening instrument for secondary school students using the Face Validity Index. Item clarity, relevance, and comprehensibility were examined through structured assessments involving expert educators and secondary school students as test takers. This investigation represents the initial stage of a broader validation process, providing empirical evidence of user acceptability and linguistic suitability before advanced psychometric analyses are undertaken. By establishing foundational validity, the study contributes to the development of accessible and contextually appropriate screening tools that support early identification and inclusive educational practices aligned with SDG 4.
Validity is a cornerstone concept in psychometrics, referring to the extent to which an assessment tool accurately evaluates the construct it is intended to measure [24,25]. It encompasses various forms, including construct validity, content validity, criterion-related validity, and face validity, each serving a unique purpose in the evaluation of an instrument’s effectiveness. Notably, construct validity ensures that the tool accurately captures the theoretical concept under investigation, while content validity ensures that the test items adequately represent the domain of interest [26]. In educational and psychological assessments, validity is critical for establishing the reliability of data-driven decisions and interventions.
Face validity, defined as the extent to which an assessment appears relevant and understandable to non-expert users, such as educators, parents, and even students, significantly influences the tool’s acceptability and implementation [27,28]. Although face validity does not provide statistical evidence of measurement accuracy, it plays a crucial role in determining user acceptance, compliance, and real-world applicability. Instruments that lack face validity may be rejected by practitioners or misunderstood by respondents, reducing their effectiveness regardless of underlying psychometric strength [22].
In educational contexts, face validity is particularly salient for screening tools used in school settings, where teachers and students are required to engage with items directly. For students with learning difficulties such as dyslexia, linguistic clarity and cognitive accessibility are essential to prevent response bias and misinterpretation [29,30]. Consequently, face validity serves as a bridge between technical measurement design and functional usability.
Despite its importance, face validity has historically been underreported or treated informally in instrument development studies, often described qualitatively without systematic evaluation. Recent methodological literature, however, has highlighted the need for structured and transparent approaches to face validity assessment, particularly during early stages of scale development [23,31].
The Face Validity Index (FVI) provides a quantitative framework for evaluating face validity through systematic ratings of item clarity and comprehensibility by experts and intended users [21]. Unlike purely descriptive approaches, the FVI offers numerical indices at both item level (Item-Level Face Validity Index, I-FVI) and scale level (Scale-Level Face Validity Index, S-FVI), allowing researchers to make evidence-informed decisions regarding item retention, refinement, or revision.
Although no single scholar is universally recognized as the originator of the FVI, its structured application has been clearly articulated in methodological work by [21], who outlined standard procedures for rating, recoding, and interpreting face validity data. The FVI is a quantitative approach to assess the degree to which test items appear effective [32], valid [33], and relevant to stakeholders, such as educators, parents, and clinicians. Furthermore, the FVI is also referred to by various synonymous terms in the literature, including face validity score [30] and expert agreement index [34], particularly in linguistic and usability studies [35]. It involves evaluators rating each item based on clarity, relevance, and comprehensibility, with the FVI calculated as the proportion of items deemed valid by the assessors [21]. Consequently, this metric provides an empirical basis for determining whether an instrument is likely to be accepted and utilized effectively in real-world settings. In the context of dyslexia screening, a high FVI indicates that the tool is perceived as appropriate and user-friendly, which is essential for its successful implementation in educational environments.
Moreover, several studies [36–39] have employed the FVI to enhance the development of educational and psychological assessments. For instance, the Dyslexia Adult Checklist underwent a validation process where items were evaluated for clarity and relevance, leading to a refined instrument that effectively identifies dyslexia in adults [13]. Similarly, the Teacher Teaching Quality Instrument using the Six Sigma approach (T2Qi-6σ) was subjected to face validity assessment [40], resulting in a tool that accurately measures teaching quality and supports SDGs related to quality education. These examples underscore the importance of incorporating FVI in the validation process to ensure that assessment tools are both effective and aligned with educational objectives.
Conversely, incorporating the FVI in the development of dyslexia screening instruments is particularly crucial due to the diverse range of stakeholders involved, including teachers, parents, and policymakers. A screening tool with a high FVI is more likely to be embraced by educators and parents, facilitating early identification and intervention for children with dyslexia. This supports inclusive education efforts [12]. Thus, by utilizing the FVI, developers can create culturally sensitive and contextually appropriate tools that address the specific needs of various populations, thereby promoting educational equity and supporting the global agenda for sustainable development.
Dyslexia screening instruments are brief, targeted tools designed to identify individuals who may exhibit early signs of dyslexia, allowing for timely diagnostic referral and intervention. However, unlike full diagnostic assessments, which are often resource-intensive and conducted by clinical professionals, screening instruments function as an initial gatekeeping mechanism to detect at-risk learners in educational settings [17,41]. Furthermore, effective dyslexia screening tools must demonstrate both psychometric robustness and practical relevance. This includes content that reflects the core cognitive markers of dyslexia, such as phonological processing deficits, working memory limitations, slow naming speed, and orthographic confusion [6], while remaining accessible to non-specialist users such as teachers. Moreover, face validity is a critical aspect of screening design, as it determines whether items appear, on the surface, to be relevant, understandable, and appropriate to both experts and target respondents [21,29].
Recent trends in dyslexia screening emphasize the need for tools that are developmentally appropriate beyond the early years, particularly for late-identified or late-emerging dyslexia cases in secondary schools [42,43]. Thus, screening instruments must be validated statistically, and through user-centric methods such as face validity evaluation involving both expert panels and student feedback. Additionally, such validation ensures that the instrument captures the lived realities of diverse learners and aligns with inclusive education priorities under SDG 4 [12].
The development of the dyslexia screening instrument in this study is conceptually grounded in Morton and Frith’s causal model of developmental disorders, which proposes three interrelated levels of explanation: biological, cognitive, and behavioral [44]. According to this framework, dyslexia is not attributable to a single deficit. However, it emerges from interactions across these levels, with the cognitive level serving as the most direct pathway through which difficulties in reading and related skills are expressed. The development of the dyslexia screening instrument was conceptually grounded in Morton and Frith’s causal model of developmental disorders, which emphasizes biological, cognitive, and behavioral levels of explanation. In the present study, particular attention was given to the cognitive level, as this level most directly reflects the learning processes associated with dyslexia in educational contexts. Therefore, although the broader screening framework includes multiple domains, the analysis presented in this paper focuses specifically on items representing cognitive manifestations, such as phonological awareness, working memory, and reading fluency.
The findings indicate that items representing cognitive manifestations of dyslexia were generally perceived as clear and comprehensible by experts and student respondents. This supports the linguistic and conceptual appropriateness of the cognitive-domain items at the preliminary face validity stage. However, further psychometric testing is required before conclusions can be drawn regarding the instrument’s ability to identify students at risk of dyslexia [6,45]. Within this scope, the findings suggest that Morton and Frith’s cognitive-level framework provides a relevant theoretical basis for guiding item development related to dyslexia-associated learning processes in secondary school contexts.
Moreover, situating the instrument within this theoretical framework underscores its relevance beyond item-level clarity by linking early-stage instrument development to broader educational equity concerns. Within the scope of face validity, the findings suggest that theory-driven item development can help ensure that cognitive-domain items are linguistically clear, conceptually aligned, and accessible to intended users. This is consistent with current calls to integrate cognitive-level insights into educational policy and practice to reduce systemic barriers in dyslexia identification [46,47]. In this way, the present study provides preliminary face validity evidence for cognitive-domain items and contributes to the early development of contextually appropriate assessment resources that may support inclusive education goals aligned with SDG 4 [12], pending further psychometric validation.
This study employed an instrument validation research design using a survey among experts and students as test takers, with a specific focus on face validity assessment. Face validity, a crucial psychometric property, measures the degree to which a screening instrument appears appropriate and relevant to its intended users [22]. Furthermore, given the role of early dyslexia detection in ensuring effective interventions, this study validates the FVI of a newly adapted dyslexia screening instrument. The validation process involved a panel of experts who assessed the instrument’s clarity, relevance, and comprehensibility, ensuring its alignment with real-world application in educational and clinical settings [37].
Ethical considerations were observed throughout the face validity assessment process. The study involved expert review for instrument validation purposes only and did not include intervention, experimental manipulation, or collection of sensitive personal information. Ethical approval for data collection involving school-based participants was obtained from the Educational Policy Planning and Research Division, Ministry of Education Malaysia (Approval No: KPM.600-3/2/3-eras (23518), dated 16 Sept 2024), prior to participant recruitment and data collection. Data collection was conducted between September and December 2024.
Institutional ethical clearance was subsequently granted by the UKM Research Ethics Secretariat (Ref No: JEP-2025-416), covering the broader PhD research project. All procedures were conducted in accordance with approved protocols and relevant ethical guidelines. Written informed consent was obtained from all participants prior to data collection. For participants under the age of 18, written consent was additionally obtained from their parents or legal guardians. Participation was voluntary, and participants were informed that their responses would remain confidential and used solely for research purposes.
Prior to the face validity evaluation, the instruments were initially developed in English and underwent a forward-backward translation process to ensure linguistic and contextual equivalence in the Malay version. Hence, to enhance the accuracy and cultural relevance of the translated items, the instrument was reviewed by two certified Language experts, recognized for their expertise in Malay language standardization. It is important to note that the two language experts involved in the translation review were not part of the FVI evaluation panel. Their involvement ensured that the translation maintained semantic equivalence, preventing potential misinterpretations that could affect the validity of the screening tool [22]. This preliminary linguistic validation ensured that the items were semantically accurate, accessible, and appropriate for the target population. Moreover, their inclusion aligns with the principle that face validity refers to the degree to which the assessed items on an instrument accurately correspond to the intended constructs and objectives of the study [48]. Additionally, the data were collected via email over one month.
Subsequently, the verified version was distributed to a panel of ten relevant experts, comprising secondary and primary school teachers. Each had more than ten years of teaching experience [48] in the Malaysian education system and prior involvement in curriculum delivery to students with diverse learning needs. Following the linguistic validation, these expert practitioners were purposively selected based on their pedagogical expertise with at least five years of professional experience and familiarity with the learning profiles of pupils in mainstream and special education settings to refine the instrument further. Notably, the quality of the raters is important, including experience, training experience, and teaching experience [49]. Their primary role was to evaluate the clarity, appropriateness, and comprehensibility of the instrument, particularly in terms of language use and sentence structure. Specifically, each expert was provided with a validation form and rated the instrument items using a 7-point Likert scale, allowing for a quantitative assessment of face validity [21].
However, for the purposes of computing the FVI, the original 7-point Likert responses were recoded into a 4-point format in accordance with methodological recommendations for FVI analysis [20]. Responses representing higher levels of agreement were collapsed to reflect clarity and appropriateness of the items. Specifically, ratings of 1–3 were recoded as 1 (not clear/not appropriate), 4 as 2 (somewhat clear), 5 as 3 (clear), and 6–7 as 4 (very clear/very appropriate). Following established FVI procedures, only ratings of 3 and 4 were considered indicators of agreement when calculating the Item-Level Face Validity Index (I-FVI). This recoding approach ensures alignment with commonly recommended thresholds for face validity evaluation while preserving the interpretability of respondents’ judgments. This expert review process ensured that the instrument was accessible and easily understood by both educators and students, reducing the risk of ambiguity or misinterpretation. Moreover, the heterogeneity of the expert panel allows for a comprehensive evaluation of the instrument’s applicability across different contexts, supporting its usability across diverse linguistic and cultural backgrounds, a key factor in SDG 4 [50]. Table 1 below presents the details of the respective experts.
https://doi.org/10.1371/journal.pone.0353781.t001
The student panel consisted of ten secondary school students representing the target population of the screening instrument. The students were aged 14 years old and included 4 males and 6 females. The participants were selected to reflect the characteristics of the intended respondents in terms of educational level and language proficiency. Basic demographic information of the participants, including age range, gender distribution, and school level, was recorded to ensure that the feedback reflected the perspectives of the target user group. In face validity studies, a small panel of respondents is generally considered sufficient because the objective is to evaluate item clarity and interpretability rather than to perform statistical generalization. Previous methodological guidelines recommend panels of approximately 5–10 respondents for preliminary face validity assessment [21]. The instrument was subsequently evaluated by ten secondary school students representing the intended end-user group to complement expert-based face validity assessment. The primary purpose of this process was to gather insights directly from test takers regarding the clarity, appropriateness, and interpretability of the item wording, structure, and instructions. Notably, this step is critical in face validation, as it ensures the instrument resonates linguistically and cognitively with its actual users and minimizes the risk of misinterpretation or response bias [23]. Furthermore, the students were instructed to complete the instrument and identify any unclear words or phrases. This was done by a face-to-face approach. [51] mentioned that the face-to-face method is highly effective in improving response rates, while online surveys offer greater efficiency in terms of cost and time. Their feedback was gathered using a 7-point Likert scale to indicate whether they understood each item correctly. Nevertheless, to facilitate the calculation of the FVI, the original scale was transformed into a 4-point Likert format in accordance with widely accepted psychometric conventions. These ratings were used to compute the Item-Level Face Validity Index (I-FVI) and further evaluated through the Scale-Level Face Validity Index by Average Agreement (S-FVI/AVE) and Scale-Level Face Validity Index by Universal Agreement (S-FVI/UA). Moreover, feedback from the students also provided valuable qualitative insights, allowing researchers to refine wording and structure in cases where ambiguity or confusion was observed. This step was crucial in assessing whether the instrument was appropriately tailored to the cognitive and linguistic capabilities of Malaysian students, ensuring that the screening tool effectively serves its intended purpose [25,52]. Consequently, this multi-step validation process enhances the face validity of the dyslexia screening instrument by incorporating expert evaluations and direct feedback from test takers.