Submitted:
22 January 2025
Posted:
23 January 2025
You are already at the latest version
Abstract
Syntactic complexity plays a critical role in assessing language proficiency, yet identifying measures that effectively distinguish proficiency levels remains challenging. This study investigates the relationship between objective measures of syntactic complexity and subjective ratings in the Test of English for Educational Purposes speaking assessment. Focusing on three metrics Mean Length of AS-Unit, Mean Length of Clause, and Subordinate Clauses per AS-Unit the research evaluates their ability to correlate with subjective ratings of grammatical range and accuracy and differentiate proficiency levels. Using a quantitative approach, the study analyzed 89 TEEP speaking test transcriptions paired with proficiency scores ranging from 5.0 to 7.5. Statistical analyses revealed weak but significant positive correlations between MLAS and MLC with subjective ratings, suggesting that these measures contribute modestly to evaluative outcomes. However, SCP-AS showed no significant correlation. ANOVA results highlighted that while MLAS distinguished between lower and higher proficiency levels, it struggled to differentiate adjacent levels. MLC and SCP-AS showed limited variation across levels, indicating low sensitivity. The findings suggest refining TEEP rating scales to assess complexity and accuracy separately and incorporating diverse syntactic, lexical, and morphological measures to enhance assessment validity. Future research should expand metrics and task types to deepen understanding of syntactic complexity in language assessment contexts.
Keywords:
Introduction
2. Literature review
2.1. Conceptualizing L2 Complexity in SLA
2.2. Defining L2 Complexity
2.3. Measures of Syntactic Complexity
2.4. Complexity in Language Proficiency Assessments
2.5. Research in L2 Speech Complexity and Proficiency
- To what extent do objective measures of syntactic complexity in dialogic tasks correlate with subjective ratings of range and accuracy in the TEEP speaking assessment?
- To what extent can objective measures of syntactic complexity reliably distinguish between different proficiency levels in the TEEP speaking paper?
3. Methodology
3.1. Research Design
3.2. Data Set
3.3. Participants
3.4. Data Preparation and Coding
3.5. Complexity Measures
3.6. Data Analysis Tools
4. Analysis and Results
4.1. Introduction
4.2. RQ1: Correlation with Subjective Ratings
4.3. RQ2: Differentiation of Proficiency Levels
4.3.1. Mean Length of AS-Unit
4.3.2. Mean Length of Clause
4.3.3. Subordinate Clauses per AS-Unit
5. Discussion
5.1. Summaryof Findings
5.2. RQ 1: Correlation between Syntactic Complexity and Subjective Ratings
5.3. RQ 2: Differentiation of Proficiency Levels by Syntactic Complexity Measures
5.4. Implications of Findings
5.4.1. TEEP Rating Scale Enhancement
- Enhanced diagnostic value, enabling more precise identification of specific language proficiency strengths and weaknesses.
- Improved fairness by ensuring that both complexity and accuracy are evaluated independently, preventing one from overshadowing the other.
- Tailored instructional feedback, which would offer more actionable insights for students.
- Increased reliability in assessment by reducing subjective bias and improving consistency across raters and test administrations.
5.4.2. Enhancing TEEP Task Validity
5.4.3. Advancing Proficiency Assessments
6. Conclusion
6.1. Summary of Findings
6.2. Contributions and Implications for Future Research
6.3. Limitations of the Study
- This study is limited by its narrow focus and context-specific nature, which may restrict the generalizability of the findings to a broader population.
- The sample size, particularly within each proficiency subgroup, was small, which may affect the stability and representativeness of the results. Future studies should aim to include larger sample sizes to enhance the generalizability of the findings.
- Additionally, this research focused on a limited set of syntactic complexity measures, and future studies could explore a broader range of fine-grained measures to capture complexity more effectively across different proficiency levels and task types.
- The study also relied on data derived from a dialogic speaking task, which may differ from tasks that require more formal or prepared speech. Future research could examine a variety of task contexts to provide a more comprehensive understanding of syntactic complexity’s role in proficiency assessment.
Acknowledgments
Conflicts of Interest
References
- Bachman, L. (1990). Fundamental considerations in language testing. Oxford: Oxford University Press.
- Bardovi-Harlig, K. (1992). A second look at t-unit analysis: Reconsidering the sentence. TESOL Quarterly 26, 390–395.
- Biber, D., Gray, B., & Staples, Sh., (2016). Predicting patterns of grammatical complexity across language exam task types and proficiency levels. Applied Linguistics 37(5). 639–668.
- Bulte, B., & Housen, A. (2012). Defining and operationalising L2 complexity. In A. Housen, F. Kuiken, & I. Vedder (Eds.), Dimensions of L2 performance and proficiency. Investigating complexity, accuracy and fluency in SLA (pp. 21-46). Amsterdam: John Benjamins.
- Bulte, B., & Roothooft, H., (2020) Investigating the interrelationship between rated L2 proficiency and linguistic complexity in L2 speech. Elsevier. [CrossRef]
- Council of Europe, (2014). The Common European Framework of Reference for Languages: Learning, teaching, assessment. https://rm.coe.int/1680459f97.
- Dörnyei, Z. (2007). Research methods in applied linguistics: quantitative, qualitative, and mixed methodologies. Oxford University Press.
- Ehret, K., Berdicevskis, A., Bentz, C., & Blumenthal-Dramé, A. (2023). Measuring language complexity: Challenges and opportunities. Linguistics Vanguard, 9(s1), 1–8. [CrossRef]
- Ellis, R. and G. Barkhuizen. 2005. Analysing learner language. Oxford: Oxford University Press.
- Foster, P., Tonkyn A. and G. Wigglesworth. 2000. Measuring spoken language: A unit for all reasons. Applied Linguistics 21: 354–75.
- Foster, P., Tonkyn A. and G. Wigglesworth. 2000. Measuring spoken language: A unit for all reasons. Applied Linguistics 21: 354–75.
- Foster, P., and Skehan, P. (1996). The influence of planning and task type on second language performance. Studies in Second Language Acquisition, 18 (3) (1996), pp. 299-323.
- Fulcher, G. (2010). Practical language testing. London: Hodder Education.
- Green, R. (2013). Statistical analyses for language test developers. Basingstoke, UK: Palgrave Macmillan.
- Harsh, C. (2016). Proficiency. ELT Journal. 71(2). [CrossRef]
- Housen, A. and F. Kuiken. 2009. Complexity, accuracy, and fluency in second language acquisition. Applied Linguistics 30(4): 461–473.
- Housen, A., & Pierrard, M. (2009). Complexity, Accuracy and Fluency in Second Language Acquisition. Applied Linguistics. https://www.academia.edu/28771030/Complexity_Accuracy_and_Fluency_in_Second_Language_Acquisition.
- Housen, A., Kuiken, F., & Vedder, I. (2012). Complexity, accuracy, and fluency: Definitions, measurements and research. In A. Housen, F. Kuiken & I. Vedder (Eds.), Dimensions of L2 performance and proficiency: complexity, accuracy and fluency in SLA (pp. 1-20). Amsterdam: John Benjamins.
- Hulstijn, J. H. (2011) Language Proficiency in Native and Nonnative Speakers: An Agenda for Research and Suggestions for Second-Language Assessment. Language Assessment Quarterly. 8:3, 229-249.
- Iwashita, N., Ortega, L., Rabie, S., & Norris, J. M. (2008). Syntactic complexity and oral proficiency in crosslinguistic perspective. Honolulu, HI: University of Hawai‘i National Language Resource Center.
- Kang, O., Yan, X., 2018. Linguistic Features Distinguishing Examinees’ Speaking Performances at Different Proficiency Levels. Journal of Language Testing & Assessment. 1:29-31. [CrossRef]
- Kuiken, F. (2023). Linguistic complexity in second language acquisition. Linguistics Vanguard, 9(s1), 83–93. [CrossRef]
- Kuiken, F., Vedder, I., Housen, A., & De Clercq, B. (2019). Variation in syntactic complexity: Introduction. International Journal of Applied Linguistics, 29(2), 161–170. [CrossRef]
- Larsen-Freeman, D. (2009). Adjusting Expectations: The Study of Complexity, Accuracy, and Fluency in Second Language Acquisition. Applied Linguistics. 30(4), 579–589. [CrossRef]
- Norris, J. M. and L. Ortega., (2009). Towards an organic approach to investigating CAF in instructed SLA: The case of complexity. Applied Linguistics 30(4): 555–578.
- Ortega, L. (2003).. Syntactic Complexity Measures and their Relationship to L2 Proficiency: A Research Synthesis of College-level L2 Writing. Applied Linguistics. 24/4: 492-518.
- Pallotti, G. (2009). CAF: Defining, refining and differentiation constructs. Applied Linguistics. 30(4), 590-601. [CrossRef]
- Pallotti, G., (2015). A simple view of linguistic complexity. Second Language Research. 31(1) 117–134.
- Plonsky, L., & Oswald, F. L. (2014). How big is “big”? Interpreting effect sizes in L2 research. Language Learning. 64, 878–912. [CrossRef]
- Seedhouse, P., Harris, A., Naeb, R., & Üstünel, E. (2014). The relationship between speaking features and band descriptors: A mixed method study. IELTS Research Reports Online Series, 2, 1e30.
- Skehan, P. (2009). Modelling Second Language Performance: Integrating Complexity, Accuracy, Fluency, and Lexis. Applied Linguistics, 30(4), 510–532. [CrossRef]
- Skehan, P. 1998. A Cognitive Approach to Language Learning. Oxford University Press.
- TEEP Candidate Handbook, (2023). University of Reading. https://www.reading.ac.uk/isli/english-language-tests/teep.
- Vercellotti, M. L. (2015). The Development of Complexity, Accuracy, and Fluency in Second Language Performance: A Longitudinal Study. Applied Linguistics. [CrossRef]
- Wolfe-Quintero, K., Inagaki, S., & Kim, H. Y. (1998). Second language development in writing: Measures of fluency, accuracy, and complexity (Tech. Rep. No. 17). Honolulu: National Foreign Language Resource Center.
| TEEP | IELTS | TOEFL iBT |
| Fluency and coherence | Fluency and coherence | Delivery |
| Lexical resource | Lexical resource | Language use |
| Grammatical range and accuracy | Grammatical range and accuracy | Topic development |
| Pronunciation | Pronunciation | - |
| Interactive communication | - | - |
| Part | Task | Mode | Description | Planning time | Response time |
| 1 | Focus/topic introduction | Silent preparation | Examinees prepare silently before responding to a prompt question |
20 seconds |
__ |
| 2 |
Individual talk (role plays) |
Monologue |
Examinees engage in role plays discussing advantages or disadvantages of a given topic |
4 minutes |
3 minutes |
| 3a | Scenario discussion | Dialogue | Examinees discuss specific scenarios | 2 minutes | 4 minutes |
| 3b | Further discussion | Dialogue | Examinees analyze and discuss the focus question | __ |
__ |
| PROFICIENCY LEVEL | NUMBER OF PERFORMANCES |
| 5.0 | 13 |
| 5.5 | 12 |
| 6.0 | 9 |
| 6.5 | 15 |
| 7.0 | 19 |
| 7.5 | 21 |
| TOTAL | 89 |
| Complexity measures | Calculation |
| Mean length of AS unit (MLAS) | Total number of words divided by total number of AS units |
| Mean length of clause (MLC) | Total number of words divided by total number of clauses |
| Subordinate clauses per AS-unit | Total number of clauses divided by total number of AS units |
| N | Minimum | Maximum | Mean | Std. Deviation | |
| MLAS | 89 | 6.87 | 19.20 | 12.2803 | 2.34848 |
| MLC | 89 | 4.23 | 10.00 | 6.5671 | 1.07250 |
| SCP-AS | 89 | 1.17 | 3.42 | 1.8920 | .40881 |
| Proficiency Level | 89 | 5.0 | 7.5 | 64.38 | 8.849 |
| Level | MLAS | MLC | SCP-AS | ||
| Level | Pearson Correlation | 1 | .298** | .231* | -.018 |
| Sig. (2-tailed) | .005 | .029 | .865 | ||
| N | 89 | 89 | 89 | 89 | |
| MLAS | Pearson Correlation | .298** | 1 | .284** | .667** |
| Sig. (2-tailed) | .005 | .007 | <.001 | ||
| N | 89 | 89 | 89 | 89 | |
| MLC | Pearson Correlation | .231* | .284** | 1 | -.427** |
| Sig. (2-tailed) | .029 | .007 | <.001 | ||
| N | 89 | 89 | 89 | 89 | |
| SCP-AS | Pearson Correlation | -.018 | .667** | -.427** | 1 |
| Sig. (2-tailed) | .865 | <.001 | <.001 | ||
| N | 89 | 89 | 89 | 89 |
| N | Mean | Std. Deviation | Minimum | Maximum | |
| 50 | 13 | 10.7608 | 1.93790 | 6.87 | 14.00 |
| 55 | 12 | 12.1417 | 2.32045 | 9.40 | 16.30 |
| 60 | 9 | 11.9622 | 2.64345 | 9.30 | 18.00 |
| 65 | 15 | 12.6133 | 2.08151 | 10.40 | 17.00 |
| 70 | 19 | 12.0158 | 2.11247 | 9.10 | 16.70 |
| 75 | 21 | 13.4381 | 2.46850 | 10.30 | 19.20 |
| Total | 89 | 12.2803 | 2.34848 | 6.87 | 19.20 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).