Submitted:
19 August 2026
Posted:
21 August 2026
You are already at the latest version
Abstract
Automatic speech recognition (ASR) evaluation for Garrusi, a Kurdish variety written in a Latin-based orthography, is complicated when an ASR system outputs Arabic script because direct scoring can confuse recognition errors with writing-system differences. This study evaluates MMS-1B-all with the Central Kurdish adapter on 1,722 speech segments from five Garrusi speakers (9,763 reference tokens; 117.9 minutes) without adaptation. A common-reference scoring design was used: the reference was folded once and kept fixed, while the hypothesis representation was changed. The unmodified Arabic-script output produced 111.70% WER and 100.92% CER. Transliteration to Latin reduced these to 102.36% and 57.89%, while additional orthographic folding reduced them to 97.45% and 51.13%. Thus, the full transformation reduced measured WER by 14.25 percentage points and CER by 49.79 points, but substantial error remained. A round-trip control showed that part of the remaining error comes from the transliteration and folding process itself, with 98.5% of its substitutions linked to two orthographic patterns. A separate comparison with a Southern Kurdish fine-tuned system showed a 10.69-point WER difference, which fell to 2.13 points when short i was treated as non-contrastive. These results show that orthographic representation can strongly affect ASR evaluation and that recognition performance should be assessed separately from differences introduced by the scoring pipeline.
Keywords:
Garrusi Kurdish
; automatic speech recognition
; Kurdish speech technology
; word error rate
; character error rate
; orthographic normalization
; transliteration
; MMS-1B-all
1. Introduction
Kurdish speech technology has extended beyond its original concentration on Central Kurdish, with dedicated evaluation sets now available for Northern Kurdish, Badini, and Southern Kurdish [10,11], and with recent work addressing variation within Central Kurdish itself [1]. For Southern Kurdish the most substantial of these is the resource of Mohammadamini and Tahon [10], which provides 30 hours of validated read speech from 208 speakers together with a multi-domain benchmark of 100 sentences, translated from an existing Central Kurdish test set and read by eight speakers drawn from five vernaculars, and which releases the corpus, a fine-tuned w2v-BERT CTC model and inference code publicly. That study also reports a result of direct relevance to the configuration evaluated here: a Whisper model fine-tuned for Central Kurdish reaches 84.23 WER on their Southern Kurdish benchmark against 13.40 WER on the Central Kurdish version of the same sentences, which they read as evidence that systems developed for one Kurdish dialect transfer poorly to another. Coverage remains uneven, and the unevenness appears to follow the field’s dialect labels: varieties not straightforwardly covered by the Sorani/Kurmanji labelling are correspondingly less visible in the published record [2,10,11]. The dependence on labels is unsurprising given that the classification of Kurdish varieties is itself unsettled and is drawn on different criteria by different authors [4,8,9]; §2 sets out the background. Garrusi is one variety that falls outside the two best-covered labels: it is described in recent linguistic work as a minority Kurdish variety of Iran [5] and it is not listed separately in the classification I cite here. Garrusi is not entirely absent from that record: the Southern Kurdish benchmark of Mohammadamini and Tahon [10] includes one speaker labelled Garrusi among its eight, contributing 95 of its 773 recordings, alongside speakers of the Kermānshāhi, Kalhori, Malekshāhi and Kolyā’i vernaculars (their Table 2). That study reports character error rate per speaker and does not aggregate results by vernacular. Since four of the five vernaculars are represented by a single speaker, speaker and vernacular effects are not separable within that design, and the authors attribute the observed spread across speakers to speaker-related characteristics and not to variety. I found no evaluation reporting recognition rates specific to Garrusi, and no evaluation set assembled for the variety1.
Two obstacles stand between that gap and a usable first measurement. No Garrusi-trained recognition system exists, so an initial evaluation must transfer a model adapted to a related variety. There is also a measurement problem. The model emits Central Kurdish Arabic script, while my reference transcriptions are in a Latin field orthography (§2), so a direct comparison scores a writing-system difference as recognition error. The standard remedy is to normalize before scoring, but normalization is usually applied to both sides at once, which means the reported gain mixes an actual increase in agreement with a change in the reference tokenization. Recent work on Kurdish ASR has observed that word error rate is inflated relative to character error rate by this class of standardization difference [10,11]. Separating the increase in agreement from the change in reference tokenization is necessary for interpreting a Kurdish WER figure, not a refinement of it.
The pilot reported here addresses both. I evaluate MMS-1B-all [11] with the Central Kurdish adapter on elicited Garrusi questionnaire speech, an adapter chosen because prior linguistic analysis found substantial morphosyntactic overlap between Garrusi and Central Kurdish [5], despite clear phonological and phonetic differences, and score it under a common-reference staged normalization: the reference is folded once, fixed, and reused unchanged across three conditions that differ only in how far the hypothesis has been converted toward it. Because the reference tokenization is constant, the measured differences between conditions arise from the hypothesis-side transformation.
Two further measurements are needed before such a design can be interpreted. The first is a control: a staged design shows how much the measured rate moves under each transformation, but it does not show what the pipeline returns for a hypothesis that is already correct, and without that quantity the residual cannot be read at all. Section 4.6 supplies it, by passing a correct Arabic-script rendering of a sample of the reference through the identical path. The second is a second recognizer, scored through the same pipeline on the same audio and the same segments, since a comparison between systems under a fixed scoring design is interpretable where an absolute level is not (§4.5).
Together these motivate a third analysis, which is in my view the more general result. If a substantial share of what the pipeline measures is traceable to one unwritten vowel, then whether that vowel is treated as contrastive at scoring time is a decision with consequences for what a reported rate means. Section 4.7 makes that decision explicitly and measures its effect on the comparison between the two systems, and §4.8 measures what it costs in lexical distinctness. I present it as a sensitivity analysis rather than as an improved scoring convention, and §3.5 states why it cannot be read as a further stage of the common-reference design.
The fixed reference and the segment-level results will be released alongside the paper, subject to the data-sharing terms of the source corpus (see the Data Availability Statement), so that the design can be checked rather than taken on trust. The result is a zero-shot ASR measurement on this Garrusi evaluation set, together with an explicit statement and full reporting of the scoring design, so that the effect of each hypothesis-side transformation on the reported rate is visible rather than absorbed into a single normalized figure. This is not a benchmark in the sense of a released dataset with a shared protocol.
2. Background: Kurdish Varieties and Orthographic Setting
Descriptive work commonly distinguishes five Kurdish varieties: Northern Kurdish (Kurmanji), Central Kurdish (Sorani), Southern Kurdish, Gorani, and Zazaki [8]. Northern and Central Kurdish are the least disputed members of this grouping, and they are also the two that speech technology has concentrated on.
The boundaries between these groups are not settled. Published classifications differ from one another, combine geographic, historical, social and linguistic criteria in varying proportions, and often do not state which criterion is doing the work. There is no consensus in the literature on how Kurdish should be defined or subdivided [4,8,9]. The classifications differ in how they place Garrusi: the general classification of Haig and Öpengin [8] does not list Garrusi separately, whereas Fattah’s classification, as reported by Belelli [7], lists Bijāri, also known as Garrusi, among the Southern Kurdish subgroups (p. 78). Belelli also describes the Bijār area as a Southern Kurdish enclave in a predominantly Central Kurdish environment (p. 74) and reports lexical items shared with neighboring Central Kurdish varieties (p. 88). The model transferred here is adapted to Central Kurdish.
Central Kurdish is itself internally varied, with regional varieties including Mukri, Hewêlrî, Silêmanî, Germiyanî and Sireyî [2,6,8]. For this experiment, the consequence is interpretive: a model adapter labelled “Central Kurdish” should be understood as adapted to some portion of that range and not to all of it. I make no claim about which Central Kurdish varieties are or are not present in the model’s training data, which I cannot audit.
Garrusi is a Kurdish variety of Iran. The linguistic description I rely on for the speech evaluated here is that of Asadpour and Zarei [5,6], whose speakers were recorded in the Mehraban District of Hamadan Province, where the recordings analyzed here were also collected. The general classification cited above does not list Garrusi separately, and Southern Kurdish is there taken to cover varieties such as Kelhuri, Feyli and Kirmashani, with boundaries against neighboring varieties that are among the less settled in the literature [8]. I have not attempted a survey of how Garrusi is placed across the classification literature, and I therefore use “Garrusi” as the variety label for the speech evaluated here without attempting to resolve its broader classification.
Kurdish is written in more than one script. Central Kurdish is commonly written in an Arabic-based script, but Latin-based conventions are also in use, and orthographic conventions can differ even where the same variety is being represented [1, p. 73,2,3]. The reference transcriptions used here are in a Latin field orthography with phonemic diacritics; the model evaluated here emits Central Kurdish Arabic script. The two sides of the comparison are in different writing systems before recognition accuracy enters into it at all.
One property of the Arabic-based orthography is central to the results reported below. The script does not write the short vowel conventionally transcribed i (the bizroke), and the grapheme و is ambiguous between the vowel /u/ and the glide /w/. Neither is a defect of any particular system: both are properties of the writing system that any Arabic-script-to-Latin conversion must resolve by inference. Section 4.6 shows that these two phenomena account for almost all of the error this scoring pipeline produces on input that is already correct.
An unmodified comparison would therefore measure a writing-system difference together with recognition error. I score the same hypotheses under progressively normalized representations, holding the reference fixed, so that the change in measured agreement produced by each transformation is visible separately from the error that remains. This shows how much of the measured rate moves under orthographic processing. It does not partition the remaining error into orthographic and acoustic components, and I do not read it as doing so (§6).
3. Data and Methods
3.1. Data
The material is Phase 1 of a Garrusi Kurdish field corpus: elicited questionnaire speech, recorded as MP3 files, with time-aligned reference transcriptions in a Latin orthography using phonemic diacritics.2 The five speakers evaluated here (CZ, FI, MR, MY, SK) are those the processing run completed, of the thirty in Phase 1. The selection was neither random nor designed: the run processed speakers in a fixed order (CZ, MY, SK, FI, MR), which is not the order in which they appear in the corpus directory, and did not cover the remainder. No log of the run was kept, so the point at which it stopped is inferred from the output rather than recorded. I report the five without treating them as a sample of Garrusi speakers, and give per-speaker results without comparing them. No speaker metadata were available for this phase. No Phase 2 or Phase 3 material is used, and the evaluated set contains no free or conversational speech.
The processing script excluded segments shorter than 0.3 s and segments whose reference field was empty or a placeholder. Of 1,765 segments in the reference alignment files, 43 were excluded, all of them for falling below the 0.3 s threshold; no segment was excluded for an empty reference. After these exclusions, and after removing 247 duplicate rows produced by a resumed processing run, deduplicated on the speaker and segment identifier (1,969 rows reduced to 1,722), the evaluation set comprises 1,722 segments, 9,763 reference word tokens, and 7,073.0 s (117.9 minutes) of scored audio. That is a mean of 5.67 reference words per segment, a median of 5, and approximately 1.38 reference words per second. There is no train/development/test split, because no training or fine-tuning was performed; the entire set is an evaluation set.
Two segment sets appear in this paper, and every reported figure belongs to exactly one of them. The staged normalization analysis of §4.1 to §4.4 uses all 1,722 segments, because it involves one system. The two-system comparison of §4.5 and the sensitivity analysis of §4.7 use the 1,703 segments on which both systems returned output, because a paired comparison requires a common set. Table 1 states the two, and §4.5 reports a check on whether the difference between them affects the comparison, together with what that check does and does not establish.
Per-speaker reference token counts document the weighting implicit in a corpus-level error rate: because the estimator pools edits and reference tokens across the corpus rather than averaging per-segment or per-speaker rates, each speaker contributes in proportion to their reference token count. On the full set the counts are CZ 1,261, FI 2,030, MR 2,176, MY 1,948 and SK 2,348 (Table 4); on the matched set they are CZ 1,261, FI 2,015, MR 2,174, MY 1,946 and SK 2,346, summing to 9,742. CZ therefore carries 12.9% of the matched corpus and SK 24.1%, and a corpus-level rate weights them accordingly.
3.2. Recognition
I used facebook/mms-1b-all [11] with the Central Kurdish (ckb) adapter loaded and the tokenizer target language set accordingly. MMS uses small per-language adapter modules inserted into the pretrained backbone [11]; all other components were left as released. Audio was loaded at 16 kHz, mono, and clipped to the reference time alignments. Decoding was greedy CTC decoding, frame-wise argmax over the model’s output distribution, with no beam search and no external language model. Pratap et al. [11] report that the multilingual benchmark results for their multi-domain model were obtained using n-gram language models trained on Common Crawl at inference (their Table 5). My figures are therefore not comparable to those results and should not be read as the best obtainable from this model.
I performed no training, fine-tuning, or adaptation on Garrusi data of any kind; the model and adapter were used as released. I cannot audit the contents of the MMS-1B-all pretraining corpus, so “zero-shot” here describes my procedure rather than a verified absence of related material upstream.
The mismatch between the model’s training material and my data is not only one of variety. MMS-1B-all is fine-tuned on a mixture of MMS-lab, FLEURS, CommonVoice, VoxPopuli and MLS, and MMS-lab is derived from New Testament recordings [11]. I do not know which of these corpora contribute to the Central Kurdish component. My data are field recordings of elicited questionnaire speech distributed as MP3, which I take to differ from these sources in speaking style and recording channel as well as in linguistic variety. The reported rates should therefore not be attributed to variety mismatch alone, and I do not attempt to separate the three.
A second system was subsequently run on the same audio: aranemini/southern-kurdish-asr, the w2v-BERT CTC model fine-tuned on Southern Kurdish released by Mohammadamini and Tahon [10], used as released and without adaptation on Garrusi data. It is included because Garrusi is placed among the Southern Kurdish subgroups by at least one classification (§2) and is represented by one speaker in that study’s benchmark, and because the checkpoint is publicly available. Like MMS-1B-all it emits Arabic script, and its hypotheses were passed through the same transliteration and folding path and scored against the same folded reference text. The normalization function distributed with the model’s inference script was not applied, so the hypothesis received no treatment beyond the path described in §3.3. The character ۊ (U+06CA) represents a vowel specific to Southern Kurdish and is not part of the Central Kurdish inventory the transliterator was written for; it occurs in the raw hypothesis of 170 of the 1,703 scored segments. The folding table of §3.3 maps it to u; a table that did not cover it would convert it to whitespace and split the token containing it, which is why the table is stated in full there. Inference failed on 19 of the 1,722 segments; §4.5 reports what was scored and what the difference makes.
3.3. Text Transformations
The two sides of the comparison start in different representations. The reference is in the Latin field orthography of the source corpus, which uses diacritics to mark phonemic distinctions. The hypothesis is in the Central Kurdish Arabic script emitted by the model. Two transformations bring them toward a common representation, and they are applied to different sides.
Transliteration converts Arabic script to Latin, and is applied to the hypothesis only, using KLPT 0.1.7 [1]. The transliterator was instantiated as Transliterate(“Sorani”, “Arabic”, target_script=“Latin”) and applied per segment via its transliterate method; other constructor arguments were left at their defaults. Its purpose is to remove the writing-system difference, so that the two sides are at least in the same alphabet. The reference requires no transliteration, as it is already Latin. KLPT’s Latin target is a romanization convention and is not the same convention as the field orthography of the reference. Transliteration therefore reduces the representational gap without closing it.
Folding maps a Latin string onto a reduced inventory. The input is NFC-normalized and lowercased. Each character is then replaced according to the mapping below, every remaining character outside [a–z and space] is converted to whitespace, and whitespace is collapsed.
Replaced: ā→a ē→e ī→i ō→o ū→u â→a ê→e î→i û→u ü→u ô→o é→e è→e ë→e ł→l ř→r ḷ→l ṟ→r ħ→h ḧ→h ẍ→x ç→c ş→s š→s ž→j ە→e ۊ→u
Deleted: ʕ ʔ ʼ ʾ ʿ ʻ ‘ ’ ‘ ` ئ U+FFFD and the digit 1
All other characters outside [a–z and space]: → whitespace
Two properties of this table matter for what follows. It covers every character that the transliterator emits on this material, so no hypothesis token is split by an unmapped character: the Southern Kurdish vowel ۊ and the Central Kurdish ە, together with the Latin diacritics ë and ẍ that KLPT emits, are all mapped rather than discarded. A table that omitted any of them would silently fragment the tokens containing it, and the resulting error would be attributed to the recognizer. The exception is U+FFFD, which the transliterator emits where it fails to convert a character at all; that material is irrecoverable and is deleted, and §6.4 states what this costs. The mapping has 40 entries in total, 27 replacements and 13 deletions, and is reproduced here from the implementation used for scoring.
The mapping is deliberately lossy: it collapses diacritic distinctions and deletes pharyngeal and glottal marks, so pairs that differ phonemically between reference and hypothesis may be counted as matches after folding. Any character not listed and not already in [a–z and space] becomes whitespace, which splits the token containing it; I verified the character inventory of both the reference and the transliterated hypothesis against this table. On the reference side the only characters outside the table are punctuation, so folding leaves the reference at 9,763 tokens and 53,017 characters.
Folding is applied to the reference once and the result is then held fixed. Across the three conditions of the main analysis only the hypothesis representation changes: RAW leaves it in Arabic script, TRANSLIT applies transliteration, and FOLDED applies transliteration and then folding. The pipeline is therefore not identical normalization on both sides, since the hypothesis undergoes a transliteration step that the reference does not. Describing the two sides as receiving identical treatment would misdescribe the design.
3.4. The Common-Reference Scoring Design
The reference is folded once and thereafter held constant. All three conditions of the main analysis score the same 1,722 segments against the same 9,763-token folded reference; only the hypothesis representation changes:
RAW: the Arabic-script hypothesis, unmodified.
TRANSLIT: the hypothesis after transliteration to Latin.
FOLDED: the hypothesis after transliteration and folding.
Scoring used jiwer 3.0.3 for word and character error rate, with the library’s default transformations. WER is the word-level edit distance (substitutions plus deletions plus insertions) divided by the number of reference word tokens. CER is the character-level edit distance divided by the number of reference characters, with inter-word spaces counted as characters. Both are pooled at the corpus level, total edits divided by the total reference count, rather than averaged over segments. This estimator is used everywhere in the paper, including for the round-trip control of §4.6; no reported figure is a mean of per-segment rates.
Because the reference is fixed, S + D + H = 9,763 in every condition. A reader can confirm from this that the reference length is constant across conditions, a necessary condition for the design. Confirming that the reference tokens themselves are identical requires the reference file itself, which is part of the data release described in the Data Availability Statement.
3.5. A Representation That Is Not Part of the Staged Ladder
Section 4.6 shows that most of what this pipeline measures on already-correct input is attributable to the letter i, which the source script does not write. That raises an obvious question: what would the measured rates look like if short /i/ were simply not treated as contrastive at scoring time? I answer it in §4.7 under a representation I call FOLDED-i, defined as the FOLDED representation with every instance of the character i deleted from both the reference and the hypothesis, after which empty tokens are dropped and whitespace is collapsed.
FOLDED-i is not a fourth stage of the design set out in §3.4, and it should not be read as one. The three conditions of the main analysis share a single fixed reference, which is what allows their differences to be attributed to the hypothesis-side transformation. FOLDED-i modifies the reference as well. Two consequences follow, and both are measured in §4.7 rather than assumed. The denominator changes: the reference falls from 9,742 to 9,684 tokens on the matched set, a reduction of 58 tokens or 0.60%, because some reference tokens consist of the letter i alone and disappear entirely. More importantly, what counts as a word changes: forms distinguished only by i, such as bani and ban, become the same string, so the effective vocabulary shrinks and some distinct lexical items are no longer distinguishable. Section 4.8 measures how often that happens.
For these reasons I present FOLDED-i as a sensitivity analysis on the comparison of §4.5 and not as an improved scoring convention, and I do not report a FOLDED-i rate alongside the RAW, TRANSLIT and FOLDED rates in the same table. A WER computed under FOLDED-i and a WER computed under FOLDED are not two measurements of the same quantity.
3.6. Statistical Procedure
Where an interval or a p-value is reported for a difference between systems, it comes from a paired bootstrap with 10,000 replicates at seed 42, computed at the 95% level and never clamped to exclude values outside the observed range. Point estimates are always the value observed on the real data, never the bootstrap mean; bootstrap means are retained only as convergence diagnostics. Replicate arrays will be released with the segment-level results.
Two resampling units are reported. The segment-level bootstrap resamples the 1,703 matched segments and is the primary interval. The speaker-cluster bootstrap resamples the five speakers with replacement, carrying all of a speaker’s segments together, and is reported as a robustness check on between-speaker variability. It should be read with a specific caveat. Because the estimator is a corpus-level rate, a cluster replicate depends only on how many times each speaker was drawn and not on the order of the draws, so the statistic takes at most 126 distinct values, this being the number of multisets of size five drawn from five speakers. The bootstrap distribution is therefore a lattice with at most 126 atoms, and its tails are driven by replicates dominated by a single speaker. Five clusters is few. The cluster interval is indicative of between-speaker variability; it is not a precise interval and I do not treat it as a population estimate.
Where an interval’s bound falls close to zero, I report the proportion of bootstrap replicates on each side of zero alongside the interval, because a bound within a small fraction of a percentage point of zero is not meaningfully distinguishable from zero and its printed sign depends on rounding.
4. Results
4.1. Main Result
Figure 1 shows the composition of the three conditions against the fixed reference; Table 2 gives the exact counts.
The RAW condition returns zero word-level hits across all 1,722 segments. This follows from the comparison itself: a Latin-script reference and an Arabic-script hypothesis have disjoint token inventories, so no reference word can match. The resulting 111.70% is a scoring baseline for the ablation, not a recognition rate for the system.
Table 2.
Common-reference staged normalization. All three conditions score the same 1,722 segments against the same folded reference of 9,763 word tokens and 53,017 characters (inter-word spaces included, as counted by the scoring library); only the hypothesis representation differs. S = substitutions, D = deletions, I = insertions, H = hits. Rates are computed from the counts shown and rounded to two decimals, so every row can be recomputed from the table. S + D + H = 9,763 at the word level in every row, confirming a constant reference length; identity of reference content across conditions is established by the released reference file, not by the table. Folding is applied only in the FOLDED condition, so the RAW and TRANSLIT rows are independent of the table in §3.3.
Table 2.
Common-reference staged normalization. All three conditions score the same 1,722 segments against the same folded reference of 9,763 word tokens and 53,017 characters (inter-word spaces included, as counted by the scoring library); only the hypothesis representation differs. S = substitutions, D = deletions, I = insertions, H = hits. Rates are computed from the counts shown and rounded to two decimals, so every row can be recomputed from the table. S + D + H = 9,763 at the word level in every row, confirming a constant reference length; identity of reference content across conditions is established by the released reference file, not by the table. Folding is applied only in the FOLDED condition, so the RAW and TRANSLIT rows are independent of the table in §3.3.
| Condition | Ref. tokens | WER | CER | S | D | I | H |
|---|---|---|---|---|---|---|---|
| RAW (Arabic hypothesis) | 9,763 | 111.70% | 100.92% | 8,054 | 1,709 | 1,142 | 0 |
| TRANSLIT (transliterated) | 9,763 | 102.36% | 57.89% | 7,065 | 1,749 | 1,179 | 949 |
| FOLDED (transliterated + folded) | 9,763 | 97.45% | 51.13% | 6,457 | 1,891 | 1,166 | 1,415 |
Table 3.
Change in the measured rates between conditions.
| Transition | ΔWER | ΔCER |
|---|---|---|
| RAW → TRANSLIT | −9.34 pts | −43.03 pts |
| TRANSLIT → FOLDED | −4.91 pts | −6.76 pts |
| RAW → FOLDED | −14.25 pts | −49.79 pts |
The final condition gives the main result: 97.45% WER and 51.13% CER, from 9,514 word-level edits against 9,763 reference tokens. Of those reference tokens, 1,415 (14.49%) are aligned as exact matches. This figure belongs to the full 1,722-segment set. The same system, in the same condition, under the same mapping, scores 97.43% on the 1,703-segment matched set used for the two-system comparison of §4.5. The two are not competing estimates of one quantity: they differ only in that the second omits the 19 segments on which the second system returned no output, and the resulting 0.016-point difference is a property of the segment set, not of the scoring. Every WER in this paper is stated with the segment set it belongs to, and Table 1 lists the two.
4.2. Error Composition
All figures in this subsection are for the full 1,722-segment set; the corresponding figures on the 1,703-segment matched set are given in §4.5 and differ slightly. In the FOLDED condition the 9,514 edits comprise 6,457 substitutions (67.9% of edits), 1,891 deletions (19.9%), and 1,166 insertions (12.3%). The hypothesis contains 9,038 word tokens, 92.6% of the reference length. The edits are substitution-dominated, and the system produces slightly less text than the reference instead of over-generating. Empty hypotheses were produced for 75 segments (4.4%). In the RAW condition, WER exceeds 100% because no reference token matches at all and insertions are added to a full complement of substitutions and deletions, not because insertions predominate.
4.3. Speaker-Level Results
Word error rate ranges from 90.35% to 101.41% and character error rate from 43.74% to 57.51%, so the pooled figure does not rest on a single outlying speaker. I do not interpret the differences between speakers. With five speakers, no metadata for this phase, and no analysis of which elicitation items each speaker contributed, the variation cannot be attributed to speaker characteristics, elicitation content, or recording conditions.
Table 4.
Speaker-level results for MMS-1B-all in the FOLDED condition, over the full 1,722-segment set. Figures computed per speaker from the segment-level data.
Table 4.
Speaker-level results for MMS-1B-all in the FOLDED condition, over the full 1,722-segment set. Figures computed per speaker from the segment-level data.
| Speaker | Segments | Ref. tokens | WER | CER |
|---|---|---|---|---|
| CZ | 230 | 1,261 | 101.03% | 53.89% |
| FI | 398 | 2,030 | 99.31% | 57.51% |
| MR | 307 | 2,176 | 90.35% | 43.74% |
| MY | 363 | 1,948 | 96.36% | 47.15% |
| SK | 424 | 2,348 | 101.41% | 54.19% |
| All | 1,722 | 9,763 | 97.45% | 51.13% |
4.4. Segment Length
This is a secondary observation and is reported descriptively; it does not bear on the main argument. Per-segment WER correlates negatively with the number of reference tokens in the segment (Spearman rho = −0.389, p = 1.69 × 10−62; Pearson r = −0.250, p = 1.07 × 10−25, computed on the matched set). Its association with segment duration is weak, and the two coefficients do not agree in sign, so I do not treat duration as showing a consistent association in this set. Per-segment WER tends to be higher for segments with fewer reference tokens (Figure 2).
Three qualifications apply. First, because per-segment WER is defined with the reference token count as its denominator, part of this association is a property of the metric and not of recognition difficulty. In very short segments the rate takes few, widely spaced values and exceeds 1.0 readily, whereas longer segments yield values concentrated nearer the corpus rate. I therefore report the relationship descriptively and do not treat it as an estimate of a length effect on recognition. Second, the token-count and duration coefficients are not independent measurements: reference token count and duration are themselves strongly correlated in this set. The contrast between them indicates which of two correlated length measures tracks per-segment WER more closely, not that duration is unrelated to recognition difficulty. Third, with more than 1,700 segments the p-values reflect sample size rather than the magnitude of the association.
The set is composed of short segments: median duration 3.0 s, a median of 5 reference tokens, and 476 of the 1,703 matched segments (28.0%) containing three or fewer reference tokens.
4.5. A Southern Kurdish Fine-Tuned System
The Southern Kurdish fine-tuned system described in §3.2 was run on the same audio and scored through the same path: the same transliteration, the same folding table, and the same folded reference, over the 1,703 segments on which both systems returned output. Holding all of these constant is what makes the comparison interpretable, given that §4.6 shows the pipeline itself to be a substantial contributor to any absolute rate. What varies between the rows of Table 5 is the recognizer and nothing else.
Inference failed on 19 of the 1,722 segments (FI 13, MR 2, MY 2, SK 2). These are inference failures rather than exclusions under the criteria of §3.1, and no empty hypothesis was substituted for them in the matched set. Because both systems are scored over the same 1,703 segments against the same 9,742-token reference, the comparison in Table 5 is paired.
Table 5.
The two systems in the FOLDED condition, over the 1,703 segments common to both, against the same folded reference of 9,742 word tokens and 52,942 characters. Both rows are computed under the folding table of §3.3, with the corpus-level estimator of §3.4; CER denominators are reference characters including inter-word spaces. S + D + H = 9,742 in both rows. The MMS row is the FOLDED row of Table 2 recomputed over 19 fewer segments: 97.43% here against 97.45% there, a difference due entirely to the segment set and not to any change in scoring.
Table 5.
The two systems in the FOLDED condition, over the 1,703 segments common to both, against the same folded reference of 9,742 word tokens and 52,942 characters. Both rows are computed under the folding table of §3.3, with the corpus-level estimator of §3.4; CER denominators are reference characters including inter-word spaces. S + D + H = 9,742 in both rows. The MMS row is the FOLDED row of Table 2 recomputed over 19 fewer segments: 97.43% here against 97.45% there, a difference due entirely to the segment set and not to any change in scoring.
| System | Segments | Ref. tokens | WER | CER | S | D | I | H |
|---|---|---|---|---|---|---|---|---|
| MMS-1B-all, ckb adapter | 1,703 | 9,742 | 97.43% | 51.06% | 6,447 | 1,880 | 1,165 | 1,415 |
| aranemini/southern-kurdish-asr | 1,703 | 9,742 | 108.12% | 55.78% | 6,460 | 686 | 3,387 | 2,596 |
The Southern Kurdish system trails the Central Kurdish adapter by 10.69 percentage points of WER (95% CI [8.65, 12.80] by segment-level bootstrap; [7.87, 14.75] by speaker-cluster bootstrap; no replicate of either bootstrap fell at or below zero) and by 4.72 points of CER. It does so on all five speakers under both metrics. This is the central factual claim of this section.
It is worth stating why this cannot be dismissed as an artefact of the scoring path. The transliterator used here was written for Central Kurdish, and the Southern Kurdish system emits characters outside that inventory, notably ۊ; a pipeline that discarded those characters would fragment this system’s tokens and inflate its rate for reasons having nothing to do with recognition. The folding table of §3.3 maps them instead, so that mechanism is not operating on the figures in Table 5. A residual asymmetry may remain, since the transliterator’s inference of unwritten material was tuned on a different variety, and §6.5 states why its size is not established here. What can be said is that the difference does not arise from characters being discarded before alignment.
The error composition differs more than the pooled rate does. Of the Southern Kurdish system’s 10,533 edits, 6,460 are substitutions (61.3%), 686 deletions (6.5%), and 3,387 insertions (32.2%); its hypothesis contains 12,443 word tokens, 127.7% of the reference length, and it returned an empty hypothesis for none of the 1,703 segments. The corresponding figures for MMS-1B-all on the same segments are 9,492 edits comprising 67.9% substitutions, 19.8% deletions and 12.3% insertions, a hypothesis 92.7% of reference length, and 66 empty hypotheses. The Southern Kurdish system matches a substantially larger share of reference tokens, 2,596 against 1,415, while producing considerably more text than the reference; the Central Kurdish adapter matches fewer while producing slightly less. Word error rate above 100% arises here for a different reason than in the RAW condition of §4.1: there it followed from the absence of any word-level agreement, whereas here insertions are added to a substantial complement of matches.
The 19 segments. Because they are segments on which one system produced no output and the other did, their exclusion is a possible source of bias. Two things can be said about it. Their content is short: they hold 21 reference tokens between them, 17 of the 19 consisting of a single token, and the tokens are predominantly short responses and discourse particles (aha seven times, xob three times, eri twice, beli, keft, and five others). They are the shortest segments to survive the 0.3 s exclusion of §3.1. I do not claim to know why inference failed on them. A minimum input length in the second system’s feature extractor is consistent with what is observed, but I did not establish it, and I report the pattern rather than the cause.
Within each representation the two treatments differ by 0.03 percentage points under FOLDED and 0.02 under FOLDED-i. What this establishes is bounded and worth stating precisely: under the only two treatments available to me, excluding these segments or scoring them as empty output, the estimated difference between the systems is insensitive to the choice. It is also worth noting that the direction is the opposite of the one that motivated the concern: on those 21 tokens the second system scores exactly 100% WER, all deletions, while the first scores 104.76% with no hits at all, so excluding them very slightly favors the first system rather than the second.
Three things remain unknown, and the sensitivity check does not address them. I do not know why inference failed on these segments, so I cannot rule out that the failure mechanism also degrades the second system’s output on segments where it did not fail outright, which no treatment of the 19 would reveal. I do not know whether an empty hypothesis is the right counterfactual for a failed inference; it is the protocol-consistent one, but a system that had returned output might have scored better or worse than empty. And 21 tokens is 0.2% of the reference, so this check has little power to detect a small bias even in principle. I report the matched set as primary because it is the paired comparison, and Table 6 as a bounded check rather than as a demonstration that the exclusion is harmless.
Two asymmetries remain, and neither is removed by the rescoring. The transliterator was written for Central Kurdish and is applied here to a system whose output is Southern Kurdish; and the two systems differ in architecture, training data and decoding behavior as well as in target variety. The comparison establishes that on this evaluation set, under this scoring pipeline, the Southern Kurdish fine-tuned system does not outperform the Central Kurdish adapter. It does not establish why.
4.6. A Round-Trip Control: What the Pipeline Returns for a Correct Hypothesis
Every rate reported so far is measured through a pipeline that converts between two writing systems, and none of them establishes what that pipeline returns for a hypothesis that is already correct. Without that quantity the residual error cannot be interpreted, because there is no way to tell how much of it the measuring apparatus contributes. This section supplies it. A correct Arabic-script rendering of 148 reference segments was passed through the identical transliteration and folding path and scored against the folded reference, using the same code and the same estimator as everywhere else in the paper. Any error it returns is produced by the pipeline, because there is no recognizer in the loop.
The control set is provisional in one specific respect. It comprises 148 segments and 880 reference tokens, and the material retained for it carries the reference and round-trip hypothesis text but not speaker or segment identifiers. I therefore cannot state which speakers are represented, whether the selection was random or purposive, or whether these segments are a subset of the 1,703 used elsewhere, and the estimate below should be read as an estimate of the floor for material of this kind rather than as a calibrated figure for the evaluation set. The qualification bears unevenly on the two things this section reports: the magnitude of the floor depends on which segments were sampled, whereas its composition is a property of the string pairs themselves and would be unchanged by a different sample of the same orthography. Section 6.6 states what would be required to remove the qualification.
Table 7.
Round-trip control. A correct Arabic-script rendering of 148 reference segments, passed through the same transliteration and folding path and scored against the folded reference. Intervals are segment-level bootstrap, 10,000 replicates, seed 42.
Table 7.
Round-trip control. A correct Arabic-script rendering of 148 reference segments, passed through the same transliteration and folding path and scored against the folded reference. Intervals are segment-level bootstrap, 10,000 replicates, seed 42.
| Condition | Ref. tokens | WER | 95% CI | S | D | I |
|---|---|---|---|---|---|---|
| Standard (as in FOLDED) | 880 | 47.05% | [43.58, 50.57] | 410 | 4 | 0 |
| With i deleted from both sides | 876 | 5.59% | [4.00, 7.34] | 49 | 0 | 0 |
The pipeline returns 47.05% WER on input that is correct by construction. It is, in my view, the most consequential number in this paper, because it constrains what any rate reported in §4.1 or §4.5 can be read as measuring, although it does not bound it arithmetically (§5). A system scoring 97.45% is not making twice as many errors as a system scoring 47.05%; it is being measured on a scale whose zero point is not zero.
The floor is not diffuse noise, and its composition can be stated precisely. The 414 edits consist of 410 substitutions and 4 deletions, with no insertions at all. Of the 410 substitutions:
359 (87.6%) are pairs in which the reference and hypothesis forms differ only in the presence or absence of the character i. This is the operational definition used throughout: the two forms are identical after deleting every i from each. It is not the same as, and is weaker than, requiring that the hypothesis be the reference with all its i characters removed; under that stricter reading the count is 294 (71.7%). The difference between the two is 65 substitutions in which some but not all instances of i are absent, such as diriyaye against diryaye and kirili against kirl. I report the looser figure and state the definition rather than describing these as pure i-deletion pairs, which would be ambiguous between the two readings. The most frequent single pair is ki against k, 58 times.
45 of the remaining 51 (88.2% of the residual) are alternations between u and w, arising on the grapheme و: for example zaruege against zarwege ten times, ruweyn against rueyn, duweye against dueye.
6 substitutions fall outside both classes. Three of these are alignment artefacts around a single-character token i, paired with the four deletions. The remaining three are muy against mui, twice, and riyi against ri.
The two identified phenomena therefore account for 404 of 410 substitutions, or 98.5%. Deleting i from both sides reduces the floor from 47.05% to 5.59%, an 88.11% reduction, which is the same fact expressed as a rate.
Both phenomena have the same structural cause, and it is a property of the writing system rather than of any component of the pipeline. The Arabic-based orthography does not write the short vowel i, and و is ambiguous between /u/ and /w/ (§2). A converter from that script to a Latin representation must therefore infer both, and KLPT’s own documentation reports its detection of the unwritten vowel at 39% accuracy [1, p. 77]. The two remaining unexplained pairs are consistent with a third instance of the same pattern, ی being ambiguous between /i/ and /y/, but three tokens is too few to assert it.
One thing this control does not show. It establishes that these string pairs differ systematically when a correct transcription is passed through the pipeline. It does not establish that the recognizers make errors of this kind, and it does not license reading any particular substitution in §4.1 or §4.5 as an i-deletion or a u/w alternation rather than as a recognition error. Distinguishing what the script fails to encode, what the transliterator infers, what the reference orthography writes, and what a system actually recognized would require an aligned analysis of the system outputs themselves, which I have not done. What the control supports is a statement about the measuring instrument, not about the systems being measured.
4.7. Sensitivity of the System Comparison to the Treatment of Short /i/
Section 4.6 shows that the letter i accounts for most of what the pipeline measures on correct input. Section 4.5 reports a difference of 10.69 points between two systems measured through that pipeline. It is therefore reasonable to ask how much of the measured difference survives if short /i/ is treated as non-contrastive. This section answers that question under the FOLDED-i representation defined in §3.5. It is a sensitivity analysis on the comparison of §4.5, not a further condition of the staged design, for the reasons given there.
Table 8.
The system comparison under the two scoring representations, matched set (1,703 segments). FOLDED-i deletes every i from both reference and hypothesis, which is why its reference token count differs. The two columns are therefore not two measurements of the same quantity and the WERs in them are not directly comparable across columns (§3.5); what is compared across columns is the difference between systems, each computed within its own representation. Intervals are 95% paired bootstrap, 10,000 replicates, seed 42, unclamped. Differences are Aranemini minus MMS.
Table 8.
The system comparison under the two scoring representations, matched set (1,703 segments). FOLDED-i deletes every i from both reference and hypothesis, which is why its reference token count differs. The two columns are therefore not two measurements of the same quantity and the WERs in them are not directly comparable across columns (§3.5); what is compared across columns is the difference between systems, each computed within its own representation. Intervals are 95% paired bootstrap, 10,000 replicates, seed 42, unclamped. Differences are Aranemini minus MMS.
| FOLDED | FOLDED-i | |
|---|---|---|
| Reference tokens | 9,742 | 9,684 |
| MMS-1B-all WER | 97.43% | 94.16% |
| edits / hits | 9,492 / 1,415 | 9,118 / 1,759 |
| aranemini WER | 108.12% | 96.28% |
| edits / hits | 10,533 / 2,596 | 9,324 / 3,840 |
| Difference | 10.686 pp | 2.127 pp |
| Segment-level 95% CI | [8.646, 12.803] | [−0.0003, 4.306] |
| Replicates ≤ 0 (of 10,000) | 0 | 251 |
| Segment-level p (two-sided) | < 0.0001 | 0.0502 |
| Speaker-cluster 95% CI | [7.87, 14.75] | [−0.22, 5.18] |
| Speaker-cluster p (two-sided) | < 0.0001 | 0.0816 |
Figure 3.
The difference between the two systems under the two scoring representations, matched set (1,703 segments), with both bootstrap intervals. The point estimate falls by 80.1%. The FOLDED-i segment-level interval’s lower bound is −0.0003 percentage points, which is not visually distinguishable from zero at this scale and should not be read as a bound away from it.
Figure 3.
The difference between the two systems under the two scoring representations, matched set (1,703 segments), with both bootstrap intervals. The point estimate falls by 80.1%. The FOLDED-i segment-level interval’s lower bound is −0.0003 percentage points, which is not visually distinguishable from zero at this scale and should not be read as a bound away from it.

The estimated difference falls from 10.686 to 2.127 percentage points, a reduction of 80.1%, and under both resampling units the interval includes zero. The reduction in the point estimate is the result. I deliberately do not frame this as a significant difference becoming non-significant. The segment-level two-sided p-value is 0.0502, which sits on the conventional threshold closely enough that it could fall on either side of it under a different seed, and the segment-level interval’s lower bound of −0.0003 percentage points is not distinguishable from zero: printed to two decimals it would appear as 0.00 and could be read as excluding zero. The more informative statement is the direct one, that 251 of 10,000 replicates fell at or below zero under segment-level resampling, and that the speaker-cluster p-value is 0.0816. An 80% reduction in the point estimate is a claim about magnitude that does not depend on where a threshold is drawn; a change in the verdict of a hypothesis test at p = 0.0502 is not.
What this does and does not show. It shows that the measured difference between these two systems is reduced by approximately 80% when short /i/ is treated as non-contrastive in reference and hypothesis alike. It does not show that 80% of the recognition error disappears, and the two statements are not equivalent: FOLDED-i is a different measurement, over a smaller reference and a coarser vocabulary, not a more accurate measurement of the same quantity. Nor does it show that the systems are equivalent under FOLDED-i. An interval including zero is not evidence of no difference, the point estimate remains 2.13 points in the same direction, and with five speakers the cluster analysis is a check on between-speaker variability rather than a population estimate. The defensible conclusion is that under FOLDED-i the data do not provide strong evidence of a remaining difference, and that the magnitude of any remaining difference is small relative to the one measured under FOLDED.
Two further observations bear on interpretation. First, the reduction is not symmetric between systems. Deleting i removes 374 edits from the first system, 3.9% of its total, and 1,209 from the second, 11.5% of its total. The gap does not close because both systems improve equally; it closes because the second system’s excess errors under FOLDED are disproportionately i-related. Second, the direction of the difference is not uniform across speakers under FOLDED-i. Under FOLDED all five speakers favor the first system, by between 7.03 and 20.22 points. Under FOLDED-i two of the five reverse sign (MR by −1.06 points, MY by −0.15). With five speakers and no per-speaker intervals this is an illustration of where the cluster interval’s width comes from and not evidence of a speaker effect.
4.8. What Deleting i Costs: Lexical Ambiguity in the Reference
Section 4.7 might be read as showing that i should simply be deleted before scoring, since doing so removes most of the pipeline floor and most of the measured difference between systems. This section is the reason I do not draw that conclusion. Deleting i is not a neutral normalization: it merges forms that the reference orthography distinguishes, and the merged forms are not rare or marginal.
Measured on the 9,742-token matched reference, deleting i reduces 1,465 distinct types to 1,384, eliminating 81 types (5.53%). Those 81 eliminated types correspond to 81 collision classes, each containing exactly two original types, so 162 types (11.06% of the vocabulary) become non-distinct. The type-level figure understates the effect, because the affected types are frequent ones: the 162 ambiguous types account for 1,629 of the 9,742 reference tokens, or 16.72%. Under an oracle that resolves every collision by guessing the more frequent member of its class, 425 tokens (4.36% of the reference) would still be assigned the wrong form.
The collisions are not arbitrary. They are overwhelmingly the contrast between a bare noun and the same noun carrying the ezafe, which in this orthography is written with the same i that the source script does not encode: bani against ban (133 and 62 tokens), naw against nawi (93 and 5), des against desi (71 and 53), bed against bedi (73 and 1), seri against ser (48 and 6), rengi against reng (28 and 23). A scoring representation that deletes i is therefore one in which a morphosyntactic distinction of the language is not measurable. That is a defensible choice for a sensitivity analysis, whose purpose is to ask what the comparison looks like without it. It is not a defensible default, because a system that never produced the ezafe and a system that always produced it would score identically on 16.72% of the corpus.
The same statistic computed on the 148-segment control subset is 8.18% rather than 16.72%. The two are not in conflict: collision rates rise with vocabulary size, and the control subset has a much smaller lexicon. I report both so that the difference is not mistaken for an inconsistency.
A note on the denominator. I report the token-weighted figure as the primary statistic because the type-level figure answers a different question. Whether 11.06% of the vocabulary is affected tells a reader about the lexicon; whether 16.72% of running tokens are affected tells a reader what share of the scored material is affected, which is the quantity that bears on a corpus-level rate.
5. Discussion
The measured rate
The 97.45% WER is a transfer result: a model adapted to Central Kurdish, applied without any Garrusi data, on elicited speech recorded in the field and distributed as MP3. It characterizes one off-the-shelf configuration on this evaluation set, not the range of systems that might be applied to this variety, and it does not indicate what Garrusi ASR could achieve with in-variety data. It should now also be read against §4.6: a pipeline that returns 47.05% WER on a correct hypothesis cannot support a reading of 97.45% as a measure of recognition failure. The absolute rate is close to uninterpretable on its own, and I would discourage citing it as a Garrusi recognition rate. What the staged design supports is the comparison between conditions, and what the control adds is a floor against which any of them must be read.
The closest published work in experimental shape is the Southern Kurdish benchmarking of Mohammadamini and Tahon [10], which compares models tuned on Southern Kurdish against a Central Kurdish-tuned baseline. My configuration is likewise a Central Kurdish adapter evaluated on a variety outside the Central Kurdish standard, but I do not restate their figures here and do not attempt a numerical comparison with them, as the corpora, genres, and recording conditions differ. What is comparable is their released system, not their reported rate. Running their Southern Kurdish fine-tuned model on this evaluation set (§4.5) holds the audio, the segmentation and the transcriptions constant and varies only the recognizer. Their released checkpoint supports that comparison where their published figures do not.
There is a further reason for not attempting a numerical comparison, and it bears directly on the design adopted here. Mohammadamini and Tahon [10] observe that the absence of a settled Southern Kurdish orthography allows the same word to be written in several forms, that this spreads small character-level differences across many word tokens, and that it therefore inflates WER relative to CER. Their error typology assigns a separate category to such standardization mismatches, in which both the reference form and the recognized form are attested spellings, and their reported figures are consistent with this account at 24.26 WER against 4.09 CER.
Their analysis is, in effect, independent evidence that a substantial share of a Kurdish word error rate is a property of the representation in which agreement is measured, not of recognition. Section 4.6 is a direct measurement of that share for one pipeline: 47.05 points of WER on input that is correct by construction, of which 98.5% of the substitutions are attributable to two identifiable properties of the writing system. What the published description of their own rates does not state is the representation in which those rates were computed: no normalization procedure, tokenization rule, treatment of punctuation and Unicode format controls, or scoring implementation is specified. This does not affect the validity of the figures within their stated evaluation setting, where the comparisons they draw involve differences of fifty WER points or more, but it does place two specific limits on what can be done with those figures. The first concerns reproduction: a reader who downloads the released corpus, checkpoint and inference script can transcribe the benchmark audio with the same model, but cannot establish that the processing applied to hypothesis and reference before alignment was the processing that produced 24.26 and 4.09, so the published rate cannot be recomputed from the released artefacts and checked. The second concerns the authors’ own analysis. Having identified standardization variation as a source of the WER–CER divergence, the study establishes that this variation contributes to the reported WER without establishing how much of it would remain under a normalization that neutralized the category. That residual quantity is not recoverable from a single reported rate, and it is the quantity that a staged design plus a round-trip control measures directly.
The structure of that benchmark separately determines which of its comparisons its evidence supports. It comprises 100 sentences read by eight speakers, of which 773 recordings passed validation, so its content dimension is one hundred sentences and the same text is scored eight times over. For the contrasts drawn between systems fine-tuned on Southern Kurdish and systems trained on other varieties, which involve differences of fifty word error rate points or more, this is not a material constraint. Two further comparisons are more tightly bounded. The first is the ranking of the two fine-tuned systems, where w2v-BERT reaches 24.26 WER and Whisper-Turbo 34.43 from single runs reported without repeated seeds or intervals; the authors offer two candidate explanations and do not separate them. The second is variety: four of the five vernaculars are represented by a single speaker, so the per-speaker character error rates, which range from 3.02 to 6.49, cannot be attributed to vernacular rather than to speaker, and the authors do not attribute them.
The staged normalization measures
The experiment measures how word- and character-level agreement changes as the same hypotheses are progressively converted from Arabic script to a Latin representation and then folded into the reference’s reduced orthography, with the reference tokenization held fixed throughout. Four observations follow directly. First, the RAW condition has no word-level agreement at all, so cross-script comparison without conversion cannot be interpreted as a recognition measurement. Second, the transliteration step recovers character-level agreement while leaving most word forms unmatched: character error rate falls from 100.92% to 57.89%, and exact word matches rise from none to 949 of the 9,763 reference tokens, with word error rate still above 100% at 102.36%. Third, folding reduces the measured rates by a further 4.91 WER and 6.76 CER points. Fourth, after both transformations the measured WER is still 97.45%, and 85.51% of reference tokens are not matched.
I do not describe any part of this as isolating the causal share of error attributable to script mismatch. The transformations change the representation in which agreement is measured; they do not partition the underlying causes. Folding in particular is lossy: it collapses diacritic distinctions and removes pharyngeals, so some pairs that differ phonemically between reference and hypothesis are counted as matches after folding, and some that differ only in an unmapped character are split into separate tokens. The FOLDED rates should therefore be read as performance under the stated reduced orthography, alongside the TRANSLIT condition rather than in place of it.
The residual, and the control
The residual is large and substitution-dominated: on the full 1,722-segment set, substitutions account for 67.9% of the edits, hypothesis length is 92.6% of reference length, and 14.49% of reference tokens are aligned as exact matches. I do not read this residual as recognition failure alone, and §4.6 shows why. A pipeline floor of 47.05% means that a substantial part of the residual is produced by the measuring apparatus, and that almost all of that part is attributable to two properties of the source writing system: an unwritten short vowel and one ambiguous grapheme.
Two cautions attach to that statement. The floor is measured on 148 segments whose provenance is not yet documented (§6.6), so its precision as an estimate for the full evaluation set is not established. And the floor is not subtractive: 97.45% minus 47.05% is not a corrected recognition rate, because the errors a recognizer makes and the errors the pipeline introduces are not independent and do not compose additively. What the floor supports is a qualitative conclusion, that rates measured through this pipeline are not on a scale whose zero is zero, and a comparative one, that differences measured through it are more interpretable than levels.
A single segment illustrates one error type that is not attributable to the pipeline, a word-boundary difference counted as substitutions. It was selected to show that pattern and is not a typical segment: at 44.44% WER it is well below the corpus rate. In segment MR.141 (10.0 s), the folded reference
ew masin esbabbazige ki we tenabo kisirili ewe kewe
is recognized as
ew masin espebazige u tenaw bo kisirili ewe kewe
The recognizer’s tenaw bo corresponds to the reference’s tenabo: one reference token is aligned against two hypothesis tokens, which the metric counts as a substitution plus an insertion regardless of how close the underlying string is. Neither of the two phenomena identified in §4.6 accounts for this, and word-boundary placement is not something folding or i-deletion addresses.
Representation dependence as the general point
The two systems compared in §4.5 differ by 10.69 percentage points under one scoring representation and by 2.13 under another, where the second representation differs from the first in a single decision: whether a vowel that the source script does not write is treated as contrastive. Both representations are defensible. The first preserves a morphosyntactic distinction that the language makes and the reference orthography records; the second neutralizes a distinction that the input script cannot express and that the transliterator infers at 39% accuracy. Neither is obviously the right choice, and §4.8 shows that the second is not free: it makes 16.72% of reference tokens lexically ambiguous.
The consequence for reporting is narrow and, I think, generalizable beyond Kurdish. Where a reference orthography encodes distinctions that the input writing system does not, the scoring representation is not an implementation detail that can be left to a normalization function and omitted from the write-up. It is a modelling decision that can change the measured difference between two systems by a factor of five. This is why I report the representation in full, release the reference, and present FOLDED-i as an alternative measurement rather than as a correction: a reader who thinks the other choice is right should be able to see what follows from it.
Segment length
That per-segment WER is higher for segments with fewer reference tokens is reported descriptively, with the qualifications in §4.4, and it is a secondary observation rather than a finding. Two explanations are available and neither is tested here. One is arithmetic: in short segments the rate moves in large steps and a single error produces a disproportionate value, and 28.0% of this set contains three or fewer reference tokens. The other is that shorter segments may offer less surrounding context to the recognizer; this is a possible explanation, not an observed result. Practically, the relationship suggests that segmentation policy is a variable worth holding constant in comparative Kurdish ASR evaluation, since two systems evaluated over differently segmented versions of the same audio will not be straightforwardly comparable.
6. Limitations
6.1. Choices in the Scoring Design, and What They Cost
Three choices shape every figure in this paper, and each carries a cost that a reader should be able to weigh.
The folding table is deliberately lossy. It collapses diacritic distinctions and removes pharyngeal and glottal marks, so pairs that differ phonemically between reference and hypothesis can be counted as matches, and the reference tokenization after folding is not the transcriber’s. The alternative is worse for this material: a table that omitted characters the transliterator emits would convert them to whitespace and split the tokens containing them, and the resulting fragmentation would be scored as recognition error. Folding therefore buys comparability between two Latin conventions at the price of phonemic resolution, and the FOLDED rates should be read as performance under the stated reduced orthography rather than as performance simpliciter.
Two segment sets are used, and the choice is forced rather than free. The staged normalization of §4.1 involves one system and uses all 1,722 segments. The two-system comparison of §4.5 requires a set on which both systems returned output, which is 1,703 segments; scoring a system on segments where it produced nothing, or scoring the two over different sets, would each introduce a different distortion. Table 6 reports what the choice is worth, and §4.5 states what that check does not establish. Because the sets differ, a rate quoted from §4.1 and a rate quoted from §4.5 differ slightly for the same system and condition; every figure in this paper is therefore stated with the set it belongs to.
FOLDED-i is a sensitivity analysis and not a scoring convention. It changes the reference as well as the hypothesis, which places it outside the common-reference design (§3.5), and §4.8 measures what it costs: 16.72% of reference tokens become lexically ambiguous, chiefly by merging ezafe-marked and bare forms of frequent nouns. I report the comparison under both representations because neither is obviously correct, and because the difference between them is itself the more general finding of this paper.
6.2. Scope of the Evaluation Set
Five speakers were processed for the analysis reported here, of thirty in Phase 1, because a run that processed speakers in a fixed order did not continue past the fifth; no log was kept, so the interruption is inferred from the output rather than recorded. A later run of the same configuration over all thirty speakers, covering 8,240 of 8,650 segments and 47,160 reference tokens after excluding segments returning an empty hypothesis, gave a pooled word error rate of 102.15% and a character error rate of 57.94%. That run was not scored under the procedure of §3.3 and its rate is therefore not directly comparable with the rates reported here; it also excludes its 410 empty hypotheses whereas the present set retains its 75. It indicates that the five speakers are not unrepresentative of Phase 1 in the aggregate and no more. They are neither a random nor a designed sample, no speaker metadata were available for this phase, and no demographic or speaker-level interpretation is offered. The material is elicited questionnaire speech only, so the results do not extend to spontaneous Garrusi.
The five-speaker limit also bounds the statistical treatment. The speaker-cluster bootstrap is reported as a robustness check and is not a population estimate: with five clusters the statistic takes at most 126 distinct values (§3.6), and its tails are determined by replicates dominated by a single speaker. Generalization beyond these five speakers is not supported by either resampling unit, and the cluster interval should not be read as though it were.
6.3. Coverage of the Folding Table
Folding reduces every character outside the reduced Latin inventory to whitespace, so any character absent from the mapping splits a token rather than normalizing it. This makes the table’s coverage a substantive property of the design rather than an implementation detail, and it is why §3.3 states the table in full and §6.1 explains the trade-off it embodies. On the reference side the effect is bounded: the only characters outside the table are punctuation, and ü, which occurs 179 times across 165 segments and is mapped to u, 170 of those occurrences being word-internal where an unmapped character would have split a token. On the hypothesis side the table covers everything the transliterator emits except U+FFFD, which is discussed below. Folding remains lossy in the other direction, as §6.1 sets out: the reference tokenization after folding is not identical to the transcriber’s.
6.4. Transliteration Is Imperfect
KLPT 0.1.7 does not fully convert the hypothesis to Latin. It emits 502 U+FFFD replacement characters on this material, marking positions where conversion failed outright. The folding table deletes them, because there is nothing to map them to, so that material is lost rather than scored; the alternative, converting them to whitespace, would split the surrounding token and score the loss twice. Separately, the transliterator’s detection of the unwritten vowel i was evaluated by its author at 39% accuracy [1, p. 77], and §4.6 shows what that costs in this pipeline: it is the dominant component of a 47.05% floor. Both belong to the scoring pipeline and not to the recognizer, and both mean the reported rates overstate recognition error by an amount that §4.6 quantifies for one component and leaves unmeasured for the others.
A separate check compared two independently produced transliterations of the same system output over the 1,722 segments, to establish that the figures reported here reflect a single inference run scored twice rather than two inference runs. Of 8,636 aligned tokens, 8,624 (99.86%) were identical, and 12 were not. Two qualifications belong with that figure. The artefact recording the check enumerates only 8 of the 12 differing tokens, and those 8 are word-initial e-epenthesis (iwt against eiwt, u against eu, and five similar) together with one el against xel; I therefore cannot state that all 12 are of these kinds, only the 8 that are listed, and the remaining 4 have not been recovered, so the enumeration should not be treated as complete. Second, 61 of the 1,722 segments (3.54%) were skipped because the two transliterations produced different token counts, and the 8,636 aligned tokens represent roughly 88% of the reference tokens in that set rather than the 96.5% the segment-level rate would suggest, which means the skipped segments are longer than average and the check does not cover them. Neither qualification affects the conclusion that the two are the same inference, but the check is weaker than the headline rate implies.
6.5. The Second System and the Scoring Pipeline
The comparison in §4.5 holds the audio, the segmentation, the reference and the scoring path constant and varies the recognizer, which is what makes it interpretable. One asymmetry nonetheless remains, and it is not removed by holding the path constant. The transliterator was written for Central Kurdish and is applied here to a system whose output is Southern Kurdish, and its task includes inferring material that the source script does not write (§4.6). There is therefore a structural reason to expect it to serve the two systems unequally, even though the folding table covers the characters both emit.
I have not measured the size of that asymmetry, and I do not claim it is negligible. The measurement that would settle it is the one §4.6 performs for the reference orthography, repeated on Southern Kurdish material: a correct Arabic-script rendering of Southern Kurdish text, passed through the same path, would give a floor for that variety against which the second system’s rate could be read as this system’s rate is read against 47.05%. Until that is done, the difference reported in §4.5 should be taken as a difference measured under a specific pipeline, and not as a variety-independent statement about the two recognizers. This is the single most useful extension of the present design.
6.6. Open Items
Three qualifications remain open, and I state them rather than leave them to be discovered. First, the composition of the 148-segment round-trip control is not documented: the material retained for it carries reference and hypothesis text but no speaker or segment identifiers, so its selection criteria, speaker composition, and relation to the 1,703-segment set cannot be stated, nor can the procedure by which the correct Arabic-script rendering was produced. The floor estimate of §4.6 is provisional in that respect, and reproducing the control with identifiers retained would remove the qualification. Second, the reference file must be identified by hash in the release (see the Data Availability Statement). Third, the four unenumerated tokens of the transliteration-divergence check (§6.4) have not been recovered; the eight that are listed are not a complete inventory of the twelve, and the characterization in §6.4 covers only those eight. Until the remaining four are recovered, no claim should be made about the composition of all twelve.
The staged normalization supports no causal claim. The design varies the hypothesis representation and holds the reference fixed. It establishes how much the measured agreement changes under each transformation, but not what causes the remaining errors, and it does not partition the reported rate into orthographic and acoustic components. Section 4.6 constrains the interpretation of the residual without partitioning it, for the reasons given there.
7. Conclusion
I report an ASR measurement for Garrusi Kurdish: MMS-1B-all with the Central Kurdish adapter, used as released and without adaptation on Garrusi data, yields 97.45% WER and 51.13% CER on 1,722 segments of elicited questionnaire speech from five speakers. Scoring under a common-reference design, in which the reference tokenization is fixed across conditions and only the hypothesis representation varies, the RAW-to-FOLDED transformation reduces the measured WER by 14.25 percentage points and the measured CER by 49.79 points, with the folding step beyond transliteration contributing 4.91 and 6.76.
The more consequential results concern what that rate can be read as measuring. A round-trip control, in which a correct Arabic-script rendering of 148 reference segments is passed through the identical pipeline, returns 47.05% WER against a reference it should match exactly. That floor is not diffuse: 404 of its 410 substitutions (98.5%) fall into two classes, pairs differing only in the presence or absence of the unwritten short vowel i, and alternations between u and w on the ambiguous grapheme و. Both are properties of the source writing system rather than of any component of the pipeline. Reported rates from this evaluation are therefore measured on a scale whose zero point is not zero, and I would discourage citing the absolute figures as Garrusi recognition rates.
Against that background, the comparison between systems is more informative than either level. A Southern Kurdish fine-tuned system, scored through the same pipeline over the 1,703 segments common to both systems, trails the Central Kurdish adapter by 10.69 percentage points of WER (95% CI [8.65, 12.80]; speaker-cluster [7.87, 14.75]), on all five speakers and on both metrics. The folding table covers the characters both systems emit, so the difference does not arise from material being discarded before alignment; whether a residual pipeline asymmetry remains is not established here (§6.5).
The difference is nonetheless strongly dependent on the scoring representation. Treating short /i/ as non-contrastive in reference and hypothesis alike reduces it from 10.69 to 2.13 percentage points, a reduction of 80.1%, with intervals that include zero under both resampling units. That manipulation is not a further stage of the common-reference design, because it changes the reference: the denominator falls from 9,742 to 9,684 tokens, and 16.72% of reference tokens become lexically ambiguous, chiefly by merging the ezafe-marked and bare forms of frequent nouns. I therefore report it as a sensitivity analysis and not as a better scoring convention, and I do not conclude that the systems are equivalent: an interval containing zero is not evidence of no difference, the point estimate remains 2.13 points, and there are five speakers.
What follows for practice is narrow. Where a reference orthography encodes distinctions that the input writing system does not, the scoring representation can change a measured difference between two systems by a factor of five, and a scoring pipeline can impose a floor of nearly fifty points of WER on input that is already correct. I suggest that cross-script Kurdish ASR evaluation report the substitution, deletion, insertion and hit counts, state the scoring representation in full, release the fixed reference alongside the rate, and where possible report a round-trip control, so that a reader can see how much of a reported rate is attributable to the measuring instrument before attributing any of it to the recognizer.
Three further steps remain. The round-trip control should be reproduced with segment identifiers retained, so that its floor estimate can be tied to a documented sample (§6.6). A corresponding control constructed from Southern Kurdish material would measure the pipeline floor for the second system and is the correct way to settle the question §6.5 leaves open. And one practical obstacle to the vernacular comparison is worth recording, because it is straightforward to remove: contrasting Garrusi against the other Southern Kurdish vernaculars on the existing benchmark requires grouping its recordings by variety, and the vernacular labels are reported in Table 2 of Mohammadamini and Tahon [10] as a description of the eight speakers rather than carried as a grouping variable in the evaluation data as distributed. A per-vernacular rate is therefore a reconstruction, obtained by matching a printed table to the speaker identifiers in the release, not a computation over the released fields. The information exists and the reconstruction is not difficult. What matters is that the analysis the paper describes and the analysis the download supports are not the same, and carrying speaker-level vernacular labels in the distributed metadata would make per-vernacular evaluation, including for Garrusi, directly available to anyone using the benchmark.
Author Contributions
Conceptualization, H.A.; methodology, H.A.; software, H.A.; validation, H.A.; formal analysis, H.A.; investigation, H.A.; resources, H.A.; data curation, H.A.; writing, original draft preparation, H.A.; writing, review and editing, H.A.; visualization, H.A.; supervision, H.A.; project administration, H.A.; funding acquisition, H.A. Masoumeh Zarei contributed to the fieldwork and to the initial transcription and translation of the recordings used as source material for this study. The present study's conceptualization, methodology, technical implementation, analysis, interpretation, and writing were carried out by H.A. The author has read and agreed to the published version of the manuscript.
Institutional Review Board Statement
No institutional ethics committee or institutional review board approval was obtained for the original fieldwork. The recordings were collected independently by the author during linguistic fieldwork in Iran between 2017 and 2026 and were collected specifically for linguistic research. The fieldwork was conducted in a sensitive geopolitical and social context in which creating or retaining written records identifying participants could have created additional risks to their privacy and safety. For this reason, the fieldwork was conducted without collecting written participant-identifying documentation. Before each recording, participants were informed orally that they were being recorded for linguistic research and that their speech could be analyzed for linguistic purposes and subsequently used in research on speech technologies. Participation was voluntary, and oral consent was obtained before recording. No participant-identifying information is reported in this article, and speakers are identified only by anonymized two-letter codes. The author reports the absence of institutional ethics approval explicitly and does not represent the present procedure as equivalent to formal institutional review. The present study is a secondary analysis of these previously collected recordings; no new recordings were made and no participants were recontacted for this study.
Informed Consent Statement
Oral informed consent was obtained from all participants before recording. Participants were informed that their speech was being recorded for linguistic research, that the recordings would be used for linguistic analysis, and that the resulting material could subsequently be used in research on speech technologies, including speech recognition. Participation was voluntary. Written consent forms were not used because the creation and retention of written participant records could have created additional privacy and safety risks in the field context. No identifying participant information is reported in this article, and speakers are referred to only by anonymized two-letter codes. No new recordings or participant contact were undertaken for the present study.
Data Availability Statement
The audio recordings analyzed in this study are not publicly available because of the privacy, safety, and data-sharing considerations associated with the circumstances under which the field recordings were collected. No participant-identifying information is released. Subject to the data-sharing terms applicable to the source corpus and the protection of participant confidentiality, the fixed reference used for evaluation, segment-level scoring results, and the scripts required to reproduce the reported scoring procedure may be made available where this can be done without exposing restricted source material or participant information.
Acknowledgments
I thank Masoumeh Zarei for her contribution to the Garrusi Kurdish fieldwork and for carrying out the initial transcription and translation of the recordings. I also thank her for her work in the initial preparation of the field materials. The transcription and translation used in the present study were subsequently checked, reviewed, and revised as part of the data-quality control process. I also thank the speakers who contributed their time and speech to the linguistic field corpus. Because of the sensitive context in which the recordings were collected, participants are not individually identified in this article.
Conflicts of Interest
Conflicts of Interest: The author declares no conflicts of interest.
Abbreviations
The following abbreviations are used in this manuscript:
| ASR | Automatic speech recognition |
| WER | Word error rate |
| CER | Character error rate |
| MMS | Massively Multilingual Speech |
| CTC | Connectionist temporal classification |
References
- Ahmadi, S. KLPT—Kurdish Language Processing Toolkit. Proc. Second Workshop NLP Open Source Software (NLP-OSS), 2020; pp. 72–84. [Google Scholar] [CrossRef]
- Ahmadi, S.; Jaff, D.Q.; Alam, M.M.I.; Anastasopoulos, A. Language and Speech Technology for Central Kurdish Varieties. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Torino, Italy, 20–25 May 2024; ELRA and ICCL: Torino, Italy, 2024; pp. 10034–10045. [Google Scholar]
- Ahmadi, S. A Rule-Based Kurdish Text Transliteration System. ACM Trans. Asian Low.-Resour. Lang. Inf. Process. 2019, 18, 18:1–18:8. [Google Scholar] [CrossRef]
- Asadpour, H. Cross-Dialectal Diversity in Mukrī Kurdish I: Phonological and Phonetic Variation. J. Linguist. Geogr. 2021, 9, 1–12. [Google Scholar] [CrossRef]
- Asadpour, H.; Zarei, M. Corpus and Experimental Analysis of Passive Structures in Garrusi Kurdish. Languages 2026a, 11, 63. [Google Scholar] [CrossRef]
- Asadpour, H.; Zarei, M. Passive Constructions in Garrusi Kurdish. In Passivisation in Semitic, Iranian, Armenian, and Beyond;Cambridge Semitic Languages and Cultures; Noorlander, P.M., Asadpour, H., Eds.; Open Book Publishers: Cambridge, UK, 2026b; pp. 233–276. [Google Scholar] [CrossRef]
- Belelli, S. Towards a Dialectology of Southern Kurdish: Where to Begin? In Current Issues in Kurdish Linguistics; Gündoğdu, S., Öpengin, E., Haig, G., Anonby, E., Eds.; University of Bamberg Press: Bamberg, Germany, 2019; pp. 73–92. [Google Scholar] [CrossRef]
- Haig, G.; Öpengin, E. Introduction to Special Issue—Kurdish: A Critical Research Overview. Kurd. Stud. 2014, 2, 99–122. [Google Scholar] [CrossRef]
- Matras, Y. Revisiting Kurdish Dialect Geography: Findings from the Manchester Database. In Current Issues in Kurdish Linguistics; Gündoğdu, S., Öpengin, E., Haig, G., Anonby, E., Eds.; University of Bamberg Press: Bamberg, Germany, 2019; pp. 225–241. [Google Scholar] [CrossRef]
- Mohammadamini, M.; Tahon, M. Southern Kurdish Speech Recognition Resources and Benchmarking. In Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026), Palma de Mallorca, Spain, 11–16 May 2026; ELRA Language Resource Association: Paris, France, 2026; pp. 5538–5544. [Google Scholar] [CrossRef]
- Mohammadamini, M.; Mohammed, A.J.; Mohammed, B.H.; Abdulazeez, D.H.; Sadeeq, I.S.; Mohammed Salih, D.; Melhum, A.I.; Dheyab, A.A. Exploring the Reusability of Northern Kurdish Resources for Badini Speech Recognition. In Proceedings of the Workshop on Dialects in NLP (DialRes 2026), Palma de Mallorca, Spain, 16 May 2026; ELRA Language Resource Association; p. 2026. [Google Scholar]
- Pratap, V.; Tjandra, A.; Shi, B.; Tomasello, P.; Babu, A.; Kundu, S.; Elkahky, A.; Ni, Z.; Vyas, A.; Fazel-Zarandi, M.; Baevski, A.; Adi, Y.; Zhang, X.; Hsu, W.-N.; Conneau, A.; Auli, M. Scaling Speech Technology to 1,000+ Languages. J. Mach. Learn. Res. 2024, 25, 1–52. [Google Scholar]
| 1 | Searched Google Scholar, ACL Anthology, HAL, arXiv, IEEE Xplore, and Scopus through 12 August 2026 for the variety name and its spelling variants (Garrusi, Gerrûsî, Garrousi, Bijari, Bîcarî) in combination with speech recognition, ASR, WER, CER, corpus, MMS, wav2vec and Whisper. I found linguistic work on the variety, and one benchmark that includes a Garrusi speaker among several vernaculars [10], but no evaluation reporting recognition rates for the variety as such. |
| 2 | I am grateful to Masoumeh Zarei for her research assistance on the fieldwork and data collection on which this corpus is based. The recognition experiments, the scoring design, and the analysis reported here are my own, as are any errors that remain. |
Figure 1.
Alignment composition in each scoring condition, against the fixed 9,763-token reference. Hits, substitutions, and deletions sum to the reference length in every condition (dashed line); insertions extend beyond it, which is why WER exceeds 100% where hits are few. The two hypothesis-side transformations convert substitutions into hits while leaving deletions and insertions comparatively stable.
Figure 1.
Alignment composition in each scoring condition, against the fixed 9,763-token reference. Hits, substitutions, and deletions sum to the reference length in every condition (dashed line); insertions extend beyond it, which is why WER exceeds 100% where hits are few. The two hypothesis-side transformations convert substitutions into hits while leaving deletions and insertions comparatively stable.

Figure 2.
Per-segment WER by reference-segment length, MMS-1B-all in the FOLDED condition, matched set. Points are bin medians, bars the interquartile range (not the full range); n per bin is shown along the axis. Median WER is exactly 1.00 in every bin from one to seven reference tokens (together 1,249 segments, 73.3% of the set) and falls only for longer segments (median 0.889 for 8–9 tokens, n = 248; median 0.833 for 10 or more, n = 206). The qualifications in the text apply: part of this pattern follows from the definition of the metric rather than from recognition difficulty.
Figure 2.
Per-segment WER by reference-segment length, MMS-1B-all in the FOLDED condition, matched set. Points are bin medians, bars the interquartile range (not the full range); n per bin is shown along the axis. Median WER is exactly 1.00 in every bin from one to seven reference tokens (together 1,249 segments, 73.3% of the set) and falls only for longer segments (median 0.889 for 8–9 tokens, n = 248; median 0.833 for 10 or more, n = 206). The qualifications in the text apply: part of this pattern follows from the definition of the metric rather than from recognition difficulty.

Table 1.
The two segment sets used in this paper. The 1,703-segment set is the subset of the 1,722-segment set for which the second system also returned a hypothesis; §4.5 accounts for the 19 segments that differ. Reference token counts are for the folded reference of §3.3.
Table 1.
The two segment sets used in this paper. The 1,703-segment set is the subset of the 1,722-segment set for which the second system also returned a hypothesis; §4.5 accounts for the 19 segments that differ. Reference token counts are for the folded reference of §3.3.
| Set | Segments | Ref. tokens | Ref. characters | Used for |
|---|---|---|---|---|
| Full | 1,722 | 9,763 | 53,017 | §4.1–§4.4, MMS-1B staged normalization |
| Matched | 1,703 | 9,742 | 52,942 | §4.5, §4.7, §4.8, two-system comparison |
| Difference | 19 | 21 | 75 | §4.5, sensitivity check |
Table 6.
Sensitivity of the system difference to the 19 segments. The matched row excludes them, as in Table 5. The substituted row scores all 1,722 segments, treating the second system’s failures as empty hypotheses, which is how an empty output would be scored under the protocol of §3.1. Differences are Aranemini minus MMS. The two difference columns rest on different reference bases and are not comparable with each other (§3.5); each column should be read down, comparing the two treatments within one representation.
Table 6.
Sensitivity of the system difference to the 19 segments. The matched row excludes them, as in Table 5. The substituted row scores all 1,722 segments, treating the second system’s failures as empty hypotheses, which is how an empty output would be scored under the protocol of §3.1. Differences are Aranemini minus MMS. The two difference columns rest on different reference bases and are not comparable with each other (§3.5); each column should be read down, comparing the two treatments within one representation.
| Treatment of the 19 segments | Segments | FOLDED ref tokens | Δ FOLDED | FOLDED-i ref tokens | Δ FOLDED-i |
|---|---|---|---|---|---|
| Excluded (matched set) | 1,703 | 9,742 | 10.686 pp | 9,684 | 2.127 pp |
| Scored as empty hypotheses | 1,722 | 9,763 | 10.652 pp | 9,705 | 2.112 pp |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.