Submitted:
02 September 2026
Posted:
03 September 2026
You are already at the latest version
Abstract
Historical music dictionaries are now widely accessible as scans and optical-character-recognition text, but they remain difficult to reuse side by side. Their languages, editions, editorial forms, and digitisation histories differ too much to treat them as a single uniform archive. This Data Descriptor presents Lexicon C19, an English-dominant multilingual computational corpus derived from eight nineteenth-century music, dance, biographical, and general reference traditions, with processed witnesses in English, French, German, Russian, and Portuguese. The v3.2.1 release contains 3,760 rule-selected records, a complete pool of 25,186 parsed candidates, source and line provenance, a controlled music-performance taxonomy, supporting metadata, deterministic build code, and an executed analysis notebook. Structural checks confirm identifier uniqueness, complete local-file provenance, stable checksums, and deterministic rebuilding. A label-blinded 200-record single-annotator audit estimates selected-candidate precision at 97.0% and category accuracy at 89.7% among real and relevant sampled rows. Labels and confidence values remain machine-derived corpus fields rather than gold-standard annotations. The corpus is designed for source-stratified retrieval, computational lexicography, OCR research, weak supervision, and the selection of passages for verification and close reading.
Keywords:
historical musicology
; digital lexicography
; music dictionaries
; humanities data
; OCR
; provenance
; computational musicology
1. Introduction
Nineteenth-century reference works did more than define musical terms. They organised musical knowledge for students, amateurs, practitioners, critics, and scholars, and they did so through very different editorial habits. Their entries framed instruments, voices, dance steps, performance practice, musical forms, institutions, repertories, and biographies. Digitisation has made many of these works searchable, but it has not made them computationally equivalent. A collaborative music dictionary, a single-author dance dictionary, a biographical series, and a curated set of encyclopaedia articles do not share the same unit of analysis. Nor do full-volume optical-character-recognition (OCR) files, manually structured Wikisource articles, and translated editions offer the same conditions of recoverability. This Data Descriptor documents the citable Lexicon C19 v3.2.1 release [1].
Histories of nineteenth-century lexicography show that dictionaries were not neutral containers of words. They took part in projects of classification, standardisation, scientific description, national representation, and imperial or transregional knowledge-making, while depending on varied forms of authorship and collaboration [3,4,5]. Lexicon C19 brings music-specific and music-adjacent reference works into this wider history. Its computational design therefore keeps the processed witnesses visible instead of smoothing them into interchangeable observations from one historical population. The publication package follows the same principle by treating provenance, reuse conditions, and claim boundaries as part of the dataset description, in line with FAIR data principles and data-statement practice [14,15].
Lexicon C19 was created to support transparent retrieval across this heterogeneous material. The resource combines selected records from eight source traditions: John Weeks Moore’s Complete Encyclopaedia of Music [6]; George Grove’s four-volume Dictionary of Music and Musicians [7]; Hugo Riemann’s Musik-Lexikon as represented by J. S. Shedlock’s 1908 English translation [8]; François-Joseph Fétis’s Biographie universelle des musiciens together with material from Arthur Pougin’s supplement [9]; G. Desrat’s Dictionnaire de la danse [10]; music articles from the fourth edition of Meyers Konversations-Lexikon [11]; music articles from the Brockhaus-Efron Encyclopaedia [12]; and Ernesto Vieira’s Diccionario biographico de musicos portuguezes [13].
The resource should not be read as a balanced sample of European musical thought. Three processed witnesses, Riemann/Shedlock, Moore, and Grove, account for 81.9% of selected records. The Fétis material lacks main-series volume 1, Vieira produces only one selected record, and the Meyers and Brockhaus-Efron inputs are curated music-article collections rather than full encyclopaedia OCR. These differences are retained in the metadata because they matter for interpretation.
The release has two linked analytical layers. The selected CSV contains 3,760 candidates assigned by transparent, ordered keyword rules to six broad domains: instruments, performance, voice, music concepts, dance, and sound. The complete JSON Lines pool contains all 25,186 parsed candidates, including spans assigned to GENERAL/OTHER. This separation lets users work with the music-performance index while still inspecting the broader behaviour of the selection procedure.
Version 3.2 corrected structural defects identified in the previously deposited exploratory v3.1 package [2]. In v3.1, 44 rows had duplicated identifiers, printed-page values were empty, and the processed Riemann witness was described as German/1882 although the local text is Shedlock’s English translation, fourth edition, 1908. Version 3.2 used deterministic identifiers across constituent source files, added local file and starting-line provenance, separated processed-witness language and date from original publication metadata, and published the complete candidate pool. Version 3.2.1 preserves those selected and complete data tables, and adds the completed 200-record audit workbook, an explicit MIT code-licence file, and refreshed publication-analysis outputs. Printed page remains unavailable and is reported as a limitation rather than inferred.
The data support source-stratified retrieval, historical OCR research, rule-based classification experiments, provenance-aware interfaces, computational lexicography, and the discovery of passages for scan verification and close reading. They should not be used as a gold-standard annotation set or as direct evidence that a term, category, place, or practice was historically more important in one national tradition than another.
2. Data Description
2.1. Dataset Identification and Release Contents
The v3.2.1 package brings together the selected CSV, complete candidate-pool JSON Lines file, source inventory and taxonomy metadata, multilingual alignment and semantic-relation tables, deterministic extraction code, clean and executed analytical notebooks, twelve derived CSV tables, nine figure designs in PNG, SVG, and PDF, release notes, citation metadata, the completed 200-record audit workbook, an explicit MIT code-licence file, and a checksum manifest. The principal data filenames retain the v3_2 stem because the selected table and candidate pool are unchanged in the v3.2.1 maintenance release.
Table 1.
Lexicon C19 v3.2.1 dataset identification.
| Element | Value |
|---|---|
| Dataset title | Lexicon C19 Export Package v3.2.1 |
| Creator | Xavier Fresquet |
| Repository and publisher | Zenodo |
| Concept DOI | https://doi.org/10.5281/zenodo.21588976 |
| Version DOI | https://doi.org/10.5281/zenodo.22207277 |
| Version history | v3.1: https://doi.org/10.5281/zenodo.21588977; v3.2: https://doi.org/10.5281/zenodo.22194770 |
| Dataset location | https://doi.org/10.5281/zenodo.22207277 |
| Data licence | CC BY 4.0 for data, metadata, documentation, and analysis outputs |
| Code licence | MIT for scripts and executable notebooks in the project repository |
2.2. Main Selected Table
The main selected table, datasets/master_performance_dataset_v3_2.csv, contains 3,760 records and 27 fields. Table 2 summarises the schema using the exact field names in the CSV.
The page, example, multilingual_alignment, and notes fields are empty throughout the selected table. They remain in the schema for compatibility and to make absence explicit. Definitions are capped at 800 characters; 116 selected records reach that cap.
Table 3.
Illustrative records from the selected table. Definitions are shortened here for display.
| Source | Category | Headword | Definition excerpt |
|---|---|---|---|
| Riemann/Shedlock | INSTRUMENTS | Brunetti | “Gaetano, performer on the violin,”; a compact but readable translated-biographical entry. |
| Moore | VOICE | ARIABUFFA | “A comic air.” |
| Meyers | MUSIC_CONCEPTS | C dur | German article on C major and the C-major triad. |
| Desrat | DANCE | JETÉ | French dance-step definition beginning “Nom d’un pas de danse”. |
| Fétis/Pougin | INSTRUMENTS | VIDAL (B.) | Biographical entry identified through “professeur de guitare”. |
2.3. Complete Candidate Pool
The complete pool, validation/all_extracted_entries_v3_2.jsonl, contains 25,186 parsed candidates. It includes the selected-table fields and predicted_relevant, a Boolean indicating whether the ordered ontology assigned a category other than GENERAL/OTHER. This file is the best starting point for alternative classifiers, future annotation, and segmentation-error studies because it preserves candidates rejected by the active music-performance selection rules.
2.4. Supporting Data, Code, and Notebook
metadata/dictionary_inventory.json documents the source witnesses, including processed edition, language, source URL, local file, machine-readable status, and reuse status. metadata/performance_taxonomy.json records the broader conceptual inventory. datasets/multilingual_lexicon.csv contains 27 curated cross-language concept alignments, and datasets/semantic_relations.csv contains 539 declared relation pairs. These companion resources are aids for navigation and reuse; they are not learned topics or human-validated labels.
scripts/build_dataset.py rebuilds the corpus from local source texts. notebooks/03_publication_analysis.ipynb is the clean analytical notebook, and notebooks/03_publication_analysis_executed.ipynb retains verified outputs. The release’s publications/analysis_outputs/ directory contains the exported tables, figures, summary, captions, alt text, and JSON analysis manifest.
2.5. Dataset Composition
The complete pool contains 25,186 parsed candidates, of which 3,760 are selected by the performance ontology. Source-specific selected-record counts range from one for Vieira to 1,295 for Riemann/Shedlock. Candidate-selection rates range from 3.1% for Vieira (1 selected record from 32 parsed candidates) and 5.2% for Fétis/Pougin (141 from 2,722) to approximately 45–56% for the curated or specialist Desrat, Meyers, and Brockhaus-Efron inputs. These rates are recoverability diagnostics for the processed texts and selection rules. They are not estimates of historical prevalence. In particular, the low Vieira and Fétis/Pougin values point to under-selection or under-extraction relative to the size and biographical character of the printed works.
Figure 1.
Corpus composition and rule-based selection yield. Parsed candidates and selected records by source tradition, with the selected share of parsed candidates. Counts are diagnostic pipeline outputs rather than estimates of historical prevalence.
Figure 1.
Corpus composition and rule-based selection yield. Parsed candidates and selected records by source tradition, with the selected share of parsed candidates. Counts are diagnostic pipeline outputs rather than estimates of historical prevalence.

Figure 2.
Rule-based category profiles by source. Within-source percentages for the six selected categories. Vieira has one selected record and is shown only for completeness.
Figure 2.
Rule-based category profiles by source. Within-source percentages for the six selected categories. Vieira has one selected record and is shown only for completeness.

Figure 3.
Lexicon C19 v3.2/v3.2.1 visual overview. (A) Rule-selected records by processed witness. (B) Overall rule-derived category composition. (C) Source-level selection rate against median OCR-surface confidence, with marker area scaled by selected-record count. (D) Field completeness by source. The overview characterises the released computational dataset and its quality dimensions rather than historical prevalence.
Figure 3.
Lexicon C19 v3.2/v3.2.1 visual overview. (A) Rule-selected records by processed witness. (B) Overall rule-derived category composition. (C) Source-level selection rate against median OCR-surface confidence, with marker area scaled by selected-record count. (D) Field completeness by source. The overview characterises the released computational dataset and its quality dimensions rather than historical prevalence.

3. Methods: Implementation and Evaluation
3.1. Source Corpus and Processed Witnesses
The corpus uses public-domain historical texts acquired as OCR or structured transcription. Moore’s 1854 encyclopaedia and the four Grove volumes were processed from Internet Archive OCR. Desrat and Vieira were likewise processed from public-domain digitised books. The local Fétis/Pougin collection contains the second-edition main-series volumes 2–8, dated 1860 in the dataset metadata, together with Pougin’s Supplement I from the later 1878–1880 continuation; main-series volume 1 is absent. The Riemann witness is not the German first edition. It is the English text of J. S. Shedlock’s translation, fourth edition, 1908, based on a work first published in 1882. Meyers and Brockhaus-Efron are represented by structured music-article collections derived from Wikisource. The Brockhaus-Efron source is the wider 1890–1907 encyclopaedia, but the processed music-article subset is labelled 1896 in the dataset metadata. The latter inputs have cleaner surface structure but narrower sampling frames than full-volume OCR.
3.2. Entry Segmentation
The build script reads local inputs in a fixed order and applies source-specific regular expressions to detect candidate entry boundaries. A parsed candidate begins at a detected headword pattern and extends to the next boundary or a source-specific limit. Basic filters remove spans consisting only of page numbers, obvious short all-capital fragments, or recurrent alternating-case OCR noise, but this filtering is intentionally conservative rather than exhaustive. Each retained candidate stores its extracted headword, a normalised lowercase form, an excerpt of the definition, a source-text excerpt, and the local starting line.
Segmentation rules differ because the processed witnesses differ. Grove’s volume files, Desrat’s dictionary, the Riemann/Shedlock translation, and the biographical materials do not mark entries in identical ways. The Wikisource-derived inputs use heading-like separators that are more explicit than the typography recovered in OCR. The full candidate pool is retained so alternative selection or annotation procedures can work from the same segmented material.
3.3. Rule-Based Selection and Classification
An ordered multilingual ontology maps lexical cues to a category and subcategory. Its six selected categories are INSTRUMENTS, PERFORMANCE, VOICE, MUSIC_CONCEPTS, DANCE, and SOUND. Subcategories include instrument families, vocal genres and types, theatrical, sacred, public, courtly, and concert contexts, dance types, theoretical concepts, forms, and performance directions. Candidates that match no selected cue are assigned GENERAL/OTHER and remain available in the full pool.
The assignment is deterministic, inspectable, order-sensitive, and single-label. When several cues occur, the first applicable ontology rule determines the category. The labels should therefore be read as computational indexing decisions, not as concepts transcribed from the historical sources or as human-validated ground truth.
Additional regular expressions derive optional semantic tags, geographic references, historical references, etymological phrases, and cross-references. Semantic tags include signals such as sacred, courtly/elite, festive, ancient/medieval, and language- or place-associated terms. These fields support retrieval but inherit the limits of surface matching. The current taxonomy has no dedicated biographical class; therefore category distributions for Fétis/Pougin and Vieira should be read only as retrieval labels for musically relevant biographical entries, not as substantive descriptions of those sources.
confidence_score is a heuristic estimate based on OCR surface characteristics; it is bounded between zero and one but is not calibrated as a probability. For a retained text span of length N, let r be the number of repeated-space runs and a the number of characters outside ASCII, extended Latin, Cyrillic, and whitespace. The implemented score is:
Very short all-capital spans of one to four characters that survive the preceding entry-boundary filters return 0.1 directly. Scores are rounded to three decimals.
3.4. Provenance and Witness Metadata
Version 3.2 assigned identifiers deterministically within each source tradition across all constituent files; v3.2.1 retains those identifiers unchanged. The identifier is an opaque legacy key plus a zero-padded sequence, for example EN-MOORE-1854-00001, RU-BROCKHAUS-00001, or DE-RIEMANN-1882-00019. The Riemann key preserves a legacy shorthand for the German original and must not be read as the processed witness language or date. Every selected record includes source_file and source_line_start, allowing a user to return to the processed local witness. source_text_language records the language actually processed. publication_year records the processed witness date, and original_publication_year records the date of an earlier work when the processed edition is later. This distinction is crucial for Riemann/Shedlock and prevents an English translation from being analysed as German-language text.
The article_type field distinguishes specialised music reference works, scholarly or biographical works, a dance dictionary, and Konversationslexikon article collections. It permits analyses to be stratified by sampling frame without treating the field as a historical ontology.
3.5. Deterministic Implementation
The active build uses Python’s standard library for extraction and writes the selected CSV, complete JSON Lines candidate pool, multilingual alignment table, semantic relation table, and quality report. Inputs are processed in sorted order. Identifiers are assigned after source-level aggregation rather than reset for each constituent file. The v3.2 selected table is written as a separate versioned output and is the unchanged selected table used in the public v3.2.1 Zenodo package.
The publication notebook uses pandas, NumPy, Matplotlib, and Seaborn. It loads the frozen v3.2/v3.2.1 CSV and candidate pool, checks schema and provenance, computes source and category profiles, exports twelve derived CSV tables, and writes nine figure designs in PNG, SVG, and PDF. An executed copy retains all cell outputs, while the JSON manifest records software versions, counts, checksums, missingness, and generated files.
3.6. Structural and Reproducibility Evaluation
The v3.2/v3.2.1 selected CSV contains 3,760 records and 27 fields. Its SHA-256 checksum is a82a89cde3125edd5fa2009fcb2a4da6acc1c2509bed22c03eccb67211054374. All identifiers are unique; source_file, source_line_start, and article_type are complete. The candidate-pool JSON Lines file contains 25,186 records and has SHA-256 checksum 4dce413d982c9765c90f12c8d815349a81e30cd557f9aa2a264855f4d5ab5259.
The build was executed twice from the same local OCR inputs and local software environment on a single machine. Both the selected CSV and full candidate-pool checksums were identical across runs.
3.7. Witness-Level Evaluation
All 1,295 Riemann/Shedlock selected records identify English as the processed language, 1908 as the processed publication year, and 1882 as the original publication year. Grove records are distributed across volumes 1–4. Fétis/Pougin records identify the second-edition main-series volumes 2–8 and Supplement I, while the absence of main-series volume 1 is documented. Brockhaus-Efron records identify a Russian structured music-article subset labelled 1896, within a reference work published across 1890–1907. The Meyers input is identified as German Wikisource text from the fourth-edition Konversations-Lexikon rather than undifferentiated German OCR.
The executed notebook verifies the controlled category vocabulary and confirms that all confidence values fall between zero and one. It also shows why confidence should not be treated as a universal quality score. Meyers and Brockhaus-Efron have median values of 1.000 and 0.952 because their structured Wikisource-derived inputs are clean at the surface level. Median values for full-volume OCR range from 0.333 for Grove to 0.415 for Moore. These differences mainly reflect digitisation mode and surface regularity.
Exact normalised-headword overlap is low. The largest pairwise Jaccard value is 4.98% between Grove and Riemann/Shedlock, followed by 3.41% between Moore and Grove. Cross-language overlap is generally lower. These values provide an interoperability diagnostic but cannot separate conceptual difference from translation, morphology, OCR error, or editorial genre.
3.8. Manual Sample Audit
In addition to structural and reproducibility checks, Xavier Fresquet completed a label-blinded, source-stratified single-annotator audit of 200 selected records on 31 August 2026. During annotation, the pipeline category, subcategory, confidence band, stratum population, sample weight, and pipeline-vs-gold scoring fields were hidden. The annotator saw the record identifier, source, local file and line provenance, headword, and definition excerpt, then recorded real-entry status, relevance, broad gold category, subcategory, boundary quality, OCR readability, confidence, and notes. This is label-blinding rather than full independent validation: the annotator was also the ontology author, and the source and identifier remained visible.
The audit used near-proportional source allocation with deliberate minimum coverage for small witnesses and proportional allocation within each source across pipeline category and confidence band. Table 4 reports the source allocation and real-plus-relevant counts. All 200 sampled rows were completed. Of these, 194 were marked Yes for both real-entry and relevance. No row was marked explicit No; six rows were marked Unclear for both fields and were conservatively scored outside the real-plus-relevant numerator. The unweighted sample selected-candidate precision is therefore 97.0% (194/200; 95% Wilson interval: 93.6–98.6%). The corresponding source-weighted precision estimate, using the source populations in Table 4, is 96.9%.
Among the 194 real and relevant rows, 174 matched the pipeline’s broad category, giving sample category accuracy of 89.7% (95% Wilson interval: 84.6–93.2%). The 20 broad-category mismatches were distributed across Moore (5), Riemann/Shedlock (5), Brockhaus-Efron (4), Grove (3), Meyers (2), and Fétis/Pougin (1). By pipeline category, mismatches involved MUSIC_CONCEPTS (8), PERFORMANCE (5), VOICE (4), DANCE (2), and INSTRUMENTS (1). The most common gold-side correction was toward INSTRUMENTS (11 cases), followed by VOICE (5), DANCE (2), and MUSIC_CONCEPTS (2).
The audit also reports quality fields that qualify the precision result. Boundary quality was marked Complete for 199 rows and Fragment for one row; OCR readability was marked Good for all 200 rows. In this audit, Good means readable enough to make the annotation decision; it does not mean clean OCR. This does not contradict the low median surface-confidence values for Grove (0.333) and Moore (0.415), which reflect character-level and spacing noise that often leaves headwords and opening clauses legible. The single fragment row was Brockhaus-Efron Violonchel’; it was also the only row carrying an error flag, Truncated definitions, and the only row assigned low decision confidence. Table 5 reports these distributions. These counts show that the precision estimate measures whether sampled selected candidates point to real and musically relevant entries, not whether every definition excerpt is complete enough for quotation. Brockhaus-Efron illustrates the same distinction: all eight sampled rows were real and relevant, but four of those eight had broad-category mismatches.
3.9. Evaluation Scope and Ethical Considerations
The audit supports sample-level precision and broad-category accuracy for selected records. The Wilson intervals reported above are descriptive binomial intervals for the completed audit sample, not full design-based confidence intervals for all possible source-stratified samples. The audit does not estimate segmentation recall, relevance recall, inter-annotator agreement, or the completeness of any source witness. A larger double-annotation and page-recall campaign remains deferred to a later student internship and should be published as a separate versioned validation layer. Users should treat the selected table as a transparent machine-generated index and verify passages against scans before quotation or historical interpretation.
The corpus contains public-domain historical reference works and no living-person or sensitive personal data. Machine-derived categories should not be used to infer stable national, ethnic, or cultural traits. Place- and language-associated tags identify textual matches, not identities or attitudes.
4. User Notes
Analyses should normally be stratified by source_dictionary and, where relevant, by article_type. Aggregate category totals are dominated by three English-language processed witnesses and combine non-equivalent sampling frames. Meyers and Brockhaus-Efron should not be compared directly with full dictionaries unless their pre-filtered music-article sampling frames are made explicit.
The Riemann records represent Shedlock’s English translation and should not be used to study Riemann’s German wording. Fétis coverage is incomplete, and Vieira’s single selected record makes it unsuitable for quantitative comparison. Because the page field is empty, quotations require return to the digitised scan or another authoritative edition. source_file and source_line_start locate the processed text but do not substitute for printed-page citation.
The definition field is an excerpt and may omit later parts of long articles. Category, subcategory, semantic, geographic, and historical fields are retrieval aids produced by ordered keyword and regular-expression rules. Network centrality based on the companion relations would inherit these design decisions and should not be interpreted as historical importance.
Recommended uses include:
- source-specific and provenance-aware search;
- teaching reproducible historical-data criticism;
- developing alternative multilingual classifiers;
- weak supervision and candidate selection for later annotation;
- measuring OCR and segmentation behaviour;
- building provenance-aware interfaces;
- identifying passages for scan verification and close reading.
Random train-test splits are not recommended for machine-learning experiments because source-specific OCR and editorial patterns may appear in both partitions. Source-held-out evaluation is preferable. The completed 200-record audit may guide exploratory reuse, but any future gold-standard annotation should be published as a separate versioned layer with sampling documentation, annotator guidance, agreement statistics, uncertainty, and adjudication records.
Author Contributions
Conceptualization, X.F.; methodology, X.F.; software, X.F.; validation, X.F.; formal analysis, X.F.; investigation, X.F.; resources, X.F.; data curation, X.F.; writing—original draft preparation, X.F.; writing—review and editing, X.F.; visualization, X.F.; project administration, X.F.; funding acquisition, X.F. The author has read and agreed to the published version of the manuscript.
Funding
This research was funded by ANR MELODY, grant number ANR-24-IAS1-0001.
Institutional Review Board Statement
Not applicable. The study did not involve humans or animals.
Informed Consent Statement
Not applicable. The dataset contains public-domain historical reference works and no human participants, living-person records, or personal data requiring consent.
Data Availability Statement
Lexicon C19 v3.2.1 is publicly available on Zenodo at https://doi.org/10.5281/zenodo.22207277 under concept DOI https://doi.org/10.5281/zenodo.21588976. Previous versions remain citable at https://doi.org/10.5281/zenodo.22194770 for v3.2 and https://doi.org/10.5281/zenodo.21588977 for v3.1. The v3.2.1 deposit contains Lexicon_C19_Export_Package_v3.2.1_2026-08-31.zip, including the selected CSV, complete candidate-pool JSON Lines file, metadata, deterministic build code, clean and executed notebooks, analytical tables, figures, release notes, citation metadata, licence statement, checksum manifest, the completed 200-record audit workbook, and LICENSE-CODE. Data, metadata, documentation, and analysis outputs are released under CC BY 4.0. Scripts and executable notebooks are released under the MIT License. Raw OCR source files are excluded; source-provider terms remain applicable to materials obtained from the linked repositories.
Acknowledgments
This work acknowledges support from ANR MELODY (ANR-24-IAS1-0001). It was developed within the musicological, digital-humanities, and computational research environments of IReMus, SCAI, and SAFIR. During preparation of this manuscript, the author used OpenAI Codex (accessed August 2026) to assist with manuscript structuring, LaTeX conversion, and figure integration. The author reviewed and edited the output and takes full responsibility for the content of the publication.
Conflicts of Interest
The author declares no conflict of interest. The funder had no role in the design of the corpus; in the collection, analysis, or interpretation of the data; in the writing of the manuscript; or in the decision to publish the results.
References
- Fresquet, X. Lexicon C19 Export Package v3.2.1: A Multilingual Dataset of Nineteenth-Century Music Performance Lexicography; Version 3.2.1; Zenodo: Geneva, Switzerland, 2026. [CrossRef]
- Fresquet, X. Lexicon C19 Export Package; Version 3.1; Zenodo: Geneva, Switzerland, 2026. [CrossRef]
- Ogilvie, S.; Safran, G., Eds. The Whole World in a Book: Dictionaries in the Nineteenth Century; Oxford University Press: Oxford, UK, 2020. [CrossRef]
- Considine, J., Ed. The Cambridge World History of Lexicography; Cambridge University Press: Cambridge, UK, 2019. [CrossRef]
- Haß, U., Ed. Große Lexika und Wörterbücher Europas: Europäische Enzyklopädien und Wörterbücher in historischen Porträts; De Gruyter: Berlin, Germany, 2012. [CrossRef]
- Moore, J.W. Complete Encyclopaedia of Music; Oliver Ditson: Boston, MA, USA, 1854.
- Grove, G., Ed. A Dictionary of Music and Musicians (A.D. 1450–1889); Macmillan: London, UK, 1879–1889; Volumes 1–4.
- Riemann, H. Dictionary of Music, 4th ed.; Shedlock, J.S., Trans.; Augener: London, UK, 1908. Original work published 1882.
- Fétis, F.-J. Biographie universelle des musiciens et bibliographie générale de la musique, 2nd ed.; Firmin-Didot: Paris, France, 1860–1865.
- Desrat, G. Dictionnaire de la danse, historique, théorique, pratique et bibliographique; Librairies-Imprimeries Réunies: Paris, France, 1895.
- Meyers Konversations-Lexikon, 4th ed.; Bibliographisches Institut: Leipzig and Vienna, 1885–1890.
- Brockhaus, F.A.; Efron, I.A., Eds. Brockhaus and Efron Encyclopedic Dictionary; Semenovskaya Tipolitografiya: St. Petersburg, Russian Empire, 1890–1907.
- Vieira, E. Diccionario biographico de musicos portuguezes; Lambertini: Lisbon, Portugal, 1900.
- Wilkinson, M.D.; Dumontier, M.; Aalbersberg, I.J.; Appleton, G.; Axton, M.; Baak, A.; Blomberg, N.; Boiten, J.-W.; da Silva Santos, L.B.; Bourne, P.E.; et al. The FAIR Guiding Principles for Scientific Data Management and Stewardship. Scientific Data 2016, 3, 160018. [CrossRef]
- Bender, E.M.; Friedman, B. Data Statements for Natural Language Processing: Toward Mitigating System Bias and Enabling Better Science. Transactions of the Association for Computational Linguistics 2018, 6, 587–604. [CrossRef]
Table 2.
Selected-table data dictionary.
| Fields | Description |
|---|---|
| id | Deterministic opaque record identifier constructed from a legacy source key and a zero-padded source-level sequence; authoritative language and date values are stored in the metadata fields. |
| language; source_text_language | Working language label for the selected record and language actually processed. |
| headword; normalized_headword | Extracted entry heading and lower-case normalised heading used for overlap diagnostics. |
| category; subcategory | Ordered rule-derived music-performance labels. |
| definition; source_text | Compact excerpts from the candidate entry and processed source text. |
| example; etymology; semantic_tags; geographic_refs; historical_refs; cross_refs | Optional rule-derived or reserved fields for reuse and retrieval. |
| source_dictionary; author; publication_year; original_publication_year; volume; page | Bibliographic and processed-witness metadata. |
| source_file; source_line_start | Local processed-text provenance. |
| confidence_score | Uncalibrated OCR-surface heuristic in the range 0–1. |
| multilingual_alignment; article_type; notes | Alignment, sampling-frame, and free-text fields. |
Table 4.
Manual audit source allocation and selected-candidate precision.
| Source | Population | Sampled | Real + relevant | Precision |
|---|---|---|---|---|
| Riemann/Shedlock (1908) | 1295 | 62 | 60 | 96.8% |
| Moore (1854) | 1053 | 52 | 51 | 98.1% |
| Grove (1879–1889) | 733 | 37 | 35 | 94.6% |
| Desrat (1895) | 233 | 15 | 14 | 93.3% |
| Meyers (1885–1890) | 214 | 14 | 14 | 100.0% |
| Fétis/Pougin (1860–80) | 141 | 11 | 11 | 100.0% |
| Brockhaus-Efron (1896) | 90 | 8 | 8 | 100.0% |
| Vieira (1900) | 1 | 1 | 1 | 100.0% |
Table 5.
Manual audit quality-field distributions.
| Field | Value | Rows | Interpretation |
|---|---|---|---|
| Entry boundary | Complete | 199 | Candidate boundary adequate for audit decision. |
| Entry boundary | Fragment | 1 | Relevant entry, but excerpt is truncated. |
| OCR readability | Good | 200 | Surface text readable enough for the audit task. |
| Error flags | Truncated definitions | 1 | One selected record requires return to scan or fuller source text. |
| Decision confidence | High | 199 | Decision supported by boundary and OCR fields. |
| Decision confidence | Low | 1 | Same fragment row as above. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.