Submitted:
21 September 2026
Posted:
22 September 2026
You are already at the latest version
Abstract
Decades of thermoelectric-materials research are recorded in tens of thousands of papers, but the quantitative results — Seebeck coefficients, conductivities, figures of merit, doping strategies, synthesis conditions — remain locked in unstructured prose, inaccessible to systematic analysis or machine learning. We present a literature knowledge graph that extracts and structures this information at scale. Using large-language-model information extraction over the parsed full text of 10,158 thermoelectric papers, the graph captures 166,178 records across 13 entity types: 22,317 material mentions, 24,121 property measurements, 7,443 dopant/modification records, 32,306 scientific claims, 14,304 experimental conditions, and derived research-coverage, novelty, and failure signals, each attributed to its source paper by DOI and aligned to an OWL ontology. An independent audit on a stratified 300-paper sample gives an extraction precision of 0.97 and a recall of ~0.85 on content-bearing papers. The graph is released in five interoperable formats (CSV, JSON, JSON-LD, RDF/Turtle with 1.11 million triples, and GraphML with 166,178 nodes and 156,020 edges), enabling literature-scale trend analysis, research-gap discovery, and the fusion of literature evidence with computational materials features.
Keywords:
material-science
; thermoelectric materials
; computational material science
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.