fig5

LLM-driven materials knowledge extraction: multimodal parsing, ontology, and agentic systems

Figure 5. Evolution of LLM-driven knowledge extraction paradigms in materials science. The input “literature corpus” denotes corpus-scale processing (103-104 papers, yielding ~1.6 × 105 nodes from a decade of literature[62]). Across all paradigms, the LLM operates in either (i) prompted, in-context inference without weight updates; or (ii) parameter fine-tuning or domain-adaptive pretraining[8,69]. The progression from schema-based to schema-free to agentic systems reflects increasing autonomy and decreasing direct human control; in practice, these paradigms are hybridized through schemas, retrieval, tools, and verification loops. LLM: Large language model; MOF: metal-organic framework; APIs: application programming interfaces.

Journal of Materials Informatics
ISSN 2770-372X (Online)
Follow Us

Portico

All published articles are preserved here permanently:

https://www.portico.org/publishers/oae/

Portico

All published articles are preserved here permanently:

https://www.portico.org/publishers/oae/