CorpusLD
Dual-Layer Academic Linked Data Engine & Deep Knowledge Graph Studio
An open-core extraction engine converting unstructured scientific papers, technical reports, and patents into Schema.org JSON-LD, W3C RDF Turtle, and Neo4j Cypher with live ROR, Wikidata, and Crossref authority linking.
| Architecture | Dual-Layer (Layer 1: 4-Tier Ingestion & 5-Agent Map-Reduce | Layer 2: Authority Resolvers & Multi-Format Graph Lake) |
|---|---|
| Parser Pipeline | 4-Tier Fallback (PyPDF → LlamaParse → Unstructured → Stateful Cross-Page Table Stitcher) |
| Authority Resolvers | Live ROR v2 (Research Institutions) + Wikidata QID & MeSH URIs + Crossref & OpenAlex DOI Reconciliation |
| Unit Ontology | Universal Unit Ontology (Standardized parsing for SI, Biomedical, Energy, and Compound Units) |
| Semantic Exports | Schema.org JSON-LD, W3C RDF Turtle (.ttl), Neo4j Cypher (.cql), BibTeX (.bib), RIS (.ris), CSL-JSON, Google Scholar Meta |
| Security & QA | 109 Passed Tests + SSRF Loopback Defense + Filename Path Traversal Protection + 100% Schema.org Validator Pass |